Senior AI Engineer (f/m/d) - Remote in Germany

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Bremen, DEU
In-Office or Remote
77K-97K Annually
Senior level
3D Printing • Aerospace • Artificial Intelligence • Automotive • Software • Automation • Manufacturing
The Role
Lead development and ownership of end-to-end evaluation and agentic systems. Build agent graphs, create golden datasets, implement LLM-as-judge pipelines, CI gating, reliability fixes, and production tracing to ensure trustworthy agent deployments.
Summary Generated by Built In

Most AI teams talk about evaluation. At Synera, you'd own it — end-to-end, in production, for two real agentic products used by companies like BMW, Airbus, and NASA.


WHAT YOU WILL DO

You’ll join the Agentic Ants team as our second native AI engineer. You’ll build and extend our agentic systems alongside the rest of the team — new tools, sub-agents, prompt iterations — and co-own the evaluation framework that gates every change: golden datasets, LLM-as-judge pipelines, regression suites, production-trace mining. This isn’t about adding scores to a notebook — it’s about shipping agents the team can trust, and proving it.

You’ll work across both Synera MAS and Synera Assistant, picking up reliability concerns (error handling, retries) as they show up in your work. Partnering with Ruben and the wider engineering team, you’ll also help the team go deeper on customer insights data.

 
👉 What a week at Synera could look like:
  • Monday: Kick off sprint planning with the Agentic Ants team, review Langfuse traces from the weekend, and flag any new failure modes worth triaging.

  • Tuesday: Work on the golden dataset for the supervisor routing surface — curating examples, versioning the set, and writing evaluators with Ahmed.

  • Wednesday: Join a cross-team sync with QA and product to align on new eval coverage for an upcoming agent feature, then push a CI integration so eval regressions block the next PR.

  • Thursday: Pair with the AI and software engineers on extending our agentic system — a new tool, a routing tweak, or a prompt iteration — then write the eval that gates the change before it ships.

  • Friday: Review a calibrated LLM-as-judge output alongside human labels, refine the rubric, and share findings in the eng review.

⚡ In 6 months: 
  • The evaluation framework is live with golden datasets across multiple agent surfaces, CI gates are blocking on regressions, and the team actually trusts the results.

  • You’ve shipped multiple meaningful changes to the agent graphs — new tools, sub-agents, or routing improvements — that are measurably better in production.

⚡ In 12 months: 
  • Online evaluation is a habit — production traces feed continuously back into datasets, and the improvement loop is running without manual intervention.

  • You co-own at least one agent surface with the team and have become the go-to person for evaluation methodology at Synera. The combination has measurably improved product quality.

Here is our team at the summer event - join us at the next one! 👋

🌱 WHAT'S IN IT FOR YOU?

We believe in transparent conversations about compensation from the start. For this role, our planned salary ranges are:

  • Experienced Level: 77,000 - 97,000 EUR annually

We determine the level we hire you for based on your experience, the scope of responsibilities you'll take on, and the impact you can drive. While we typically hire within these bands, we're open to some flexibility for candidates who bring exceptional value to the role.

Check out this page to learn how we approach salary, career growth, and creating an environment where everyone can shine.

What we offer beyond salary:
  • Flexible working: you decide when & where to work (as long as you have a residency in Germany).

  • Flexible public holidays: swap days off according to your values and beliefs!

  • Home office setup support + access to our office in Bremen.

  • Personal development budget of €2,000 to attend conferences and trainings or buy interesting books to improve in an area of your choice.

  • We don't count your vacation days, as we trust all our team members to decide what's best for them and the company.

  • Prefer two wheels over four? 🚲 We’ve got you covered with JobRad.

  • To support your personal and professional well-being, we offer company fitness with Wellpass and mental health platform nilo.

  • Regular team events, virtual coffee breaks, and spontaneous afterworks. We also get together as a whole company for 2-3 day off-sites twice a year! 🌈

🚀 WHAT YOU NEED TO SUCCEED

Even if you don't meet these criteria perfectly but believe you have lots to bring to the role, we encourage you to apply. We know it's tough, but please keep in mind that you don't have to match all the listed requirements exactly to be considered for this role. 

  • You resonate with Synera's Core Values - they're central to how we work, and we'll explore them together in your first interview.

  • You’ve designed and shipped agent graphs in production with LangGraph or equivalent — supervisor / sub-agent patterns, tool design, prompt iteration.

  • You’ve built or meaningfully contributed to an LLM evaluation pipeline — LLM-as-judge design, calibration against human labels, dataset versioning, and the statistical reasoning behind it (CIs, sample sizes, false positives).

  • You write production-grade Python (FastAPI, PostgreSQL).

  • You’ve worked with at least two of Anthropic, OpenAI/Azure, Bedrock, or Vertex AI.

  • You’re familiar with Langfuse, LangSmith, or similar tracing tools.

  • You handle reliability when it shows up in your work — retries, error handling, graceful degradation. SLO thinking is a plus.

  • You communicate clearly and push back when you disagree.

P.S. Synera is a place where everyone can grow. So, however you identify and whatever background you bring with you, please apply if this is a role that would make you excited to come to work every day, and be prepared to share with us how your perspective will bring something unique and valuable to our Agentic Ants team. 

Skills Required

  • Residency in Germany
  • Designed and shipped agent graphs in production with LangGraph or equivalent (supervisor/sub-agent patterns, tool design, prompt iteration)
  • Built or meaningfully contributed to an LLM evaluation pipeline (LLM-as-judge design, calibration, dataset versioning, statistical reasoning)
  • Production-grade Python experience (FastAPI, PostgreSQL)
  • Experience with at least two: Anthropic, OpenAI/Azure, AWS Bedrock, or Vertex AI
  • Familiarity with Langfuse, LangSmith, or similar tracing tools
  • Practical reliability engineering: retries, error handling, graceful degradation (SLO thinking is a plus)
  • Clear communication and willingness to push back when needed
  • SLO thinking and formal reliability/design-for-SLO approaches
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bremen
65 Employees
Year Founded: 2025

Similar Jobs

CrowdStrike Logo CrowdStrike

Sr. Intelligence Analyst, Recon+ (Remote, GBR)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
5 Locations
11000 Employees

Drata Logo Drata

Corporate Counsel

Security • Software • Cybersecurity • Automation
Remote
26 Locations
600 Employees
111K-137K Annually

SEON Logo SEON

Infrastructure Engineer

Artificial Intelligence • Cybersecurity
Remote
27 Locations
415 Employees

Akamai Technologies Logo Akamai Technologies

Solutions Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees

Similar Companies Hiring

Outpost Space Thumbnail
Aerospace • Defense
US
24 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account