Technical Lead, Senior AI Engineer - VonHalsky (m/f/n)

Posted 11 Days Ago
Be an Early Applicant
Kraków, Małopolskie, POL
Hybrid
Senior level
Logistics • Transportation
The Role
Own and advance the evaluations platform for Von Halsky, InPost’s conversational AI shopping assistant. Lead LLM-judge pipelines, evaluation datasets, Polish-language golden sets, root-cause analysis, anomaly detection, and eval-driven development. Define measurable quality standards with product and business stakeholders, set technical direction, review code, mentor engineers, and ensure releases are reproducibly evaluated and gated.
Summary Generated by Built In
Company Description

InPost Group is an innovative European out of home deliveries company, revolutionizing the way parcels are delivered to customers. With operations across several countries, our network of intelligent lockers provides customers with a fast, convenient, and secure delivery option. InPost Group is a publicly traded company, with a market capitalization of about $5 billion as of March 2023. With over 10,000 people worldwide, InPost Group is one of the largest out of home delivery providers in Europe, committed to providing sustainable and efficient delivery solutions to meet the evolving needs of customers in today's rapidly changing landscape. 

We are a team that builds the best FMCG e-commerce in Poland. We already have over 300,000 products in our offer, we work with over 300 sellers, and our mobile application has already been downloaded by almost 1.5 million users in Poland. Due to the fact that our product consists of many elements that must work together efficiently and be managed effectively, we are looking for experienced consumer-focused (Web / Mobile App products) engineering leaders to join us in that journey - heavily influence our future platform build, improve processes and help us deliver best customer experience in the market across all types of devices.  

Job Description

Why this role exists

Von Halsky is InPost's conversational AI shopping assistant, live in production and serving a growing share of our customers. What decides whether it wins is not the model but whether we can tell, at release cadence, that a change made conversations better. That is the Evaluations Platform, and we are hiring the engineer who takes it to the next level.

What you will own

  • The LLM-judge pipeline. Our release-gating judge over real and golden conversations, calibrated well enough that people act on its verdicts instead of arguing about it.
  • Eval datasets and the golden set. Real coverage across intents and categories, including the Polish-language coverage generic benchmarks do not give us.
  • Root-cause analysis on real conversations. Making our conversation-mining stack diagnostic rather than descriptive, on a stable issue taxonomy.
  • The "sus" detector. Abusive, adversarial and anomalous sessions, next to our guardrails and red-team work.
  • Evals-driven development. The eval comes before the feature, and writing it is as cheap as writing the code.
  • The interface to product and business. Vague asks in, measurable quality definitions out, and results stakeholders can act on.

What "leadership aspirations" means here concretely

This is a technical lead role, not a people-management role, and you will not carry line-management duties on day one. You will set and defend the technical direction for the platform, act as reviewer of record for the area, mentor other engineers, scope work with our PM and EM, present results to stakeholders, and hold the line against ad-hoc requests crowding out platform work.

Engineering management later, or a Staff-level hands-on track, are both paths we will build with you. Either way we need someone accountable for an area rather than for a ticket.

How we define success in this role

  • Judge pass rate is calibrated against human labels and is a number people trust and cite.
  • A stable, versioned issue taxonomy is live and week-over-week trends are comparable.
  • Every release is gated by an eval run the team can reproduce.
  • The golden set has documented coverage and named blind spots.
  • At least two engineers besides you can operate and extend the platform.

Qualifications

What we are looking for

Required

  • 5+ years building and running production software, with 2+ years on LLM-based systems that real users hit.
  • Strong engineering fundamentals, plus the habits that go with production ownership: testing, CI/CD, containers, observability, and working in cloud. We work primarily in Python.
  • Demonstrable experience evaluating generative systems, not only building them: LLM-as-judge, human-label calibration, inter-annotator agreement, regression suites, offline-versus-online divergence. You should have opinions about what makes an eval worthless.
  • AI engineering fundamentals. Prompt and context engineering as a discipline (context-window budgeting, structured outputs, failure-mode taxonomies); agentic primitives in production (tool use, multi-turn state, MCP, agent-to-agent integration patterns); and eval and LLM-observability tooling (LangFuse, Braintrust, Weave or equivalent, including things you built yourself).
  • Comfort with data at scale: SQL, working with a lake or warehouse, and building a metric someone else can reproduce.
  • Fluency with AI-assisted development tooling (Claude Code, Cursor, Copilot). We use it daily and expect it.
  • Ability to make a technical argument to a non-technical audience and be understood.
  • English B2 and Polish. Our users converse in Polish and you will read their conversations; judging quality you cannot read is not possible.

Nice to have

  • Harness and loop engineering. Building the scaffolding around models rather than only calling them: agent loops, retries and fallbacks, tool-call orchestration, deterministic replay, and the plumbing that makes a non-deterministic system testable.
  • Auto-improving systems. Closing the loop from production signal back into the product: mining failures into cases, using eval results to drive prompt, retrieval and routing changes, and automating the parts of that cycle that people do by hand today.
  • Adversarial robustness, jailbreak testing, red-teaming, or abuse and fraud detection.
  • E-commerce, search or recommendation domain experience.

Additional Information

What we offer

  • A product that has already been released to millions of users, with a real business case, not a lab pilot.
  • Direct access to frontier models at committed capacity across multiple providers, plus an open-source track we run ourselves.
  • A quality mandate with executive attention.
  • Ownership of a platform that is greenfield in practice inside a company with production traffic, which is the rarest combination in this market.
  • Hybrid working from Warsaw or Kraków, in a team growing fast enough that early hires shape how it works.
  • Fulfilling careers with a range of benefits for people and investing in providing training opportunities for their development. 
  • You will feel a part of the InPost community that makes an impact on sustainability, convenient deliveries, and the circular economy every day. 
  • Excellent working environment and flexible hours
  • We offer B2B type of contract

Skills Required

  • 5+ years building and running production software
  • 2+ years working on LLM-based systems used by real users
  • Strong engineering fundamentals, including testing, CI/CD, containers, observability, and cloud experience
  • Production software development experience primarily using Python
  • Experience evaluating generative systems, including LLM-as-judge, human-label calibration, inter-annotator agreement, regression suites, and offline-versus-online divergence
  • AI engineering experience with prompt and context engineering, structured outputs, failure-mode taxonomies, agentic primitives, tool use, multi-turn state, MCP, and agent-to-agent integration patterns
  • Experience with LLM evaluation and observability tooling such as LangFuse, Braintrust, Weave, or equivalent
  • Experience working with data at scale, including SQL, data lakes or warehouses, and reproducible metrics
  • Fluency with AI-assisted development tools such as Claude Code, Cursor, or Copilot
  • Ability to communicate technical arguments clearly to non-technical audiences
  • English proficiency at B2 level
  • Polish language proficiency
  • Experience building model harnesses and loops, including agent loops, retries, fallbacks, tool-call orchestration, deterministic replay, and testing infrastructure
  • Experience building auto-improving systems that use production signals and evaluation results to improve prompts, retrieval, or routing
  • Experience with adversarial robustness, jailbreak testing, red-teaming, abuse detection, or fraud detection
  • E-commerce, search, or recommendation experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Kraków
2,812 Employees
Year Founded: 2006

What We Do

InPost is the most successful operator of automated parcel lockers in Europe and also the one and only company in the world that is both a heavy operational user of APM machines as well as their manufacturer. InPost Parcel Lockers revolutionized the e-commerce delivery by providing a convenient way to send, collect and return parcels whenever the customer chooses, through a network of conveniently located and easy to use terminals.It is important for us to protect your personal data, our privacy policy can be found under the link: https://inpost.pl/ochrona-danych-osobowych

Similar Jobs

Pfizer Logo Pfizer

Director R&D EHS Program Lead

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
36 Locations
121990 Employees
177K-294K Annually

SailPoint Logo SailPoint

Consultant

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
2461 Employees

Capco Logo Capco

Business Analyst

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Ericsson Logo Ericsson

Senior Devops Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Kraków, Małopolskie, POL
88000 Employees

Similar Companies Hiring

Blissway Thumbnail
Computer Vision • Fintech • Hardware • Internet of Things • Machine Learning • Software • Transportation
Denver, CO
24 Employees
Toro TMS Thumbnail
Cloud • Enterprise Web • Sales • Software • Transportation
Chicago, IL
80 Employees
Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account