About the Role
We are looking for an experienced Staff AI Engineer to help build the engineering foundations behind Blue Yonder's next generation of AI-powered products.
As part of Blue Yonder's Autonomy Labs, you will work alongside model-training teams, researchers, data engineers, and product engineers to develop agents that operate within real supply chain and retail workflows. You will build the services, APIs, evaluation systems, and development environments that allow us to train, test, deploy, and continuously improve these systems.
This is a hands-on software engineering and technical leadership role close to the model-development lifecycle. You do not need to be a model-training researcher, but you should understand how production behavior, evaluation results, traces, and data can be turned into better models and more reliable products.
Your Mission
Help turn increasingly capable models into dependable software systems. You will own important architectural decisions, build critical parts of the platform, and create the engineering loops through which agents are evaluated and improved after they begin operating in real workflows.
The challenge is not simply to make a model sound knowledgeable about supply chain. It is to build agents that can use tools, manage state, follow constraints, recover from failures, and complete valuable work reliably.
What You'll Do
- Architect and build robust software systems supporting the development, evaluation, deployment, and ongoing improvement of AI agents.
- Develop services and pipelines for collecting and processing agent traces, tool calls, workflow outcomes, user feedback, and other production signals.
- Build systems that evaluate and verify model behavior at scale using deterministic checks, model-based evaluation, statistical analysis, and human review where appropriate.
- Create reliable feedback loops between production systems and model-training teams, turning observed failures and successful behaviors into evaluation cases, regression tests, and candidate training data.
- Build training and evaluation environments that represent realistic business workflows through stable APIs, tools, simulations, resettable scenarios, and reproducible state.
- Implement the integrations and supporting services required for agents to interact with enterprise systems, domain data, and multi-step workflows.
- Design for reproducibility across models, prompts, tools, datasets, environments, and application versions so that changes can be measured with confidence.
- Partner with model-training teams to integrate new models, investigate behavior failures, define engineering requirements, and determine whether problems are best addressed through software, tools, data, evaluation, or model changes.
- Improve reliability and operability through observability, automated testing, failure recovery, safe rollout mechanisms, and clear launch criteria.
- Set a high engineering bar through system design, code review, documentation, mentoring, and pragmatic technical leadership across teams.
What We're Looking For
- Around eight or more years of software engineering experience, or equivalent depth, with a track record of designing and delivering complex production systems.
- Strong Python and backend engineering expertise, including API design, distributed systems, data processing, automated testing, and maintainable service architecture.
- Hands-on experience building LLM-powered applications, AI agents, machine learning platforms, or similarly complex systems.
- Practical understanding of AI evaluation, including test-case design, trace analysis, failure classification, regression testing, and measurement of nondeterministic systems.
- Strong understanding of software engineering fundamentals, including code quality, security, observability, CI/CD, and operational reliability.
- Experience with cloud platforms such as Azure, AWS, or GCP, along with containerization and modern deployment practices.
- The ability to work effectively with researchers and model-training engineers, translating experimental needs into well-designed software and production evidence into actionable feedback.
- Strong technical judgment, clear communication, and the ability to provide direction in ambiguous problem spaces without becoming a bottleneck for delivery.
Preferred Qualifications
- Familiarity with model-training or post-training approaches such as supervised fine-tuning, preference optimization, reinforcement learning, reward or verifier design, and checkpoint evaluation.
- Experience building model-training environments, simulators, evaluation harnesses, or agent development platforms.
- Understanding of training and evaluation data practices, including curation, synthetic data, provenance, versioning, quality control, and prevention of data leakage.
- Experience with supply chain, retail, or other enterprise environments involving complex workflows, permissions, business constraints, and human approval paths.
What Makes This Role Different
This role sits between model development and production engineering. You will not be asked to conduct research in isolation or simply wrap a model in an application. You will build the systems that make model behavior measurable, reproducible, improvable, and useful in real operational environments.
Our Values
If you want to know the heart of a company, take a look at their values. Ours unite us. They are what drive our success – and the success of our customers. Does your heart beat like ours? Find out here: Core Values
All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or protected veteran status.
Skills Required
- Around eight or more years of software engineering experience, or equivalent depth, designing and delivering complex production systems
- Strong Python and backend engineering expertise
- Experience with API design, distributed systems, data processing, automated testing, and maintainable service architecture
- Hands-on experience building LLM-powered applications, AI agents, machine learning platforms, or similarly complex systems
- Practical understanding of AI evaluation, including test-case design, trace analysis, failure classification, regression testing, and measurement of nondeterministic systems
- Strong software engineering fundamentals, including code quality, security, observability, CI/CD, and operational reliability
- Experience with cloud platforms such as Azure, AWS, or GCP, containerization, and modern deployment practices
- Ability to work effectively with researchers and model-training engineers
- Strong technical judgment, clear communication, and ability to provide direction in ambiguous problem spaces
- Familiarity with supervised fine-tuning, preference optimization, reinforcement learning, reward or verifier design, and checkpoint evaluation
- Experience building model-training environments, simulators, evaluation harnesses, or agent development platforms
- Understanding of training and evaluation data practices, including curation, synthetic data, provenance, versioning, quality control, and prevention of data leakage
- Experience with supply chain, retail, or other enterprise environments involving complex workflows, permissions, business constraints, and human approval paths
Blue Yonder Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Blue Yonder and has not been reviewed or approved by Blue Yonder.
-
Leave & Time Off Breadth — PTO is described as generous or “unlimited” in the U.S., alongside paid holidays, sick time, and two paid volunteer days. These policies are often highlighted as strengths that support work–life balance.
-
Flexible Benefits — Remote-work options and flexible arrangements are emphasized as part of the package. This flexibility is valued alongside compensation and can help offset middling pay for some roles.
-
Healthcare Strength — Medical, dental, and vision coverage are provided, with mental health/EAP support and HSA/FSA options referenced. These core coverages are portrayed as solid and comprehensive.
Blue Yonder Insights
What We Do
Blue Yonder is the world leader in digital supply chain and omni-channel commerce fulfillment. Our intelligent, end-to-end platform enables retailers, manufacturers and logistics providers to seamlessly predict, pivot and fulfill customer demand. With Blue Yonder, you can make more automated, profitable business decisions that deliver greater growth and re-imagined customer experiences. Blue Yonder - Fulfill your Potential Blue Yonder’s tagline “Fulfill Your Potential” reflects the company’s mission to empower every organization and person on the planet to fulfill their potential. Each day, our global teams of associates and business partners work together to accelerate global economic growth, increase sustainability and prosperity with a Sonoran Spirit.









