You'll own our inference and model-serving infrastructure end to end. This isn't a research role. It's a build role: you're setting up and scaling the systems that let our agents actually run in production, fast and reliably, at increasing concurrency.
You report to Sofus and work closely with our ML and infra teams.
What You’ll Own
- Set up and scale inference/Ray Serve for ML and LLM model serving, integrated with our data analysis and agent workflows
- Scale agent GPU infrastructure for concurrency and efficiency across multiple agent workloads
- Optimize and improve the engine builder and model server that power scalable agent orchestration
What You Need
- Proven ability to build scalable ML/AI platforms from scratch, end-to-end, for production use cases. You've owned a zero-to-one build before, or can show you're capable of it
- Deep understanding of the inference stack: vLLM, KV cache, and the optimization layers underneath model serving
- Experience building distributed systems for AI/ML workloads at scale, connecting them to real product or vertical integrations
- 3+ years of relevant experience. We care about capability, not tenure
Nice to Have
- Ray / Ray Serve experience
- Familiarity with AIBrix
Why Join
- A rare chance to shape both company and product direction as an early team engineer
- Work alongside engineers and researchers from LinkedIn, Visa, Meta, and Branch
- Onsite culture in San Mateo, built for deep collaboration and high-velocity building
- Full benefits (medical, dental, vision, 401k)
- We sponsor H-1B visas and assist with immigration
Skills Required
- PhD in CS, ML, or related field OR MS with 4+ years of relevant experience
- Background in LLM optimization: inference efficiency, quantization, memory layout
- Ability to read, navigate, and debug LLM source code and underlying runtime libraries
- Comfortable in Rust and/or C++ at the systems level
- Strong Python skills
- Strong algorithmic fundamentals: data structures, complexity, distributed systems
- Hands-on experience with model serving infrastructure (vLLM, Baseten, Triton, etc.)
- Experience setting up and scaling ML pipelines end-to-end
What We Do
Zaimler is an AI infrastructure company that provides a platform for enterprise AI agents. It focuses on discovering domain knowledge, mapping relationships, and providing semantic understanding to enable autonomous agents to reason over fragmented enterprise data. Founded by industry veterans, the company aims to build the infrastructure layer for the agentic era, supporting real-time inference and precision at scale for enterprises in sectors like insurance, travel, and technology.









