Senior Machine Learning Developer
Location: Montreal, Quebec, Toronto, or Ontario
About Numa
Numa is building the platform to power AI-native dealerships, rearchitecting automotive service and sales with advanced AI agents that automate customer interactions, streamline operations, and reimagine how dealerships work. Numa integrates AI into every aspect of dealership functions—from rescuing customer calls and voicemails that generate more revenue, to reducing customer resolution times that drive overall customer satisfaction (CSI), to improving dealership team productivity and accountability. Numa has raised $50 million from leading investors (Google, Threshold, Costanoa, Mitsui, and Touring Capital).
The Role
We’re hiring a Senior Machine Learning Developer to build and ship ML/AI systems that interact with real customers thousands of times a day. Our voice agents book service appointments, rescue missed calls, and route callers through natural conversations. You’ll work across products and platforms. You’ll ship AI features including prompts, agents, tools, and production ML models, while building the evaluations and tooling that help teams ship with confidence. You’ll also contribute to our ML platform, including model serving, LLM infrastructure, and production observability. At Numa, we believe great ML is about more than building bigger models. It’s about knowing whether a change is good enough to ship. Our evaluation-first approach makes that measurable in CI and production.
What You’ll Do
- Build conversational AI systems for phone and SMS that understand customer needs, take action, and know when to act autonomously
- Develop tooling such as memory, knowledge graphs, and validated customization that help agents reason and adapt to dealership needs
- Train, evaluate, and deploy ML models using Ray Serve and Dagster for prediction, classification, ranking, and capacity forecasting, and keep them healthy in production
- Create offline and online evaluations, simulations, and CI gates that catch regressions and continuously measure quality in production
- Implement shared evaluation, observability, and LLM tooling that helps product squads ship faster and more reliably
- Raise the engineering bar through design and code reviews, technical writing, and mentorship
- Work autonomously when goals are clear and help create clarity when they are not
What You Bring
- 6+ years of software or machine learning engineering experience, with a track record of shipping ML or LLM powered systems to production
- Strong Python and solid software engineering fundamentals; you write code others can build on, test, and trust
- Hands-on experience building AI systems; training and deploying ML models (prediction, classification, ranking) and/or working with LLMs (prompting, tool use, agents, retrieval) with a real intuition for where they break
- A rigorous, evaluation-first instinct, you reach for measurement to tell a real improvement from a lucky one, and a real regression from noise
- Comfort operating with meaningful ambiguity in a product environment, and the judgment to know where to invest depth versus speed
- Ownership and communication; you can carry a piece of work from problem framing through shipped, monitored, and improved
Nice To Have
- Experience building or operating ML platform and tooling: evaluation frameworks, model serving, feature/prompt registries, or ML observability
- Familiarity with our stack or its neighbors: LiveKit, Ray Serve, Dagster, Vertex AI, GCP, Kubernetes, Pulumi, or provider APIs (Anthropic, OpenAI, Deepgram, ElevenLabs)
- Comfortable taking a model or agent from prototype to production and owning it there
- Proficient with real-time or streaming systems: voice, telephony (SIP/WebRTC), or low-latency inference
- Familiarity working in a startup or high-growth environment with evolving conditions
What Success Looks Like
- Your squad ships AI features faster and with more confidence because the evals and tooling around them are solid
- No AI change you own reaches customers without an evaluation standing behind it and regressions are caught before they ship
- You've turned at least one hard-won squad lesson into platform leverage the rest of the team now relies on
- Production issues in your area are understood, measured, and closed with statistical confidence that the fix actually worked
- Other engineers are better at building AI systems because of your reviews, your writing, and how you work
Why Join Numa
- We believe in everyone's growth
- Be part of an industry leader named the #1 fastest-growing AI Automotive company by Inc. 5000
- Represent category-defining AI technology that is transforming the automotive industry
- Join a high-growth organization with significant opportunities for career advancement
Compensation & Benefits
- Base Salary Range: $200,000 - $250,000 CAD
- Equity Packages
- Flexible PTO
- Fully Covered Group Insurance
At Numa, you’ll have the chance to represent game-changing technology in an industry that’s ready for innovation. If you love being in the field, thrive on solving problems in real time, and want to make an impact at a high-growth company, we’d love to hear from you.
Skills Required
- 6+ years of software engineering or machine learning engineering experience
- Track record of shipping machine learning or LLM-powered systems to production
- Strong Python skills
- Solid software engineering fundamentals, including writing testable and maintainable code
- Hands-on experience building AI systems, training and deploying ML models, and/or working with LLMs, prompting, tool use, agents, or retrieval
- Experience with evaluation-driven development and measuring model or AI system quality
- Ability to operate effectively with ambiguity in a product environment
- Ownership and communication skills, including carrying work from problem framing through production monitoring and improvement
- Experience with ML platform or tooling, such as evaluation frameworks, model serving, feature or prompt registries, or ML observability
- Familiarity with LiveKit, Ray Serve, Dagster, Vertex AI, GCP, Kubernetes, Pulumi, or AI provider APIs
- Experience taking models or agents from prototype to production and operating them
- Experience with real-time or streaming systems, voice, telephony, SIP, WebRTC, or low-latency inference
- Experience working in a startup or high-growth environment
What We Do
We fix why calls go uanswered, why advisors & reps spend their day on status updates instead of selling, why managers get blindsided by bad CSI reviews, and why every department operates in its own silo. Numa's AI agents answer and route every call, book appointments automatically, send proactive RO status updates, flag heat cases before CSI takes the hit, follow up on declined services and equity opportunities, and give managers full visibility into every customer communication across every department. All min one inbox and one shared context.
Why Work With Us
Built for People Who Move Fast. We hire drivers, not passengers. The people who thrive here bring their own urgency - they notice what's broken, ask why, and move toward the problem. We're a company that adapts. New information changes the plan. A better approach replaces the old one. We are energized by uncertainty, not rattled by it.









