The Role
Lead the architecture and performance engineering of Nurix AI’s production infrastructure for real-time voice and chat agents. Design scalable, multi-cloud distributed systems supporting ASR, TTS, LLM orchestration, agentic RAG, and self-learning workflows. Establish reliability, observability, security, compliance, and low-latency standards while optimizing GPU and TPU utilization. Translate AI research into production platforms, mentor engineers, and evaluate advanced inference technologies.
Summary Generated by Built In
At Nurix AI, we are pioneering the Autopilot Enterprise. Our conversational AI agents handle workflows, drive outcomes, and deliver measurable impact for businesses. Born from the belief that enterprises need a new playbook, we build autonomous, multilingual agents capable of complex reasoning, contextual understanding, and end-to-end workflow ownership. Backed by $27.5M in funding from Accel, General Catalyst, and Meraki Labs, and led by Mukesh Bansal, we are India’s first scaled enterprise AI company, delivering cutting-edge AI solutions that integrate seamlessly into workflows across industries like Retail, Insurance, Education & Home Services. Join us in shaping the future of enterprise AI - where every interaction is smarter, faster, and human-like.
As Principal Engineer at Nurix AI, you will be the cornerstone of our technical infrastructure, enabling our AI agents to scale reliably and securely in production. You will design and oversee distributed systems that deliver low-latency, high-availability voice and chat AI, while meeting enterprise-grade security and compliance requirements. This is a hands-on leadership role focused on architecture, systems design, and performance engineering - ensuring that Nurix’s groundbreaking AI research translates into robust, real-world deployments.
Systems Architecture & Scalability
- Design and evolve the end-to-end infrastructure supporting ASR/TTS, LLM orchestration, Agentic RAG, and self-learning workflows.
- Architect low-latency pipelines for real-time conversational AI, ensuring sub-second response times across voice and chat.
- Build multi-cloud, distributed systems (AWS, GCP, Azure) with elastic scaling to handle spiky workloads.
Reliability & Performance Engineering
- Define and enforce SLAs around latency, uptime, and throughput for AI services.
- Drive observability, monitoring, and resilience strategies to handle failures gracefully.
- Optimize GPU/TPU utilization for cost-effective training and inference.
Security & Compliance
- Partner with InfoSec to embed security-by-design across all AI/ML workloads.
- Implement controls to protect sensitive enterprise data while meeting global compliance standards (SOC2, ISO 27001, GDPR, DPDP).
Collaboration & Leadership
- Work closely with the Head of AI to translate cutting-edge research into production-grade platforms.
- Provide technical mentorship to engineering teams, ensuring best practices in distributed systems and infra design.
- Evaluate and adopt emerging technologies (e.g., SSMs, inference optimizers like Triton, Riva, vLLM) to stay ahead of the curve.
- 10 - 15 years of experience in large-scale systems architecture, with at least 5 years in principal architect-level roles.
- Proven expertise in distributed systems, cloud-native architectures, and real-time pipelines.
- Hands-on experience with containerization, orchestration (Kubernetes), and microservices.
- Strong background in scalable ML infrastructure, including model serving, GPU/accelerator utilization, and CI/CD for ML.
- Demonstrated ability to architect systems with low latency (<300ms), high throughput, and enterprise reliability.
- Experience in conversational AI, speech systems, or real-time inference workloads.
- Deep knowledge of MLOps platforms (Kubeflow, MLflow, VertexAI, SageMaker).
- Familiarity with state-of-the-art inference optimization frameworks (e.g., Triton, Nvidia Riva, vLLM, SGLang).
- Open-source contributions or patents in distributed systems, infra, or ML tooling.
Skills Required
- 10–15 years of experience in large-scale systems architecture
- At least 5 years in principal architect-level roles
- Expertise in distributed systems, cloud-native architectures, and real-time pipelines
- Hands-on experience with containerization, Kubernetes orchestration, and microservices
- Experience with scalable ML infrastructure, model serving, GPU or accelerator utilization, and ML CI/CD
- Ability to architect systems with latency below 300 milliseconds, high throughput, and enterprise reliability
- Experience with conversational AI, speech systems, or real-time inference workloads
- Deep knowledge of Kubeflow, MLflow, Vertex AI, or SageMaker
- Familiarity with Triton, NVIDIA Riva, vLLM, or SGLang
- Open-source contributions or patents in distributed systems, infrastructure, or ML tooling
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Nurix AI (now NuPlay AI) develops enterprise conversational AI agents and workflow automation software. Its platform delivers voice and chat agents that engage customers, qualify leads, resolve support requests, and execute back-office operations across connected business systems. Designed for production use, the company helps enterprises automate end-to-end workflows across sales, support, retail, insurance, financial services, and other customer-facing and internal functions.


.jpeg)





