Senior Engineer, Machine Learning Ops (India)

Posted 3 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Agency • Software • Consulting
Living at the intersection of creativity and technology.
The Role
Design, deploy, and operate scalable AI/ML infrastructure and CI/CD for production ML systems. Build and maintain cloud-native, containerized platforms supporting LLM agents, data pipelines, model serving, telemetry, and messaging; ensure performance, reliability, cost optimization, security, and collaboration with ML engineers and data scientists.
Summary Generated by Built In

We are looking for a highly skilled Senior Engineer – MLOps to build, deploy, and operate scalable, reliable AI/ML infrastructure that powers our next-generation AI platforms. This role sits at the intersection of Machine Learning, DevOps, and Cloud Engineering, with a strong focus on supporting LLM agent systems, data pipelines, cloud infrastructure, and production AI workloads.

The ideal candidate will have hands-on experience deploying and managing machine learning systems in cloud environments, automating infrastructure, building CI/CD pipelines, and ensuring the performance, scalability, and reliability of AI/ML platforms.

WHAT YOU'LL DO

  • Design, deploy, and manage scalable AI/ML infrastructure for production environments
  • Build and maintain infrastructure supporting LLM agent systems, machine learning models, and data pipelines
  • Deploy and operate ML/AI workloads across major cloud platforms such as AWS, GCP, or Azure
  • Develop and maintain containerized applications using Docker and serverless container platforms such as Cloud Run, ECS Fargate, or Azure Container Apps
  • Design and manage cloud databases, data warehouses, and storage solutions for AI applications
  • Build and optimize ETL/ELT pipelines to support machine learning workflows and analytics
  • Implement Infrastructure as Code (IaC) using Terraform or similar tools
  • Design and maintain CI/CD pipelines for automated model deployment and infrastructure provisioning
  • Implement monitoring, logging, alerting, and performance optimization for AI/ML systems
  • Manage event-driven architectures using messaging platforms such as Kafka, Pub/Sub, SNS/SQS, or similar technologies
  • Collaborate closely with AI/ML Engineers, Data Scientists, and Software Engineers to support model deployment and production operations
  • Optimize infrastructure for reliability, scalability, latency, and cost efficiency
  • Ensure security, governance, and operational best practices across cloud environments

WHAT YOU'LL NEED

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field
  • 6–10 years of experience in MLOps, DevOps, Cloud Engineering, or Infrastructure Engineering
  • Proven experience deploying and operating Machine Learning or AI systems in production
  • Strong expertise in cloud platforms including Google Cloud Platform (GCP), Amazon Web Services (AWS), or Microsoft Azure
  • Experience with Infrastructure as Code (Terraform preferred)
  • Strong Python programming skills
  • Excellent analytical, troubleshooting, and problem-solving skills
  • Strong communication skills with the ability to collaborate across technical and non-technical teams
  • Expert-level Python or Typescript
  • Very deep experience with at least one major cloud platform, including microservice creation and maintenance, databases and warehousing, and messaging/streaming infrastructure 
  • Strong experience with containerization
  • Strong experience with building-out telemetry and monitoring platforms
  • Strong experience with ML model serving
  • Expertise with infrastructure as code (IaC) such as Terraform
  • Experience building and maintaining CI/CD pipelines
  • Experience both implementing and advocating for MLOps best practices, model lifecycle management, and AI platform operations
  • Knowledge of Vector Databases, Agent Orchestration and Cost Optimization

NICE TO HAVE

  • Experience with LLM Agent Frameworks such as Google ADK, LangChain, LangGraph, or similar
  • Experience operating LLM/Generative AI workloads in production
  • Experience supporting enterprise-scale AI/ML platforms
  • Passion for building reliable, scalable, and automated infrastructure
  • Ability to troubleshoot complex distributed systems
  • Strong collaboration and stakeholder management skills
  • A continuous learning mindset and enthusiasm for emerging AI technologies

ABOUT US

Born in 2001, Code and Theory is a digital-first creative agency that sits at the center of creativity and technology. We pride ourselves on not only solving consumer and business problems, but also helping to establish new capabilities for our clients. With a global client roster of Fortune 100s and start-ups alike, we crave the hardest problems to solve. We have teams distributed across North America, South America, Europe, and Asia. The Code and Theory global network of agencies is growing and includes Kettle, Instrument, Left Field Labs, Create Group, Current, and TrueLogic.

Striving never to be pigeonholed, we work across every major category: from tech to CPG, financial services to travel & hospitality, government and education to media and publishing. We value the collaboration with our client partners, including but not limited to Adidas, Amazon, Con Edison, Diageo, EY, J.P. Morgan Chase, Lenovo, Marriott, Mars, Microsoft, Thomson Reuters, and TikTok.

The Code and Theory network is comprised of nearly 2,000 people with 50% engineers and 50% creative talent. We’re always on the lookout for smart, driven, and forward-thinking people to join our team.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
  • 6-10 years of experience in MLOps, DevOps, Cloud Engineering, or Infrastructure Engineering
  • Proven experience deploying and operating Machine Learning or AI systems in production
  • Strong expertise with cloud platforms (AWS, GCP, or Azure)
  • Infrastructure as Code experience (Terraform preferred)
  • Strong Python programming skills
  • Expert-level Python or Typescript
  • Strong experience with containerization and serverless container platforms (Docker, Cloud Run, ECS Fargate, Azure Container Apps)
  • Experience building and maintaining CI/CD pipelines for model and infrastructure deployment
  • Experience with monitoring, logging, alerting, and telemetry for AI/ML systems
  • Experience with ML model serving and production AI workloads
  • Experience with messaging/streaming platforms (Kafka, Pub/Sub, SNS/SQS)
  • Experience designing and managing cloud databases, data warehouses, and storage solutions
  • Experience building and optimizing ETL/ELT pipelines
  • Knowledge of Vector Databases, Agent Orchestration, and Cost Optimization
  • Experience with LLM agent frameworks (e.g., Google ADK, LangChain, LangGraph)
  • Experience operating LLM/Generative AI workloads in production or supporting enterprise-scale AI/ML platforms

Code and Theory Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Code and Theory and has not been reviewed or approved by Code and Theory.

  • Leave & Time Off Breadth Time-off policies include an “unlimited”/flexible PTO approach with paid holidays/sick time and seasonal early-close Fridays, offering notable flexibility. These options provide additional downtime beyond standard accrual models.
  • Healthcare Strength Core medical, dental, and vision coverage is part of the standard package, with employer-verified health plan information cited as current. Life and disability insurance further reinforce foundational coverage.
  • Retirement Support A 401(k) with employer match is included in the package, supporting long-term savings. Candidates are encouraged to confirm the match formula during the offer stage.

Code and Theory Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
445 Employees
Year Founded: 2001

What We Do

Code and Theory is a strategically driven, digital-first creative agency that lives at the intersection of creativity and technology. We solve consumer and business problems across the entire customer journey that flex to meet the ever-changing needs of consumer expectations. We put the user, their behaviors and needs, at the center of everything we do — from our proprietary research methodologies, to product development processes, to how we create brand, channel and messaging strategies. Our goal is simple: to solve our clients business problems. We bring big ideas to life by looking holistically at brand ecosystems where digital plays a prominent role in driving the consumer from first-touch through to conversion to relationship deepening over time. We identify gaps in the consumer journey and opportunities in culture that require products, services or communications to fill. We work across categories, ranging as far and wide as health care (Pfizer, Sanofi, Reach MD, Bioreference Laboratories) to financial services (JP Morgan Chase, Prudential, Morgan Stanley, First Data) to cpg (Mars, Unilever, Johnson & Johnson) to technology companies (Facebook, Xerox, Samsung, Comcast) to culture brands (adidas, H&M). And because our DNA is in publishing — we’ve embedded in over 135 newsrooms in the past decade — we bring unique expertise in understanding how content is created, distributed and optimized, including our work with CNN, NBC News, NBC Sports, and Bustle Digital Group. At Code and Theory, we strive to only be limited by our own ambition and creativity. We believe in pushing our creativity beyond the easy and obvious answers in order to deliver the solutions that are right for our clients, their businesses, and their consumers.

Gallery

Gallery

Similar Jobs

ZS Logo ZS

Engineering Manager

Artificial Intelligence • Healthtech • Professional Services • Analytics • Consulting
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
15000 Employees

ZS Logo ZS

Technology Leader - Life Sciences R&D

Artificial Intelligence • Healthtech • Professional Services • Analytics • Consulting
Hybrid
3 Locations
15000 Employees

ZS Logo ZS

Business Technology Solutions Manager - Salesforce Health Cloud

Artificial Intelligence • Healthtech • Professional Services • Analytics • Consulting
Hybrid
3 Locations
15000 Employees
10-14 Annually

ZS Logo ZS

PMO Associate - Visual and Content - Workflow Coordinator

Artificial Intelligence • Healthtech • Professional Services • Analytics • Consulting
Hybrid
2 Locations
15000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account