Staff MLOps Engineer

Posted 9 Days Ago
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Expert/Leader
Healthtech • Information Technology
The Role
Leads the architecture, implementation, and scaling of enterprise MLOps platforms on AWS. Builds ML CI/CD pipelines, model deployment and monitoring systems, lifecycle governance, infrastructure automation, and secure, cost-efficient production environments. Provides Staff-level technical leadership, establishes platform standards, mentors engineers and data scientists, and aligns cross-team ML architecture and roadmap decisions. Requires strong experience with AWS, SageMaker, Python, ML frameworks, Docker, Kubernetes, distributed systems, and cloud architecture.
Summary Generated by Built In
Overview

We are looking for a seasoned Staff MLOps Engineer to lead the design, implementation, and scaling of enterprise-grade machine learning platforms on AWS. This role will focus on building reliable, secure, and cost-efficient MLOps systems that enable data scientists and engineers to deploy, monitor, and manage ML models in production. As a Staff Engineer, you will provide technical leadership, define best practices, and drive cross-team alignment on ML platform architecture. 


Duties & Responsibilities

Key Responsibilities 


MLOps Platform & Architecture 


  • Architect and own scalable MLOps platforms on AWS supporting model training, deployment, monitoring, and governance. 
  • Design and maintain end-to-end ML CI/CD pipelines, including data validation, model training, testing, approval, and deployment. 
  • Establish standards for model lifecycle management, experiment tracking, versioning, reproducibility, and rollback. 

Model Deployment & Monitoring 


  • Enable real-time, batch, and asynchronous model inference using AWS-native and container-based solutions. 
  • Implement monitoring for model performance, data drift, concept drift, and operational metrics. 
  • Ensure high availability, fault tolerance, and observability for production ML systems. 

AWS Cloud & Infrastructure 


  • Lead design and implementation using AWS services, including but not limited to: 
  • Amazon SageMaker (training, hosting, pipelines, feature store) 
  • EKS, ECS, EC2, Lambda for model serving and orchestration 
  • S3, Glue, Athena, Redshift for data storage and analytics 
  • CloudWatch, X-Ray for logging and monitoring 
  • Implement Infrastructure as Code (IaC) using Terraform or AWS CloudFormation. 
  • Optimize ML workloads for cost, performance, and scalability, including GPU/spot instance strategies. 



DevOps, Security & Compliance 

  • Build and maintain CI/CD pipelines using tools such as GitHub Actions, GitLab CI, Jenkins, or AWS CodePipeline. 
  • Enforce security best practices (IAM, VPC, encryption, secrets management). 
  • Support compliance, auditability, and governance requirements for ML systems. 

Technical Leadership & Collaboration 


  • Serve as a Staff-level technical leader, influencing MLOps architecture across multiple teams. 
  • Mentor engineers and data scientists on production ML best practices. 
  • Partner with Data Science, Data Engineering, Platform, and Product teams to align ML solutions with business goals. 
  • Contribute to the long-term ML platform roadmap and strategy. 

Skills Required

Mandatory Skills Required: 


  • 11–13 years of overall experience, with 5+ years in MLOps, ML Platform, or ML Infrastructure roles. 
  • Strong experience deploying and operating machine learning models in production on AWS. 
  • Proficiency in Python and experience with ML frameworks such as TensorFlow, PyTorch, Scikit-learn. 
  • Deep hands-on experience with Docker and Kubernetes (EKS). 
  • Strong understanding of Amazon SageMaker and its ecosystem. 
  • Experience with CI/CD systems and Git-based workflows. 
  • Solid background in distributed systems, system design, and cloud architecture. 

Preferred / Nice-to-Have Skills 

  • Experience with SageMaker Feature Store, Pipelines, Model Registry, or MLflow. 
  • Exposure to LLMOps / GenAI on AWS (Bedrock, custom LLM deployment, vector databases like OpenSearch, Pinecone). 
  • Experience with streaming and real-time pipelines (Kafka, Kinesis, Spark). 
  • Experience in regulated or high-scale environments (finance, healthcare, retail, etc.). 
  • AWS certifications (Solutions Architect, Machine Learning Specialty) are a plus. 

Soft Skills 

  • Strong ownership and decision-making ability at a Staff level. 
  • Excellent communication skills across engineering, data science, and leadership teams. 
  • Ability to balance short-term delivery with long-term platform vision. 
  • Passion for building reliable, scalable, and maintainable ML systems. 





Qualifications Required: 


  • Bachelor’s degree (B.A.) from four-year college or university, or equivalent combination of education and experience.  
  • 11–13 years of overall experience, with 5+ years in MLOps, ML Platform, or ML Infrastructure roles. 


About symplr: 


As a leader in healthcare operations solutions, we empower healthcare organizations to navigate the complexities of integrating critical business operations. Our customers are at the heart of everything we do, and they rely on our mission-critical systems to drive better operations and better outcomes. 

  

We are a remote-first company with employees working across the United States, India, and the Netherlands. Guided by values, we focus on teamwork, championing our customers, being rooted in action and outcomes, overcoming challenges, and leading through equality and integrity. Read more about symplr's culture and values at symplr.com/careers. 

Skills Required

  • 11-13 years of overall professional experience
  • 5+ years of experience in MLOps, ML Platform, or ML Infrastructure roles
  • Experience deploying and operating machine learning models in production on AWS
  • Proficiency in Python
  • Experience with TensorFlow, PyTorch, and Scikit-learn
  • Hands-on experience with Docker and Kubernetes, including Amazon EKS
  • Strong understanding of Amazon SageMaker and its ecosystem
  • Experience with CI/CD systems and Git-based workflows
  • Background in distributed systems, system design, and cloud architecture
  • Bachelor's degree or equivalent combination of education and experience
  • Experience with SageMaker Feature Store, Pipelines, Model Registry, or MLflow
  • Exposure to LLMOps or generative AI on AWS, including Bedrock, custom LLM deployment, or vector databases
  • Experience with Kafka, Kinesis, Spark, or other streaming and real-time pipelines
  • Experience in regulated or high-scale environments
  • AWS certification, such as Solutions Architect or Machine Learning Specialty
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Houston, TX
1,517 Employees

What We Do

Better operations. Better outcomes. As the leader in healthcare operations solutions, anchored in governance, risk management, and compliance, symplr enables enterprise customers to efficiently navigate the unique complexities of integrating critical business operations in healthcare. Our healthcare-specific software solutions and professional services provide value far beyond single, siloed solutions and enhance customers’ ability to achieve truly connected, integrated, enterprise-wide operational efficiencies. For over 30 years, healthcare organizations have trusted our expertise and depended on our provider data management, workforce and talent management, contract management, spend management, access management, and compliance, quality, safety solutions to help drive better operations for better outcomes.

Similar Jobs

CSC Logo CSC

Accountant

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Hybrid
2 Locations
8500 Employees

Cloudflare Logo Cloudflare

Systems Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4400 Employees

Dynatrace Logo Dynatrace

Senior Data Engineer

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
5600 Employees

The Aerospace Corporation Logo The Aerospace Corporation

Principal AI Engr

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
4600 Employees

Similar Companies Hiring

OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account