AI/MLOps SRE Lead Engineer

Posted 3 Days Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
Biotech • Pharmaceutical
The Role
Leads reliability, scalability, observability, and automation for enterprise AI, ML, LLM, and cloud platforms. Designs and operates multi-cloud ML infrastructure, MLOps workflows, CI/CD pipelines, Infrastructure as Code, ChatOps, and self-healing systems. Establishes reliability standards, implements AI-powered monitoring and remediation, supports production AI agents and RAG platforms, evaluates emerging technologies, manages technical debt and risks, and mentors engineering teams across SRE, MLOps, and cloud platform strategy.
Summary Generated by Built In

Build our future together

Regeneron is founded on the belief that the right idea, combined with the right team, can lead to significant transformations. Our growing global network is dedicated to inventing, developing, and commercializing medicines that change lives for those with serious diseases. In doing so, we are pioneering innovative approaches to science, manufacturing, and commercialization, as well as redefining our understanding of health.


At Regeneron Digital & Technology, we are expanding our AI and Platform Engineering capabilities to support next-generation intelligent systems, machine learning platforms, and cloud-native technologies. We are seeking an AI/MLOps SRE Lead Engineer to drive reliability, scalability, observability, and operational excellence across our AI, ML, and cloud ecosystem. This role will lead the design and operation of resilient platforms supporting machine learning workloads, LLMs, AI Agents, and enterprise-scale automation while enabling engineering teams to innovate with speed and confidence.


When & Where

Hyderabad (Hybrid)


Discover your role

  • Drive service reliability, availability, and performance across multi-cloud environments by establishing SLOs, SLIs, error budgets, and reliability standard methodologies.
  • Design, build, and operate enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI.
  • Develop AI-powered observability capabilities using anomaly detection, predictive analytics, and automated remediation to proactively identify and resolve operational issues.
  • Lead the implementation, monitoring, and optimization of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, scalability, and operational excellence.
  • Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation solutions to improve engineering productivity and platform resilience.
  • Architect enterprise ChatOps solutions integrating operational events, observability platforms, AI workflows, and automated remediation capabilities.
  • Partner closely with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions.
  • Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value.
  • Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity.
  • Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, cloud platform strategy, and AI-enabled operations across the organization.

This role requires

  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related field; Master's degree preferred.
  • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology subject areas with enterprise-scale delivery experience.
  • Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, and Azure.
  • Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI.
  • Proven experience implementing end-to-end ML workflows including model training, experiment tracking, deployment, monitoring, and pipeline orchestration.
  • Hands-on experience applying machine learning techniques such as anomaly detection, predictive analytics, time-series modelling, and operational intelligence within enterprise environments.
  • Advanced proficiency with Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices.
  • Strong programming and scripting skills in Python, Go, Bash, or similar languages.
  • Experience designing enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, logging, and monitoring platforms.
  • Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities.
  • Proven experience designing and implementing enterprise ChatOps solutions, including operational workflow automation and AI-enabled integrations.
  • Strong ability to identify technical debt, assess platform risks, influence technical strategy, and drive modernization initiatives.
  • Experience using AI tools, AI Agents, and LLM-powered assistants to improve engineering operations, incident management, and developer productivity.
  • Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred.
  • Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred.
  • Knowledge of cloud cost optimization, policy-as-code, compliance automation, FinOps practices, and multi-cloud governance preferred.

Does this sound like you? Apply now to take your first step towards living the Regeneron Way! We are committed to building a workplace with an inclusive culture. Regeneron is an equal opportunity employer and all  qualified applicants will receive consideration for employment without regard to race, color, religion or belief (or lack thereof), sex, sexual orientation, gender identity or expression, gender reassignment, marital or civil partnership status, civil status, pregnancy or parental status, age, disability, nationality, citizenship status, ethnic or national origin, membership of the Traveler community, familial status, genetic information, military or veteran status, or any other characteristic protected under applicable law. Where required, we will provide reasonable accommodation to applicants with known disabilities or chronic illnesses during the recruitment process, unless such accommodation would impose undue hardship.


Where necessary, we disclose salary ranges for roles in all countries in which we operate.  The final offer will be determined within the relevant range based on the country of employment, specific role level, and your skills and experience. In some countries, collective bargaining agreements (CBAs) may apply and influence certain elements of pay or benefits.  Regeneron offers a competitive and comprehensive total rewards package which may include, depending on country and role: annual bonuses or other incentive plans, equity awards, pension or retirement benefits, 401(k) company match, health and wellness programs, fitness centers, insurance benefits (e.g. medical, dental, vision, life and disability), paid time off, and family support benefits. For additional information about Regeneron benefits in the U.S., please visit https://careers.regeneron.com/en/working-at-regeneron/total-rewards/. For other locations, additional information will be provided during the recruitment process.  If you have any questions, please speak with your recruiter. 


Please be advised that at Regeneron, we believe we do our best work when we are together. For that reason, many roles are required to be performed on‑site. Please speak with your recruiter and hiring manager for more information about on‑site expectations for your role and location.


As part of the recruitment process, certain background checks may be conducted in accordance with the laws of the country where the position is based. The purpose of such checks is to verify certain information prior to the commencement of employment such as identity, right to work and educational qualifications.


For jobs in Canada: this posting is for an existing position.

Skills Required

  • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related field
  • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology areas
  • Experience delivering enterprise-scale technology solutions
  • Hands-on experience with two or more major cloud platforms, including AWS, GCP, and Azure
  • Deep expertise with Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI
  • Experience implementing end-to-end ML workflows, including training, experiment tracking, deployment, monitoring, and orchestration
  • Experience applying anomaly detection, predictive analytics, time-series modeling, and operational intelligence
  • Advanced proficiency with Terraform, Pulumi, AWS CDK, Infrastructure as Code, and CI/CD automation
  • Strong programming or scripting skills in Python, Go, Bash, or similar languages
  • Experience designing observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, logging, and monitoring platforms
  • Experience with automated remediation and AI-assisted operational capabilities
  • Experience designing and implementing enterprise ChatOps solutions and operational workflow automation
  • Ability to identify technical debt, assess platform risks, influence strategy, and drive modernization
  • Experience using AI tools, AI Agents, and LLM-powered assistants for engineering operations and incident management
  • Master's degree
  • Experience with Kubernetes and EKS, GKE, or AKS
  • Familiarity with Kubeflow, Feast, MLflow, LangSmith, RAGAS, Evidently AI, or Weights & Biases
  • Knowledge of cloud cost optimization, policy-as-code, compliance automation, FinOps, and multi-cloud governance

Regeneron Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Regeneron and has not been reviewed or approved by Regeneron.

  • Healthcare Strength — Medical, dental, and vision coverage is positioned as comprehensive, with Regeneron prescription drugs covered at 100% for those enrolled in the medical plan. Mental health support is also emphasized through EAP access and tools like Talkspace and the Journey app.
  • Equity Value & Accessibility — Stock grants are described as available to all employees, strengthening the overall total-rewards package beyond base pay. Long-term incentives and stock-related rewards are repeatedly framed as meaningful components of compensation.
  • Parental & Family Support — Paid parental leave is paired with fertility/adoption assistance and childcare-related support such as discounts and nanny services. Additional family-oriented resources extend to elder care, pet care, and education support like college coaching and tutoring.

Regeneron Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tarrytown, NY
15,000 Employees
Year Founded: 1988

What We Do

At Regeneron we believe that when the right idea finds the right team, powerful change is possible. As we work across our expanding global network to invent, develop and commercialize life-transforming medicines for people with serious diseases, we’re establishing new ways to think about science, manufacturing and commercialization. And new ways to think about health. Connect with us so we can learn more about you, and you can learn more about our biopharmaceutical medicines. And join us, as we build a future we believe in. Please visit www.regeneron.com/social-media-terms for information on how to engage with us on social media. An important note about privacy: Regeneron is committed to your privacy and will not ask for sensitive personal information such as social security number, date of birth or bank account details via email or social media.

Similar Jobs

Regeneron Logo Regeneron

Site Reliability Engineer

Biotech • Pharmaceutical
In-Office
Hyderabad, Telangana, IND
15000 Employees

Regeneron Logo Regeneron

Site Reliability Engineer

Biotech • Pharmaceutical
In-Office
Hyderabad, Telangana, IND
15000 Employees

Optum Logo Optum

Software Engineering Lead-salesforce FSL

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

SOPHiA GENETICS Thumbnail
Software • Healthtech • Biotech • Big Data • Artificial Intelligence
Boston, MA
450 Employees
Pfizer Thumbnail
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
New York, NY
121990 Employees
Cencora Thumbnail
Healthtech • Logistics • Pharmaceutical
Conshohocken, PA
51000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account