At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.
The PositionThe Kubernetes Infrastructure Reliability Engineer is a highly skilled expert responsible for solving complex business problems using advanced cloud native technologies. The engineer will build and maintain a Kubernetes-based infrastructure, enabling the modernization of business applications and processes.
This role combines software and systems engineering to optimize systems, increase efficiency, and eliminate operational work through automation.
You will be part of the global CaaS infrastructure team at a leading healthcare company, working with members across different regions. The team's mandate is to deliver, maintain, and continuously improve a highly available Kubernetes platform across hybrid cloud deployments, including on-premise data centers and public clouds like AWS. In this role, you will apply software engineering principles to operations to build and run massively distributed, fault-tolerant systems, focusing heavily on automation, security, and observability.
Job Responsibilities
Service Reliability and Optimization: Focus on capacity planning and launch reviews for services before they go live. Perform blameless postmortems and proactive identification of potential outages to foster iterative improvements
Accountability/Problem Solving: Resolves complex problems in a global Kubernetes-based infrastructure through in-depth evaluation of variable factors, including inter-organizational impact, balanced with effective consultative engagement of key stakeholders. Leads end-to-end design of infrastructure solutions and maintains component standards. Evaluates promising solutions via Proof of Concept (PoCs) and feasibility studies across multiple areas, and serves as an internal escalation point for major incidents
Stakeholder Management: Acts as a bridge between engineering and operations. Communicates and presents complex information and potential solutions to cross-functional teams and the business in non-technical terms. Represents the organization as a prime contact on initiatives and interacts with senior internal and external personnel. Uses deep knowledge to influence IT infrastructure vendor product evaluations and collaborates with multiple IT partners (e.g. Enterprise Architects, Solution Owners) to integrate feedback. Mentors and shares DevOps culture, guiding developers on how to create and deploy cloud-native applications
Impact/Strategy: Provides technical leadership and direction for small-to-medium sized initiatives (projects, lifecycle work, PoCs). Ensures solutions comply with Quality/Regulatory standards and that designs adhere to the organization’s Technical Architecture Framework (TAF) policies and directions. Assists in planning technology projects, estimating engineering resources, dependencies, risks and timelines for successful delivery
Business / Technical ability: Applies extensive cloud native technical expertise, acting as a recognized expert in Kubernetes and maintaining in-depth knowledge across related cloud native technologies (containers, AWS, etc.). Demonstrates a detailed understanding of how IT infrastructure impacts respective Roche business processes and outcomes
Qualifications
Education & Professional Experience
Without Degree: 4–7 years of relevant experience
Bachelor’s Degree: 2–5 years of relevant experience
Master’s Degree: 1–3 years of relevant experience
At least 1 year of experience working in a multinational environment; healthcare industry experience is a plus
Technical Skills
Kubernetes & Containers: Strong hands-on experience navigating, managing, and hardening Kubernetes clusters and containers, including knowledge of distributed storage. A Certified Kubernetes Administrator (CKA) certification is a strong plus. Knowledge of tools like Rancher or Portworx is beneficial
Infrastructure as Code (IaC): Hands-on experience delivering and managing infrastructure automation using tools like Ansible and Terraform
Scripting & Software Engineering: Proficiency in scripting and programming languages, primarily Python, Bash, or Go, including experience with test automation (e.g., pytest) and APIs deployment and management
CI/CD Tools: Expert knowledge of implementing software delivery pipelines using tools (e.g., Jenkins, Rundeck, or GitLab)
Systems & Networking: Strong understanding of Linux operating systems and core networking principles, including DNS, load balancing, firewalls, routing, and service meshes.
Observability: Experience configuring logging, metrics, and monitoring tools, specifically focusing on setting up alerts based on symptoms rather than waiting for system outages
Cloud Infrastructure: Experience with public cloud platforms, with a strong preference for AWS, specifically involving managed services for compute, networking, security, and identity (e.g., EKS, VPC, IAM)
General and Operational Knowledge
Proven experience applying best practices in an always-up, always-available service environment utilizing Scrum and Agile methodologies
Deep understanding of Technical Architecture Frameworks (TAF) and Quality/Regulatory compliance standards
Additional Qualifications
Excellent problem-solving skills, decision-making ability, and sound judgment
A strong team-oriented mindset with the ability to function independently with low supervision and navigate ambiguity.
Highly fluent oral and written English communication skills are required.
Ability to work across multiple time zones
Strong customer & delivery focus
*** 24/7 on-call rotation is required for this role****
#RDT2026
Who we are
A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Let’s build a healthier future, together.
Roche is an Equal Opportunity Employer.
Skills Required
- Hands-on experience with Kubernetes and containers, including cluster hardening and distributed storage
- Experience with Rancher or Portworx
- Infrastructure as Code using Ansible and Terraform
- Scripting/programming in Python, Bash, or Go; experience with test automation (e.g., pytest) and APIs
- CI/CD pipeline implementation experience (Jenkins, Rundeck, GitLab)
- Strong Linux systems knowledge and networking (DNS, load balancing, firewalls, routing, service meshes)
- Observability: configuring logging, metrics, monitoring and symptom-based alerting
- Cloud infrastructure experience, preferably AWS (EKS, VPC, IAM)
- At least 1 year experience working in a multinational environment
- Healthcare industry experience
- Understanding of Technical Architecture Frameworks (TAF) and Quality/Regulatory compliance
- Experience working in Scrum/Agile environments
- Highly fluent oral and written English
- Ability to work across multiple time zones
- Availability for 24/7 on-call rotation
- Certified Kubernetes Administrator (CKA) certification
Roche Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Roche and has not been reviewed or approved by Roche.
-
Retirement Support — U.S. materials describe a 401(k) with both matching and an additional company contribution, supported by formal plan documents and true‑up features. This structure is positioned as a standout element of the total package, particularly at Genentech.
-
Leave & Time Off Breadth — Time‑off provisions include substantial vacation, a year‑end shutdown, and a paid six‑week sabbatical after six years. These elements indicate a recharge‑oriented approach within the U.S. offering.
-
Healthcare Strength — Company materials emphasize comprehensive medical, dental, vision, and mental‑health resources alongside well‑being programs. Benefits pages consistently highlight breadth across core health coverage elements.
Roche Insights
What We Do
Roche is a global pioneer in pharmaceuticals and diagnostics focused on advancing science to improve people’s lives. The combined strengths of pharmaceuticals and diagnostics under one roof have made Roche the leader in personalised healthcare – a strategy that aims to fit the right treatment to each patient in the best way possible. Roche is the world’s largest biotech company, with truly differentiated medicines in oncology, immunology, infectious diseases, ophthalmology and diseases of the central nervous system. Roche is also the world leader in in vitro diagnostics and tissue-based cancer diagnostics, and a frontrunner in diabetes management. Founded in 1896, Roche continues to search for better ways to prevent, diagnose and treat diseases and make a sustainable contribution to society. The company also aims to improve patient access to medical innovations by working with all relevant stakeholders. Thirty medicines developed by Roche are included in the World Health Organization Model Lists of Essential Medicines, among them life-saving antibiotics, antimalarials and cancer medicines. Roche has been recognised as the Group Leader in sustainability within the Pharmaceuticals, Biotechnology & Life Sciences Industry ten years in a row by the Dow Jones Sustainability Indices (DJSI).







