At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.
The PositionJob description
Our Applied AI Engineering Team is seeking a fullstack AI Data Engineer to join a newly formed, autonomous development squad dedicated to pioneering agentic, LLM-based solutions. Operating across the entire lifecycle—from initial ideation and rapid prototyping to production-grade deployment and ongoing operations—you will architect the resilient data infrastructure required to power next-generation AI. We are looking for an expert capable of orchestrating both structured and unstructured datasets, implementing high-performance vector databases, and managing real-time streams within cloud-native environments.
Description of the area
At Roche Digital Technology, we are advancing the boundaries of Applied AI. The Applied AI Engineering Team focuses on architecting, building, and operating high-value AI solutions and services to solve complex business challenges in healthcare.
In the 2026 tech landscape, we operate in small, highly autonomous agile teams (e.g., ~9 members) powered by advanced coding agents (like Claude Code) to develop and ship solutions faster than ever before. In this highly regulated environment, quality cannot be an afterthought. You will be the foundational pillar ensuring our rapidly developed agentic workflows and AI, GenAI, and agentic applications are safe, compliant, and robust before they reach the clinical or enterprise user.
Job Responsibilities
Generative AI Application Co-creation: Collaborate with AI engineers, data scientists, product owners, and other developers in Agile teams to integrate LLMs into scalable, robust, fair, and ethical end-user applications, focusing on user experience, relevance, and real-time performance
Data Infrastructure Development and Data Integration: Design and implement scalable, high-performance data pipelines for AI/GenAI applications, ensuring efficient data ingestion, transformation, storage and retrieval; integrate different databases, requiring understanding of data architectures / Domain data ecosystem
Vector Databases: work with vector databases (e.g., AWS OpenSearch, Azure AI Search) to facilitate scalable, high-speed similarity search and RAG for generative AI applications with high-dimensional data.
Graph Databases: Work with graph databases (e.g., Neo4j, AWS Neptune) to enable GraphRAG, support multi-hop logical reasoning for agentic workflows, and provide auditable explainability for enterprise AI decision-making.
Cloud-Based Data Engineering: Build and maintain cloud-based data solutions using AWS (OpenSearch, S3) or Azure (Azure AI Search, Azure Blob Storage)
Snowflake Implementation: Design and optimize data storage and processing using Snowflake for scalable, cloud-native analytics solutions
Data Processing & Transformation: Develop ETL/ELT pipelines to enable real-time and batch data processing
Support AI Model Workflows: Collaborate with AI/ML Engineers and Data Scientists to ensure seamless integration of data pipelines with AI finetuning, inference and training workflows
Performance Optimization: Optimize data storage, retrieval, and processing strategies for efficiency, scalability, and cost-effectiveness
Software Development Lifecycle: understand and leverage an agentic software development lifecycle (SDLC) in day to day work
Security & Compliance: Implement data governance, security best practices, and compliance measures aligned with Roche’s standards
Monitoring & Maintenance: Set up monitoring, alerting, and logging for data pipelines, ensuring high availability and reliability
Skills
Must have:
Experience: 7+ years in data engineering, preferably supporting AI/ML applications
Advanced Programming & SQL: Writing production-grade code in Python alongside highly optimized, complex SQL queries.
Advanced System Architecture & Modeling: Designing scalable, fault-tolerant ETL/ELT data pipelines and Lakehouse architectures (e.g., Snowflake).
Orchestration: Hands-on expertise with orchestration tools (like Airflow).
Data Engineering in AI: Developing Retrieval-Augmented Generation (RAG), AI systems powered by Vector Databases and/or LLM fine-tuning, and data preparation
Document Processing Proficiency: Extracting, transforming, and loading data from diverse file formats (PDF, DOCX, CSV, JSON, etc.), including automated parsing and information retrieval from unstructured and semi-structured documents
Version Control & DevOps: Hands-on experience with Git, CI/CD, containerization (Docker, Kubernetes), and Infrastructure as Code (Terraform, CloudFormation)
Problem Solving: Excellent analytical skills and the ability to tackle complex challenges with innovative solutions
Should have:
AWS Cloud Platforms: Hands-on experience with AWS (OpenSearch, S3, Lambda, AWS fundamentals)
GraphDB: Experience in building solutions with GraphDB
APIs & Microservices: Ability to design and integrate RESTful APIs for data exchange
Data Security & Governance: Understanding of encryption and role-based access controls
Working in an SDLC environment meeting regulatory requirements
Proficiency in best practices of software engineering
Agentic SDLC & Engineering Excellence: Leverage AI coding assistants and autonomous agents (e.g., Claude Code, Ona) daily to accelerate full-stack development and testing cycles. Conduct rigorous code reviews for both human-written and AI-generatd code.
Could have:
Regulatory Compliance: Proven experience in working within highly regulated industries.
Data Science & Classical Machine Learning: Practical background in Data Science, encompassing feature engineering, model training, and data preparation leveraging traditional ML techniques.
Distributed Data Processing: Hands-on expertise with big data frameworks (like Apache Spark or Flink).
Capabilities:
Problem-Solving Skills: Excellent analytical skills to tackle complex engineering and statistical challenges.
Ownership & Leadership: Deep sense of accountability, eager to define architectural patterns, and able to step into a Tech Lead role when necessary.
Consulting: Ability to work closely with stakeholders across the enterprise to consult on the technological approaches to their business problems.
Ethics: Strong understanding of biases, fairness, hallucination mitigation, and responsible AI deployment.
Qualifications
hold B.Sc., B.Eng., or higher, or equivalent in Computer Science, Data Engineering or related fields
have an interest in AI and stay up to date with the latest advancements in data engineering
be team-oriented, proactive, and collaborative
have strong analytical and problem-solving skills
have excellent verbal and written communication skills
be detail-oriented and highly organized
be willing to learn and expand their skill set
have the ability to work collaboratively in a fast-paced, dynamic environment
be able to communicate in English at the level of C1+
Located in Hyderabad, India, with working hours structured to capture the 'golden hours' of overlap with Central European Time (typically running through the IST evening).
#Hyd2026
Who we are
A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Let’s build a healthier future, together.
Roche is an Equal Opportunity Employer.
Skills Required
- 7+ years of experience in data engineering, preferably supporting AI or machine learning applications
- Production-grade Python programming and advanced SQL query development
- Experience designing scalable, fault-tolerant ETL/ELT pipelines and lakehouse architectures such as Snowflake
- Hands-on expertise with data orchestration tools such as Apache Airflow
- Experience developing RAG systems, vector database solutions, LLM fine-tuning workflows, or AI data preparation
- Experience extracting, transforming, and loading PDF, DOCX, CSV, JSON, and other structured or unstructured data
- Hands-on experience with Git, CI/CD, Docker, Kubernetes, and Infrastructure as Code using Terraform or CloudFormation
- Bachelor's degree in Computer Science, Data Engineering, or a related field, or equivalent experience
- Hands-on experience with AWS, including OpenSearch, S3, Lambda, and core AWS services
- Experience building solutions with graph databases
- Ability to design and integrate RESTful APIs and microservices
- Understanding of encryption, role-based access controls, data security, and governance
- Experience working in regulated software development lifecycle environments
- Proficiency in software engineering best practices and rigorous code review
- Experience using AI coding assistants or autonomous agents such as Claude Code or Ona
- Experience in regulated industries
- Practical data science or classical machine learning experience, including feature engineering and model training
- Hands-on experience with distributed data processing frameworks such as Apache Spark or Apache Flink
- Bachelor's degree or higher in a relevant field or equivalent
- English proficiency at C1 level or higher
- Located in Hyderabad, India and able to work during IST evening hours
Roche Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Roche and has not been reviewed or approved by Roche.
-
Retirement Support — U.S. materials describe a 401(k) with both matching and an additional company contribution, supported by formal plan documents and true‑up features. This structure is positioned as a standout element of the total package, particularly at Genentech.
-
Leave & Time Off Breadth — Time‑off provisions include substantial vacation, a year‑end shutdown, and a paid six‑week sabbatical after six years. These elements indicate a recharge‑oriented approach within the U.S. offering.
-
Healthcare Strength — Company materials emphasize comprehensive medical, dental, vision, and mental‑health resources alongside well‑being programs. Benefits pages consistently highlight breadth across core health coverage elements.
Roche Insights
What We Do
Roche is a global pioneer in pharmaceuticals and diagnostics focused on advancing science to improve people’s lives. The combined strengths of pharmaceuticals and diagnostics under one roof have made Roche the leader in personalised healthcare – a strategy that aims to fit the right treatment to each patient in the best way possible. Roche is the world’s largest biotech company, with truly differentiated medicines in oncology, immunology, infectious diseases, ophthalmology and diseases of the central nervous system. Roche is also the world leader in in vitro diagnostics and tissue-based cancer diagnostics, and a frontrunner in diabetes management. Founded in 1896, Roche continues to search for better ways to prevent, diagnose and treat diseases and make a sustainable contribution to society. The company also aims to improve patient access to medical innovations by working with all relevant stakeholders. Thirty medicines developed by Roche are included in the World Health Organization Model Lists of Essential Medicines, among them life-saving antibiotics, antimalarials and cancer medicines. Roche has been recognised as the Group Leader in sustainability within the Pharmaceuticals, Biotechnology & Life Sciences Industry ten years in a row by the Dow Jones Sustainability Indices (DJSI).
.jpeg)





