At Illumina, we are expanding access to genomic technology to realize health equity for billions of people around the world. Our efforts enable life-changing discoveries that are transforming human health through the early detection and diagnosis of diseases and new treatment options for patients.
Sequencing has moved genomics from single studies to population scale. Multi-million-sample cohorts, linked phenotypes and multi-omic data are now generated faster than most organizations can analyze them. Our team builds the data platform and the analytical applications that turn these assets into products: a cloud-native lakehouse for genomic and phenotypic data, large-scale statistical genetics pipelines that run reliably at population scale, and the APIs and user interfaces that let scientists ask questions of billions of records and get answers in seconds.
We are looking for a Sr. Staff Software Engineer for this effort in Singapore: someone who can take state-of-the-art analytical methods and turn them into scalable, high-performance, commercial software, and who can work fluently across bioinformatics scientists, cloud infrastructure engineers, and product management. You will work with a globally distributed team, prototyping quickly, demonstrating early, and iterating toward a releasable product.
ResponsibilitiesTurn advanced statistical genetics methods (for example, genome-wide and phenome-wide association testing for both common and rare variants) from prototypes into high-performance, scientifically accurate, commercially releasable products.
Design and implement analysis pipelines that scale to millions of samples on a cloud-native lakehouse, and establish the standardized, reusable design pattern for low-latency analysis on open table formats at petabyte scale.
Deliver the APIs, web UI and natural-language/agentic query interfaces through which scientific users explore association results and cohort/phenotype data interactively at scale.
Own the operability and unit economics of the platform: orchestrate, monitor, debug and cost-optimize workloads spanning millions of concurrent jobs, designing for fault tolerance and reproducible results.
Partner with bioinformatics scientists, cloud infrastructure teams and product management to translate scientific and customer needs into a product architecture and credible delivery plan aligned with the platform and data business strategy, applying a clear view of what distinguishes a commercial-grade offering from research and open-source community tooling.
Set technical direction and raise engineering standards through design review, mentorship and hands-on contribution, prototyping and iterating rapidly with a customer-first mindset across a globally distributed team.
Degree in Computer Science / Engineering / Bioinformatics / Mathematics or a related field.
Demonstrated experience designing and delivering large-scale, data-analysis production software in the cloud (AWS, Azure, or GCP), including containerization (Docker), orchestration, infrastructure automation, and CI/CD.
Strong experience with modern data lake / lakehouse architectures, including open table format internals rather than usage alone: Apache Iceberg or Delta Lake table specifications, catalog services, snapshot lifecycle, partitioning and pruning strategy, compaction, and schema evolution, over columnar storage (Parquet) with distributed query engines (for example Spark, Trino, Databricks).
Solid algorithms and systems engineering foundation, with proven ability to implement performance-critical code in a compiled language such as C/C++, including concurrency, memory management, vectorization and SIMD, and work with columnar in-memory formats and vectorized engines such as Apache Arrow, DuckDB, or Velox.
Proven track record of profiling and optimizing end-to-end systems for runtime, memory, I/O and cloud cost at large scale, and of designing for horizontal scalability.
Proficiency in Python for rapid prototyping, data analysis, and productionized tooling, with familiarity with the scientific and ML/DL library ecosystem.
Experience designing and operating distributed batch and workflow systems that manage very large numbers of concurrent jobs, including metrics, tracing and structured logging, retry, checkpointing and idempotency, and failure triage at million-task scale.
Experience with scientific workflow languages and engines (for example Nextflow, WDL, CWL) and with execution on cloud batch or Kubernetes.
Experience designing and delivering service APIs with an API-first approach: REST or gRPC interface design, versioning and backward compatibility, and multi-tenant access patterns; plus enough full-stack fluency to partner on (or build) modern web front ends and to hold strong opinions on UI/UX for data-heavy scientific applications.
Experience validating numerically and statistically sensitive software: concordance testing against reference implementations, calibration checks, benchmark suites, versioned reference data, and end-to-end reproducibility.
Proven technical leadership and influence at senior level, with strong verbal and written communication skills, and the ability to self-manage and manage interdisciplinary relationships.
Demonstrated experience using AI tooling to plan, manage and accelerate the full software development lifecycle, and to materially increase engineering productivity and quality.
Sr. Staff Software Engineer: Typically requires a minimum of 12 years of related experience with a Bachelor's degree; or 8 years and a Master's degree; or a PhD with 5 years experience; or equivalent experience.
Working knowledge of statistical genetics and population genomics: genome-wide and phenome-wide association studies, common and rare variant analysis (including gene-based burden and aggregate tests), mixed-model and whole-genome regression approaches, and the standard data formats and quality control practices of the field (VCF/gVCF, PLINK, BGEN, summary statistics, cohort and phenotype definition, ancestry and relatedness handling).
Hands-on experience with genomics-native storage and query stacks such as TileDB / TileDB-VCF, Hail, GLOW, or GenomicsDB.
Experience delivering commercial or clinical-grade genomics or life-science data products, including taking an internal or research tool through to a commercially released product, and a clear understanding of the differences in philosophy, obligations and trade-offs between commercial productization and academic, non-profit or research-community tooling.
Familiarity with techniques for low-latency query over very large result sets: precomputed summary indices, bitmap and zone maps, sketching, and tiered materialization.
Experience building agentic workflows and products, including tool/agent orchestration and MCP-style integrations, over enterprise data; and with machine learning, deep learning or foundation model training applied to genomic or biomedical data.
Experience implementing data governance controls: role-based and row/column-level access, audit trails, tenant isolation, and consent-driven restrictions on data use.
Experience with interactive data visualization for very large result sets, and with modern front-end frameworks.
Cloud cost engineering: cost-per-analysis modelling, spot and preemptible capacity strategy, storage tiering, and egress management.
Be curious, detail oriented, and analytical, with a proven ability to learn quickly.
Be customer-focused, team-oriented, and motivated, taking ownership of assigned tasks and of ambiguous problems.
Listed responsibilities are an essential, but not exhaustive, list of the usual duties associated with the position. Changes to individual responsibilities may occur due to business needs.
We are a company deeply rooted in belonging, promoting an inclusive environment where employees feel valued and empowered to contribute to our mission. Built on a strong foundation, Illumina has always prioritized openness, collaboration, and seeking alternative perspectives to propel innovation in genomics. We are proud to confirm a zero-net gap in pay, regardless of gender, ethnicity, or race. We also have several Employee Resource Groups (ERG) that deliver career development experiences, increase cultural awareness, and offer opportunities to engage in social responsibility. We are proud to be an equal opportunity employer committed to providing employment opportunity regardless of sex, race, creed, color, gender, religion, marital status, domestic partner status, age, national origin or ancestry, physical or mental disability, medical condition, sexual orientation, pregnancy, military or veteran status, citizenship status, and genetic information. Illumina conducts background checks on applicants for whom a conditional offer of employment has been made. Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable local, state, and federal laws. Background check results may potentially result in the withdrawal of a conditional offer of employment. The background check process and any decisions made as a result shall be made in accordance with all applicable local, state, and federal laws. Illumina prohibits the use of generative artificial intelligence (AI) in the application and interview process. If you require accommodation to complete the application or interview process, please contact [email protected]. To learn more, visit: https://www.dol.gov/ofccp/regs/compliance/posters/pdf/eeopost.pdf. The position will be posted until a final candidate is selected or the requisition has a sufficient number of qualified applicants. This role is not eligible for visa sponsorship.
Skills Required
- Degree in Computer Science, Engineering, Bioinformatics, Mathematics, or a related field.
- Large-scale data-analysis production software experience in AWS, Azure, or GCP.
- Experience with Docker, orchestration, infrastructure automation, and CI/CD.
- Experience designing modern data lake or lakehouse architectures using Apache Iceberg or Delta Lake, Parquet, and distributed query engines such as Spark, Trino, or Databricks.
- Strong algorithms and systems engineering skills, including performance-critical C or C++ development, concurrency, memory management, vectorization, SIMD, and columnar in-memory formats.
- Experience profiling and optimizing runtime, memory, I/O, cloud cost, and horizontal scalability at large scale.
- Proficiency in Python for prototyping, data analysis, and production tooling.
- Experience operating distributed batch and workflow systems with observability, retries, checkpointing, idempotency, and failure triage at million-task scale.
- Experience with Nextflow, WDL, CWL, cloud batch, or Kubernetes execution.
- Experience designing service APIs using REST or gRPC, including versioning, backward compatibility, and multi-tenant access patterns.
- Full-stack fluency and ability to build or partner on modern web front ends and data-heavy scientific user experiences.
- Experience validating numerically and statistically sensitive software through concordance testing, calibration checks, benchmarks, reference data, and reproducibility testing.
- Senior-level technical leadership, influence, communication, self-management, and interdisciplinary collaboration skills.
- Experience using AI tooling across the software development lifecycle to improve engineering productivity and quality.
- At least 12 years of related experience with a bachelor's degree, 8 years with a master's degree, 5 years with a PhD, or equivalent experience.
- Working knowledge of statistical genetics and population genomics.
- Experience with genomics-native storage and query stacks such as TileDB, TileDB-VCF, Hail, GLOW, or GenomicsDB.
- Experience delivering commercial or clinical-grade genomics or life-science data products.
- Experience with low-latency querying over very large result sets.
- Experience building agentic workflows, MCP-style integrations, and machine learning or foundation-model products for genomic or biomedical data.
- Experience implementing data governance controls, including role-based access, row or column-level access, audit trails, tenant isolation, and consent restrictions.
- Experience with interactive visualization of very large result sets and modern front-end frameworks.
- Experience with cloud cost engineering, including cost-per-analysis modeling, spot or preemptible capacity, storage tiering, and egress management.
Illumina Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Illumina and has not been reviewed or approved by Illumina.
-
Healthcare Strength — Health coverage includes comprehensive medical (HMO/PPO), dental, vision, mental health, life, and disability insurance with options such as FSAs/HSAs and wellness resources. Offerings are described as robust across regions with customization in certain locations.
-
Leave & Time Off Breadth — Time-off programs include flexible schedules and remote work, PTO or flexible time off, paid holidays and sick leave, and volunteer time. Company-wide shutdowns during summer and winter add additional paid time away.
-
Parental & Family Support — Family-focused programs include fertility assistance, adoption assistance, and reproductive health support alongside family medical leave. Paid parental leave is provided, with specifics varying by country and employment status.
Illumina Insights
What We Do
Illumina is an innovative technology and revolutionary assays aiming the analyze genetic variation and function.






