Software Engineer II (Data Engineering)

Posted 13 Days Ago
Bangalore, Bengaluru Urban, Karnataka, IND
Hybrid
Mid level
Biotech
The Role
Design, build, and own reliable, low-latency production data pipelines and APIs for a life-sciences catalog. Reduce technical debt, implement CDC, add observability and automated tests, translate business rules into durable pipeline logic, and manage deployments and incident response.
Summary Generated by Built In

About the Role

ZAGENO is hiring a Data Engineer to own the operational data pipelines powering our life sciences catalog. This is a high-ownership role: you’ll be the primary engineer responsible for the reliability, correctness, and scalability of the pipelines our operations teams depend on daily.The work spans production pipeline reliability, reducing accumulated technical debt in existing systems, building observability and testing infrastructure, and partnering with Data Science and Analytics on clean data delivery. You’ll work directly with CatalogOps and business stakeholders – turning evolving, often underspecified business rules into architectures that stay flexible without compromising data quality.

You’ll make architectural decisions, push back on requests that introduce heuristic debt, and own incident response end-to-end.

In this role you will:

  • Own reliability and performance of operational pipelines across our product catalog infrastructure.
  • Identify and reduce technical debt, replacing reactive patches with designed, testable logic.
  • Build and maintain low latency data APIs that serve downstream operational and analytics consumers.
  • Implement CDC patterns to keep catalog data synchronized across systems with minimal lag.
  • Build monitoring, observability, and automated testing so failures surface before stakeholders report them.
  • Design and implement unit standardization and master data logic at catalog scale.
  • Translate business requirements from non-technical stakeholders into durable pipeline logic.
  • Own code versioning, deployment, and incident response for your layer.
  • Leverage your expertise in the tech stack: BigQuery, Databricks, Spark/PySpark, AWS/GCP.

About you:

Required:

  • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines
  • Strong Python – data engineering, transformation logic, testing discipline
  • Strong SQL with ability to write correct queries, identify and refactor anti-patterns
  • Databricks, Delta Lake, Airflow for production orchestration
  • Experience with CDC patterns for real-time or near-real-time data synchronization
  • Experience building low latency APIs serving operational or analytical consumers
  • Test-driven development discipline – unit tests, integration tests, regression coverage as standard practice, not afterthought
  • Operates independently under ambiguity; designs systems to be maintained, not just to run

Preferred:

  • Kafka or equivalent event streaming platform experience
  • Experience with entity matching, deduplication, or master data management
  • Exposure to ML pipeline support in production
  • Familiarity with NLP techniques for entity resolution or text normalization (tokenization, similarity matching, named entity recognition)

What success looks like:

  • Engineers ship data products: Outputs are documented, versioned, and designed for reuse across Analytics, Data Science, and operational consumers.
  • Reliability: Pipelines run reliably with minimal manual intervention.
  • Performance: Data latency and downtime decrease measurably over time.
  • Data Quality: Analytics and Data Science teams receive clean, trustworthy data without ad hoc fixes.
  • Scalability: Infrastructure scales with volume growth without proportional cost increase.
  • Reduction of Debt: Technical debt in existing pipelines decreases measurably as heuristic patches are replaced with designed logic.
  • Synchronization: CDC-driven synchronization eliminates the class of cross-pipeline identity drift bugs.

Skills Required

  • 3+ years as a Data Engineer, including solo or primary ownership of production pipelines
  • Strong Python for data engineering, transformation logic, and testing discipline
  • Strong SQL with ability to write correct queries and refactor anti-patterns
  • Experience with Databricks and Delta Lake in production
  • Experience with Apache Airflow for production orchestration
  • Experience with CDC patterns for real-time or near-real-time data synchronization
  • Experience building low-latency data APIs serving operational or analytical consumers
  • Test-driven development discipline (unit, integration, regression tests)
  • Experience with Spark or PySpark
  • Experience with BigQuery and cloud platforms (AWS or GCP)
  • Ability to operate independently under ambiguity and design maintainable systems
  • Ability to own code versioning, deployment, and incident response for your layer
  • Experience with Kafka or equivalent event streaming platform
  • Experience with entity matching, deduplication, or master data management
  • Exposure to ML pipeline support in production
  • Familiarity with NLP techniques for entity resolution or text normalization
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boston, MA
147 Employees
Year Founded: 2015

What We Do

ZAGENO is on a mission to accelerate scientific innovation by streamlining biotech purchasing processes with its award-winning, first-of-its-kind e-commerce platform. With over 25 million product SKUs available from more than 5,000 unique product brands, ZAGENO makes online shopping for any research material convenient, efficient and reliable. The ZAGENO experience includes its Scientific Score, a best-in-class product rating system that offers unbiased, peer-reviewed ratings to support accurate purchasing decisions. Available on desktop, tablet, and mobile devices, ZAGENO makes biotech purchases easier than ever and is an ideal sales channel for suppliers and partners. Founded in 2015, the company is headquartered in the United States and has additional offices in Germany, and Poland. For more information, visit www.zageno.com.

Similar Jobs

In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
300 Employees

CrowdStrike Logo CrowdStrike

Senior Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
11000 Employees

Micron Technology Logo Micron Technology

Principal Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
45000 Employees

Micron Technology Logo Micron Technology

Staff Engineer

Artificial Intelligence • Hardware • Information Technology • Machine Learning
In-Office
2 Locations
45000 Employees

Similar Companies Hiring

Formation Bio Thumbnail
Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
New York, NY
150 Employees
SOPHiA GENETICS Thumbnail
Software • Healthtech • Biotech • Big Data • Artificial Intelligence
Boston, MA
450 Employees
Pfizer Thumbnail
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
New York, NY
121990 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account