ETL Data Engineer | Onsite

Posted Yesterday
Hiring Remotely in United States
Remote
42K-147K Annually
Senior level
Agency • Information Technology
The Role
Design, build, and maintain scalable batch and streaming data pipelines using Python, SQL, Kafka, and GCP. Develop REST and GraphQL data services, CI/CD automation, observability, and data-quality frameworks. Integrate ML and NLP models into production workflows, including continuous retraining and real-time inference. Support structured and unstructured healthcare datasets, distributed systems, microservices, and reliable data infrastructure while collaborating with data science and MLOps teams.
Summary Generated by Built In

Responsibilities 

  • Design, build, and maintain scalable data pipelines to support analytics, ML, and operational reporting. 
  • Develop robust data ingestion, transformation, and integration workflows using Python, SQL, and modern data engineering frameworks. 
  • Build and maintain batch and streaming data pipelines leveraging technologies such as Kafka (or similar pub/sub tools)
  • Work with Google Cloud Platform (GCP) services, including Cloud Storage, Dataflow, Pub/Sub, BigQuery, Cloud Spanner and Cloud Functions 
  • Develop and manage data APIs and interfaces (REST and GraphQL) to enable high-performance data access across microservices. 
  • Implement CI/CD automation for data pipelines using GitHub Actions, Argo CD, or equivalent tools. 
  • Collaborate with Data Scientists and MLOps teams to integrate ML/NLP models into data pipelines and production workflows. 
  • Build and operationalize NLP data pipelines for structured and unstructured data sources (e.g., Rx claims, clinical documents). 
  • Enable continuous learning and modelretraining workflows using Vertex AI, Kubeflow, or similar GCP‑native tooling. 
  • Implement frameworks for observability and data quality, ensuring ML predictions, confidence scores, and fallback events are logged into data lakes or monitoring systems. 
  • Support distributed data systems and ensure reliability, performance, and scalability of data infrastructure. 

Required Qualifications 

  • 5+ years of experience building data pipelines or backend data workflows using Python, Java, or similar languages. 
  • 2+ years of experience designing REST/GraphQL data services or integrating data APIs. 
  • Hands‑on experience working with ML/AI model integration in production (e.g., Vertex AI Endpoints, TensorFlow Serving, ML REST APIs). 
  • Experience handling structured and unstructured datasets, including healthcare data (Rx claims, clinical documents, NLP text). 
  • Familiarity with the end-to-end ML lifecycle: data ingestion, feature engineering, training, deployment, and real‑time inference. 
  • 2+ years of experience with cloud platforms (GCP preferred; AWS or Azure acceptable). 
  • 2+ years working with streaming platforms like Kafka or equivalent. 
  • 2+ years of experience with databases (Postgres or similar relational systems). 
  • 2+ years of experience with CI/CD tools (GitHub Actions, Jenkins, Argo CD, etc.). 

Preferred Qualifications 

  • Direct, hands-on experience with Google Cloud Platform, especially BigQuery, Dataflow, GKE, Composer and Vertex AI. 
  • Knowledge of Kubernetes concepts and experience running data services or pipelines on GKE
  • Strong understanding of distributed systems, microservice patterns, and data‑centric system design. 
  • Experience using Vertex AI, Kubeflow, or other ML orchestration platforms for model training and serving. 
  • Knowledge of GenAI pipelines, LLM prompt workflows, and agent orchestration frameworks (e.g., LangChain, transformers). 
  • Experience deploying Python-based ML/NLP services into microservice ecosystems using REST, gRPC, or sidecar architectures. 
  • Domain experience in healthcare, claim adjudication, or Rx data processing

Education 

  • Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or equivalent experience 
    (High School Diploma + 4 years of relevant experience acceptable)

Compensation, Benefits and Duration

Minimum Compensation: USD 42,000
Maximum Compensation: USD 147,000
Compensation is based on actual experience and qualifications of the candidate. The above is a reasonable and a good faith estimate for the role.
Medical, vision, and dental benefits, 401k retirement plan, variable pay/incentives, paid time off, and paid holidays are available for full-time employees.
This position is available for independent contractors
No applications will be considered if received more than 120 days after the date of this post

Skills Required

  • 5+ years building data pipelines or backend data workflows using Python, Java, or similar languages
  • 2+ years designing REST or GraphQL data services or integrating data APIs
  • Hands-on experience integrating ML or AI models into production
  • Experience handling structured and unstructured datasets, including healthcare data
  • Familiarity with the end-to-end ML lifecycle, including ingestion, feature engineering, training, deployment, and real-time inference
  • 2+ years of experience with cloud platforms; GCP preferred
  • 2+ years working with Kafka or equivalent streaming platforms
  • 2+ years of experience with Postgres or similar relational databases
  • 2+ years of experience with CI/CD tools such as GitHub Actions, Jenkins, or Argo CD
  • Bachelor's degree in Computer Science, Data Engineering, Information Systems, or equivalent experience
  • Hands-on experience with GCP services including BigQuery, Dataflow, GKE, Composer, and Vertex AI
  • Knowledge of Kubernetes concepts and experience running data services or pipelines on GKE
  • Understanding of distributed systems, microservice patterns, and data-centric system design
  • Experience with Vertex AI, Kubeflow, or other ML orchestration platforms
  • Knowledge of GenAI pipelines, LLM prompt workflows, and agent orchestration frameworks
  • Experience deploying Python-based ML/NLP services using REST, gRPC, or sidecar architectures
  • Healthcare, claim adjudication, or Rx data processing experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
5,017 Employees
Year Founded: 2007

What We Do

Photon.com has emerged as one of the world’s largest and fastest-growing Digital Agencies. We work with 40% of the Fortune 100 on their Digital initiatives and are known for our ability to integrate Strategy Consulting, Creative Design, and Technology at scale. Please visit www.photon.com to learn more about us, how we work, and our customer case studies. Digital Transformation Starts Here.

Similar Jobs

Rapid7 Logo Rapid7

Sales Development Representative

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
Tampa, FL, USA
2400 Employees

Rapid7 Logo Rapid7

Senior Product Manager

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
Boston, MA, USA
2400 Employees
135K-183K Annually

Wipfli Logo Wipfli

Senior Account Executive

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
Minneapolis, MN, USA
2900 Employees
88K-118K Annually

Wipfli Logo Wipfli

Senior Account Executive

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
Radnor, PA, USA
2900 Employees

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account