Data Engineer

Posted 2 Days Ago
Houston, TX, USA
In-Office
103K-170K Annually
Junior
Greentech • Energy • Utilities • Renewable Energy
The Role
Design, build, and operate scalable batch and real-time data pipelines on Databricks/Azure, implement medallion architecture with Delta Lake, ensure data quality and governance via Unity Catalog, support CI/CD and production operations, and collaborate with stakeholders to deliver analytics-ready datasets for BI and ML use cases.
Summary Generated by Built In

Description

  

Fervo is building the most cost-effective, repeatable geothermal power plants in the world. Scaling that mission depends on a trustworthy, well-governed data foundation that turns raw sensor signals, drilling and completions records, and power plant telemetry into reliable, decision-ready information. The Data Engineer, within the Data & AI team, designs, builds, and operates the pipelines, models, and platforms that move data from the field to the people and systems that act on it — engineers, operators, data scientists, and the analytical and agentic applications built on top.

The Data Engineer owns data products end to end — from ingestion and modeling, through quality, governance, and serving, to monitoring in production. Working across Data Science, AI Engineering, IT Infrastructure, domain SMEs, and business stakeholders, this role establishes reusable patterns for real-time and batch processing, IoT/historian integration, data quality and entity linkage, and self-service analytics on our Azure, Databricks, and Snowflake stack. Success requires strong hands-on engineering depth in distributed data processing, sound data modeling and architecture judgment, and pragmatism about what to ship versus what to defer.

Requirements

Responsibilities

Data Pipeline & Platform Engineering

  • Design, build, and operate scalable batch and real-time/streaming data pipelines on Databricks and Azure Data Factory, landing data in Azure Data Lake Storage (ADLS) and Snowflake
  • Implement the medallion (bronze/silver/gold) architecture using Delta Lake and Delta Live Tables, with reliable incremental processing, schema evolution, and change data capture
  • Build and tune Apache Spark jobs (PySpark/Spark SQL) for large-scale, parallel data processing — partitioning, shuffles, caching, broadcast joins, and cost/performance optimization
  • Ingest and process high-volume IoT and historian data (sensor, SCADA, time-series) via streaming frameworks (Structured Streaming, Event Hubs/Kafka) and micro-batch patterns

Data Modeling, Quality & Governance

  • Model curated, analytics-ready datasets and serving layers that are well-documented, performant, and easy for downstream consumers to use
  • Implement automated data quality frameworks — validation, profiling, anomaly detection, freshness and completeness checks — with clear alerting and remediation paths
  • Build entity resolution and record linkage logic to unify wells, pads, assets, equipment, and events across heterogeneous source systems
  • Establish and enforce data governance using Unity Catalog — access controls, lineage, data classification, and a shared semantic/metadata layer that makes business concepts queryable and trustworthy

Reliability, CI/CD & Production Operations

  • Apply software engineering discipline to data: version control, code review, automated testing, and CI/CD pipelines (Azure DevOps or GitHub Actions) for data and infrastructure
  • Implement monitoring, logging, and observability across pipelines to support debugging, SLA tracking, cost monitoring, and continuous improvement
  • Support production incidents and platform-level issues impacting data pipelines and downstream consumers; develop runbooks and reduce toil through automation

Analytics Enablement & Collaboration

  • Partner with analysts and stakeholders to deliver datasets and semantic models that power dashboards in Power BI and Spotfire
  • Collaborate with Data Science and AI Engineering to provision clean, governed, feature-ready data for ML and agentic workflows
  • Translate domain problems from drilling, completions, production, geophysics, and power plant operations into well-scoped, reliable data products with clear ownership and success metrics 

Required Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Data Engineering, Software Engineering, Information Systems, Applied Mathematics, Physics, or a related technical field — or equivalent practical experience demonstrated through a portfolio of shipped data systems. Master’s preferred.
  • 2+ years of hands-on experience building and operating production data pipelines, not just prototypes or notebooks
  • Deep understanding of the Apache Spark framework and distributed, parallel data processing — partitioning, shuffles, joins, caching, and performance tuning at scale
  • Strong programming skills in Python (PySpark) and SQL, including writing testable, maintainable production code
  • Hands-on experience with Databricks, including Delta Lake, Delta Live Tables, and Unity Catalog
  • Experience with Azure Data Factory and Azure Data Lake Storage (ADLS), or equivalent cloud data services with willingness to work in our Azure-first environment
  • Experience with cloud data warehousing on Snowflake (or equivalent: BigQuery, Redshift, Databricks SQL)
  • Experience building both real-time/streaming and batch pipelines (Structured Streaming, Event Hubs/Kafka, or similar)
  • Solid data modeling skills (dimensional, medallion/lakehouse, or normalized) and a track record of building well-documented, consumable datasets
  • Experience implementing data quality, validation, and observability for pipelines
  • Strong Git and CI/CD experience (Azure DevOps or GitHub Actions), including version control discipline, code review, and automated testing
  • Experience delivering data to BI/analytics tools such as Power BI and/or Spotfire 

Preferred Qualifications

  • Experience with data governance and cataloging — Unity Catalog, Microsoft Purview, lineage tooling, and semantic/metric layers (Snowflake Semantic Views, Databricks Metric Views, dbt Semantic Layer)
  • Experience with entity resolution / record linkage and master data management across messy, multi-source data
  • Background in time-series data, signal processing, or industrial IoT (MQTT, OPC UA, SparkplugB) and historian systems (Canary, Ignition)
  • Experience with infrastructure-as-code (Terraform preferred) and containerization (Docker)
  • Experience with orchestration tooling (Databricks Workflows, Airflow, ADF pipelines) and dbt for transformation
  • Familiarity with data security and access control patterns — Entra ID, row/column-level security, masking, Key Vault-managed secrets
  • Oil and gas or energy industry experience, including familiarity with drilling, completions, production, geophysics, or industrial historian data
  • Exposure to ML/AI data needs — feature engineering, feature stores, or provisioning data for agentic/LLM applications
  • Experience optimizing cloud data cost and performance at scale

Location

Fervo Energy is headquartered in Houston, TX, with growing offices in Golden, CO, Reno, NV, and Oakland, CA, and Salt Lake City, UT. This position will be eligible for some hybrid work flexibility, but regular in-office presence at the Oakland or Houston office will be required.


Compensation & Benefits

Fervo provides a comprehensive suite of benefits including medical, dental, vision, life, short-term and long-term disability, flexible paid time off, and paid parental leave. Additionally, Fervo offers an incentive stock options program, a bonus incentive program, and a 401(k) plan with an employer match.

Fervo Energy is providing the compensation range and general description of other compensation and benefits that the company in good faith believes it might pay and/or offer for this position based on the successful applicant’s education, experience, knowledge, skills, and abilities in addition to internal equity and geographic location. Expected Salary: $103,000 – $170,000 based on location and experience.

Fervo Energy reserves the right to ultimately pay more or less than the posted range and offer other compensation, depending on circumstances not related to an applicant’s sex or other status protected by local, state, or federal law.


Fervo Energy is an Equal Opportunity Employer and does not discriminate on the basis of race, color, creed, gender, religion, marital status, registered domestic partner status, age, national origin, ancestry, physical or mental disability, medical condition, sex, genetic information, sexual orientation, military and veteran status or any other consideration made unlawful by federal, state, or local laws. It also prohibits unlawful discrimination based on the perception that anyone has any of those characteristics or is associated with a person who has or is perceived as having any of those characteristics.ffice will be required.

Skills Required

  • Bachelor's or Master's degree in Computer Science, Data Engineering, Software Engineering, Information Systems, Applied Mathematics, Physics, or related technical field, or equivalent practical experience
  • 2+ years of hands-on experience building and operating production data pipelines
  • Deep understanding of the Apache Spark framework and distributed, parallel data processing (partitioning, shuffles, joins, caching, performance tuning)
  • Strong programming skills in Python (PySpark) and SQL, writing testable, maintainable production code
  • Hands-on experience with Databricks, including Delta Lake, Delta Live Tables, and Unity Catalog
  • Experience with Azure Data Factory and Azure Data Lake Storage (ADLS) or equivalent cloud data services; willingness to work in an Azure-first environment
  • Experience with cloud data warehousing on Snowflake (or equivalent: BigQuery, Redshift, Databricks SQL)
  • Experience building both real-time/streaming and batch pipelines (Structured Streaming, Event Hubs/Kafka, or similar)
  • Solid data modeling skills (dimensional, medallion/lakehouse, or normalized) and track record of building well-documented, consumable datasets
  • Experience implementing data quality, validation, and observability for pipelines
  • Strong Git and CI/CD experience (Azure DevOps or GitHub Actions), including version control, code review, and automated testing
  • Experience delivering data to BI/analytics tools such as Power BI and/or Spotfire
  • Master's degree (preferred)
  • Experience with data governance and cataloging (Unity Catalog, Microsoft Purview, lineage tooling, semantic/metric layers)
  • Experience with entity resolution / record linkage and master data management
  • Background in time-series data, signal processing, or industrial IoT and historian systems (MQTT, OPC UA, SparkplugB, Canary, Ignition)
  • Experience with infrastructure-as-code (Terraform) and containerization (Docker)
  • Experience with orchestration tooling (Databricks Workflows, Airflow, ADF pipelines) and dbt for transformation
  • Familiarity with data security and access control patterns (Entra ID, row/column-level security, masking, Key Vault-managed secrets)
  • Oil and gas or energy industry experience, including drilling, completions, production, geophysics, or industrial historian data
  • Exposure to ML/AI data needs — feature engineering, feature stores, or provisioning data for agentic/LLM applications
  • Experience optimizing cloud data cost and performance at scale
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
175 Employees
Year Founded: 2017

What We Do

Fervo Energy is a next-generation geothermal power developer providing 24/7 carbon-free energy. The company leverages innovations in geoscience, including horizontal drilling and subsurface analytics, to make geothermal energy a dependable and affordable source of clean power. Its mission is to accelerate the global transition to sustainable energy and decarbonize the most challenging parts of the electricity sector through the development of enhanced geothermal systems.

Similar Jobs

Applied Systems Logo Applied Systems

Data Engineer

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Remote or Hybrid
United States
3079 Employees
70K-160K Annually

Wells Fargo Logo Wells Fargo

Infrastructure Engineer

Fintech • Financial Services
Hybrid
Lewisville, TX, USA
205000 Employees

Capital One Logo Capital One

Lead Data Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
3 Locations
55000 Employees
179K-225K Annually

PwC Logo PwC

Artificial Intelligence Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
15 Locations
370000 Employees
63K-142K Annually

Similar Companies Hiring

Halter Thumbnail
Software • Machine Learning • Internet of Things • Hardware • Greentech • Business Intelligence • Agriculture
Boulder, Colorado
350 Employees
Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
108 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account