Data Pipeline Engineer

Posted 2 Days Ago
Sunnyvale, CA, USA
In-Office
150K-250K Annually
Senior level
Artificial Intelligence • Hardware • Software • Cybersecurity
The Role
Design, build, and maintain a petabyte-scale open-source data lakehouse and end-to-end data pipelines (ingest → transform → consume) ensuring performance, reliability, and data quality for analytics workloads.
Summary Generated by Built In
Your Impact

Join a small team building the next generation of cybersecurity products from the ground up. Led by industry veterans with a proven track record of success - you will get to architect, build, and deliver hugely impactful products with this world-class team. You will have the opportunity to grow your career and skills along with the company from the very start.

Role Overview

Design, build, and maintain a scalable, open-source data lakehouse architecture supporting petabyte-scale analytics workloads. Responsible for architecting end-to-end data pipelines from ingestion through transformation to consumption, ensuring high performance, reliability, and data quality.

Required Experience
  • A proven track record of success architecting, building, and running large-scale data systems (PB scale)

  • Experience with both batch and real-time processing architectures

  • Experience with open source Data Lakehouse components, including Apache Iceberg, PostgreSQL, Neo4j, Apache Parquet, etc.

  • Experience with tools for stream processing and data analytics - Apache Kafka, Spark, Flink, etc.

  • Understanding and experience with data transformation solutions

  • Excellent programming experience with Python

  • Strong communication and documentation skills

  • Experience with CSP data platforms is a plus

  • Understanding of data lineage, quality, and governance tooling is a plus

We offer competitive compensation and a comprehensive benefits package designed to support our employees’ health, well-being, and long-term success. The expected salary range for this position is $150,000 – $250,000 per year. Within this range, individual pay is determined based on job-related factors including skills, experience, qualifications, and internal equity. Most candidates can expect an offer within the range listed above. Your recruiter will provide additional details on compensation and benefits to qualified candidates during the hiring process.

We’re committed to building a diverse, inclusive workplace where everyone can do their best work. We are proud to be an equal opportunity employer and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, age, veteran status, disability, genetic information, or other applicable legally protected characteristic. If you require a reasonable accommodation during the application or interview process, please let us know — we’re happy to support you.

Skills Required

  • Proven track record architecting, building, and running large-scale data systems (PB scale)
  • Experience with both batch and real-time processing architectures
  • Experience with open source Data Lakehouse components, including Apache Iceberg, PostgreSQL, Neo4j, Apache Parquet
  • Experience with tools for stream processing and data analytics - Apache Kafka, Spark, Flink
  • Understanding and experience with data transformation solutions
  • Excellent programming experience with Python
  • Strong communication and documentation skills
  • Experience with CSP data platforms
  • Understanding of data lineage, quality, and governance tooling
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
32 Employees
Year Founded: 2026

What We Do

Cylake is developing a complete, AI-native, data-driven cybersecurity platform for regulated organizations that require operational and data sovereignty. Designed to run on-premises or in private clouds, its system combines hardware and software to provide end-to-end protection. The platform analyzes an organization’s data and operations, delivering deep visibility and agentic security without relying on public cloud or public AI infrastructure for sensitive environments.

Similar Jobs

In-Office
Santa Clara, CA, USA
993 Employees
175K-296K Annually
In-Office
Santa Clara, CA, USA
993 Employees
203K-344K Annually

42dot Logo 42dot

Senior AI Data Pipeline Engineer (Autonomous Driving)

Artificial Intelligence • Software • Transportation
Hybrid
Sunnyvale, CA, USA
739 Employees
133K-254K Annually
In-Office
Sunnyvale, CA, USA
472 Employees
150K-200K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account