Senior Software Engineer — Lakehouse Systems

Posted One Month Ago
Be an Early Applicant
Mountain View, CA, USA
In-Office
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Better Data for Better AI
The Role
Build foundational lakehouse infrastructure for exabyte-scale AI data environments. Responsibilities include metadata and transaction systems, table maintenance, schema and partition evolution, snapshot isolation, compaction, clustering, file-layout optimization, object-store performance, columnar-format optimization, and query performance across major lakehouse engines. The role also involves debugging distributed systems, implementing compression and data-efficiency algorithms, and contributing to open-source or research efforts.
Summary Generated by Built In
Senior Software Engineer — Lakehouse Systems

Location: Mountain View, CA — On-site

About Granica

Granica builds AI infrastructure for enterprises operating massive data environments.
Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.

Granica’s products include:

  • Crunch — continuous optimization for enterprise lakehouse data

  • Myelin — stateful infrastructure for long-running AI agents

  • Large Tabular Models — foundation models designed for enterprise tables

Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.

About the Role

Granica is hiring a Senior Software Engineer to build foundational lakehouse systems for AI.

You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, transaction semantics, table maintenance, object-store-backed storage layouts, file optimization, and lakehouse cost/performance across petabyte- and exabyte-scale environments.

You will own core systems that directly affect customer infrastructure cost, query performance, table reliability, and the operational health of large lakehouse environments.

This is a hands-on systems role for engineers who have gone deep on lakehouse internals, table formats, metadata systems, storage layout, or distributed storage infrastructure.

What You’ll Do
  • Build metadata, transaction, and table-maintenance systems for large-scale lakehouse datasets

  • Work with Iceberg, Delta Lake, Hudi, manifests, snapshots, transaction logs, schema evolution, and garbage collection

  • Optimize compaction, clustering, file sizing, data skipping, pruning, and physical data layout

  • Improve performance and cost efficiency across Parquet/ORC and object stores such as S3, GCS, and ADLS

  • Debug and optimize bottlenecks across metadata, storage, table maintenance, object-store access, and query execution

What We’re Looking For
  • Deep engineering experience in distributed systems, storage systems, databases, or data infrastructure

  • Production experience building, extending, or deeply optimizing lakehouse or table-format systems such as Iceberg, Delta Lake, Hudi, or similar technologies

  • Strong understanding of metadata architectures, transaction semantics, snapshots, manifests, schema evolution, and physical data layout

  • Hands-on experience with table maintenance, compaction, clustering, file sizing, Parquet/ORC, and cloud object stores such as S3, GCS, or ADLS

  • Strong programming skills in Java, Scala, Go, Rust, C++, or a similar systems-oriented language, with a pragmatic end-to-end builder mindset

Bonus
  • Contributions to Iceberg, Delta Lake, Hudi, Parquet, ORC, Spark, Trino, Flink, Velox, DuckDB, DataFusion, or related systems

  • Experience with small-file optimization, metadata scaling, delete handling, catalog consistency, indexing, caching, compression, or storage-engine internals

  • Research or open-source contributions in distributed systems, databases, storage, compression, or data processing

Why Join Granica
  • Build foundational infrastructure for enterprise data and AI

  • Work on deep systems problems across lakehouse metadata, table formats, storage layout, and object-store behavior

  • Own meaningful parts of the architecture in a small, high-caliber engineering team

  • Work directly with Product, Engineering, and company leadership

  • Have direct impact on customer performance, infrastructure cost, product direction, and company growth

Compensation & Benefits
  • Competitive salary, meaningful equity, and performance bonus for top performers

  • 401(k) with company match, comprehensive health coverage, and unlimited PTO

  • Daily catered meals in our Mountain View office

  • Support for research, publication, and conference participation

At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

 

Skills Required

  • Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
  • Production experience with modern data lake or lakehouse technologies such as Apache Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, or Unity Catalog
  • Hands-on experience with columnar formats such as Parquet or ORC
  • Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
  • Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
  • Familiarity with cloud object storage systems such as S3, GCS, or ADLS and their performance tradeoffs
  • Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
  • Curiosity about compression, entropy, information theory, and data representation effects on AI efficiency
  • Pragmatic builder mindset; rigorous, hands-on, and comfortable owning complex systems end to end
  • Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems
  • Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection
  • Experience solving the small-file problem or optimizing object-store access patterns at scale
  • Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization
  • Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing
  • Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency

Granica Compensation & Benefits Highlights

  • Healthcare Strength — Public job materials describe premium medical, dental, and vision coverage, with some postings indicating fully covered employee premiums and dependent support. This positions healthcare as a robust anchor of the total rewards package.
  • Leave & Time Off Breadth — Listings consistently advertise unlimited PTO alongside paid holidays/sick time and quarterly company‑wide recharge days. Some sources also note guidance encouraging roughly four weeks of time off under the unlimited policy.
  • Equity Value & Accessibility — Employer pages and postings highlight meaningful equity as a core component of compensation. Equity is presented alongside competitive base pay as part of a comprehensive total‑rewards design.

Granica Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, California
45 Employees
Year Founded: 2023

What We Do

Our mission is to remove inefficiency from the foundation of AI. By combining new research in information theory, probabilistic modeling, and distributed systems, we’re creating self-optimizing data infrastructure that continuously improves how information is represented and used by intelligent systems.

Why Work With Us

We’re a tight-knit team combining --> * Fundamental research in compression, data systems, and information theory * World-class systems engineering across storage, infrastructure, and research led by our Chief Scientist & Stanford Prof. Andrea Montanari * A shared obsession with performance, scale, and clean design

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Granica Offices

OnSite Workspace

Typical time on-site:
HQMountain View, California
India
Learn more

Similar Jobs

Granica Logo Granica

Senior Software Engineer

Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Hybrid
Mountain View, CA, USA
45 Employees
160K-240K Annually

Granica Logo Granica

Head of Finance — Strategic Finance & Corporate Development

Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
In-Office
Mountain View, CA, USA
45 Employees
140K-180K Annually

Granica Logo Granica

Scientist

Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
In-Office
Mountain View, CA, USA
45 Employees
160K-240K Annually

Granica Logo Granica

Scientist

Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
In-Office
Mountain View, CA, USA
45 Employees
160K-240K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account