Senior AI Data Engineer
Job Summary
We are seeking a Senior AI Data Engineer to design and operate the data foundation that our AI systems depend on. This role owns the movement, modeling, and quality of data from source systems through the warehouse and into the retrieval and feature layers that power LLM pipelines, agentic workflows, and analytical products.
The ideal candidate is a rigorous software engineer first and a data specialist second: someone who models a warehouse deliberately, writes production Python that other engineers can extend, and treats pipelines as versioned, tested, observable software rather than scripts.
This role partners closely with the AI/ML Data Scientist, who owns model behavior and retrieval strategy.
The boundary: you own the pipeline, the schema, and the guarantees; they own the algorithm, the prompt, and the evaluation.
Required Qualifications
5-10+ years in software engineering or data engineering, with substantial time in production data platform work.
Data warehousing: demonstrable command of Kimball dimensional modeling - not just familiarity with the vocabulary, but the judgment to choose a grain, resolve a many-to-many relationship, and know when to denormalize. Working knowledge of alternative approaches (Data Vault, One Big Table, Inman) and the tradeoffs against Kimball.
SQL: expert-level - window functions, CTEs, query plan reading, and performance tuning on a columnar warehouse.
Python: expert-level, production-grade - typing, packaging, dependency management, testing.
Design: SOLID and domain-driven design applied in real systems, with examples you can walk through.
Orchestration: Airflow, Prefect, Dagster, or equivalent, in production.
Cloud: expert-level on AWS, Azure, or GCP - storage, compute, IAM, networking, and cost management.
Platform: containerization, Kubernetes (EKS/AKS/GKE), and CI/CD.
Experience with lakehouse table formats (Iceberg, Delta Lake, Hudi) and their maintenance characteristics: compaction, snapshot expiry, schema and partition evolution.
Preferred Qualifications
Experience building the data layer beneath production RAG systems, including hybrid search infrastructure and index freshness guarantees.
Streaming systems: Kafka, Kinesis, Flink, or Spark Structured Streaming.
dbt or an equivalent transformation and testing framework.
Data quality tooling (Great Expectations, Soda, or similar) and catalog/lineage platforms.
Familiarity with the model-facing side of the stack - MLflow, Weights & Biases, feature stores - sufficient to collaborate credibly with data scientists.
Working knowledge of a second language: TypeScript, Java, Go, Scala, or Rust.
Experience with AI security, governance, and compliance frameworks.
Open-source contributions to data or AI infrastructure projects.
Skills Required
- 5-10+ years of software engineering or data engineering experience, including substantial production data platform work
- Expert command of Kimball dimensional modeling, including grain selection, many-to-many relationships, denormalization, and alternative modeling approaches
- Expert-level SQL, including window functions, CTEs, query plan reading, and performance tuning on columnar warehouses
- Expert-level production Python, including typing, packaging, dependency management, and testing
- Applied experience with SOLID principles and domain-driven design
- Production experience with Airflow, Prefect, Dagster, or equivalent orchestration frameworks
- Expert-level experience with AWS, Azure, or GCP, including storage, compute, IAM, networking, and cost management
- Experience with containerization, Kubernetes, and CI/CD
- Experience with Iceberg, Delta Lake, or Hudi lakehouse table formats
- Experience building data layers for production RAG systems, hybrid search, and index freshness guarantees
- Experience with Kafka, Kinesis, Flink, or Spark Structured Streaming
- Experience with dbt or an equivalent transformation and testing framework
- Experience with data quality tooling and catalog or lineage platforms
- Familiarity with MLflow, Weights & Biases, or feature stores
- Working knowledge of TypeScript, Java, Go, Scala, or Rust
- Experience with AI security, governance, and compliance frameworks
- Open-source contributions to data or AI infrastructure projects
What We Do
Strategic Systems International (SSI) is a fast-growing Advanced Analytics and Software Engineering firm that partners with tech companies to help them launch and scale their products. The company was launched in 1991 by alumni of University of Chicago and Northwestern has grown to 200 employees with presence in US, Europe and Asia. We architect a








