LLM Platform Engineer

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
In-Office
Mid level
Artificial Intelligence • Edtech • Software • Defense
The Role
Build and operate multimodal LLM platform capabilities, including knowledge ingestion, embeddings, retrieval-augmented generation, evaluation, feedback loops, deployment, monitoring, and latency optimization. The role spans cloud, on-premises, and air-gapped environments, integrating AI knowledge systems with 3D scene representations and real-time training applications. Responsibilities include productionizing AI updates, improving reliability, and collaborating with spatial computing, mobile, and product engineering teams.
Summary Generated by Built In

About Schemata

At Schemata, we are transforming the $400 B virtual‑training and simulation market by fusing 3D computer vision, neural rendering and large multimodal models inside highly regulated industries. Our platform delivers photorealistic, intelligent 3D experiences, demanding robust spatial reasoning, high‑performance data pipelines and seamless integration between traditional graphics and AI‑driven perception.

About the Role

We are seeking a highly skilled LLM Platform Engineer to join our team full‑time. You will play a foundational role in designing, building and optimizing the AI systems that turn heterogeneous knowledge — technical documentation, video, imagery, voice, and 3D scene data — into grounded, real‑time guidance for next‑generation training and maintenance applications.

This is a high‑impact, cross‑functional role: you will work end‑to‑end from evaluating emerging methods to production inference and performance optimization, ensuring our platform retrieves the right knowledge, reasons over it reliably, and responds in real time across diverse deployment environments.

Core Responsibilities
  • Design and build multimodal knowledge base ingestion: pipelines that convert documents, video, images, audio, and structured data into a unified, queryable semantic layer with rich embeddings and metadata, including bringing in external sources and messy real‑world formats as they arise.

  • Improve our retrieval‑augmented generation (RAG) systems: chunking and indexing strategies, hybrid retrieval, reranking, grounding controls, and context assembly for multimodal queries.

  • Build the knowledge base lifecycle: handle versioned documentation updates, incorporate SME corrections and annotations, and design validation and freshness mechanisms so the knowledge base improves continuously.

  • Scale the feedback loop: capture user signals (ratings, corrections, query patterns), route them into measurable improvements to retrieval and content, and build a durable picture of how the AI is performing in the field.

  • Own the deployment and delivery lifecycle of AI system updates across cloud, on‑premises, and air‑gapped environments: packaging, release management, configuration, and monitoring.

  • Strengthen the reliability and latency of AI features: profiling inference paths, caching, fallback behavior, and graceful degradation in production.

  • Collaborate with spatial computing, mobile, and product engineers to connect the knowledge base with 3D scene representations and ship mission‑critical features.

Essential Skills & Experience
  • 4+ years of software engineering experience, with at least 1–2 years building production LLM systems (RAG pipelines, agents, or LLM‑powered products with real users).

  • Hands‑on experience designing knowledge ingestion and retrieval systems: embedding models, vector and hybrid search, chunking strategies, and semantic data modeling across more than one modality.

  • Strong Python and backend engineering fundamentals: APIs, data pipelines, async systems, and working in a cloud environment (AWS preferred).

  • A metrics‑driven approach: experience building evals or benchmarks for LLM systems and using them to drive iteration, not just report scores.

  • Demonstrated ability to take ambiguous problems from prototype to reliable, maintainable production services.

  • Experience deploying and operating production systems: CI/CD, release management, and monitoring.

Nice to Have
  • Experience fine‑tuning open‑weight models (LoRA/PEFT, distillation) or deploying quantized models for offline, edge, or on‑premises environments.

  • Research background or publications in retrieval, multimodal learning, or agentic systems, or a track record of translating recent research into shipped capabilities.

  • Voice pipeline experience: streaming STT/TTS, latency optimization for real‑time conversational systems.

  • Familiarity with 3D or spatial data (scene graphs, point clouds, 3DGS) and grounding language models in spatial representations.

  • MLOps and inference infrastructure experience: GPU serving, model versioning, cost/latency optimization at scale.

  • Defense, aerospace, energy or other regulated‑industry experience; active or ability to obtain U.S. security clearance.

Why Join Us?

  • Competitive salary that reflects your experience and track record

  • Meaningful equity stake in a high-growth, venture-backed defense tech startup, so you share in the upside you help create

  • Comprehensive health coverage: medical, dental, and vision insurance

  • 401(k) plan

  • Paid parental leave

  • High visibility and real impact: Collaborate with world-class engineers and researchers in a high-ownership environment.

Skills Required

  • 4+ years of software engineering experience
  • 1–2 years building production LLM systems, such as RAG pipelines, agents, or LLM-powered products
  • Experience designing knowledge ingestion and retrieval systems across multiple modalities
  • Experience with embedding models, vector search, hybrid search, chunking strategies, and semantic data modeling
  • Strong Python and backend engineering fundamentals
  • Experience building APIs, data pipelines, and asynchronous systems
  • Experience working in a cloud environment; AWS preferred
  • Experience building evaluations or benchmarks for LLM systems and using metrics to drive improvements
  • Ability to take ambiguous problems from prototype to reliable, maintainable production services
  • Experience deploying and operating production systems, including CI/CD, release management, and monitoring
  • Experience fine-tuning open-weight models using LoRA or PEFT, or deploying quantized models
  • Research background or publications in retrieval, multimodal learning, or agentic systems
  • Experience translating recent research into shipped capabilities
  • Voice pipeline experience with streaming STT/TTS and real-time latency optimization
  • Familiarity with 3D or spatial data, including scene graphs, point clouds, or 3DGS
  • MLOps and inference infrastructure experience, including GPU serving, model versioning, and cost or latency optimization
  • Defense, aerospace, energy, or other regulated-industry experience
  • Active U.S. security clearance or ability to obtain one
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
15 Employees
Year Founded: 2022

What We Do

Schemata is a spatial intelligence company developing AI-native, immersive training and simulation for defense and enterprise organizations. Its platform uses 3D reality capture and AI-driven reconstruction to turn images, physical systems, and technical documentation into photorealistic, interactive environments. Users can deploy the system securely at the point of need for equipment guidance, procedural practice, operational readiness, and scalable training programs.

Similar Jobs

Pax Historia Logo Pax Historia

Founding Engineer - LLM Infra & Platform

Artificial Intelligence • Gaming • Software
In-Office
San Francisco, CA, USA
6 Employees
140K-220K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
5 Locations
55000 Employees
179K-246K Annually

Whatnot Logo Whatnot

Platform Engineer

eCommerce • Mobile • Retail
In-Office
San Francisco, CA, USA
1200 Employees
200K-345K Annually

Similar Companies Hiring

Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account