Multimodal AI Engineer, Document Understanding

Reposted 18 Days Ago
Hiring Remotely in San Francisco, CA, USA
Remote or Hybrid
180K-250K Annually
Mid level
Artificial Intelligence • Information Technology • Machine Learning • Software
The Role
Develop and optimize machine learning models for document understanding, handle production ML systems, and integrate innovations into APIs.
Summary Generated by Built In

Join us and help shape the future of AI by defining the narrative around document understanding.

About the Role:

We are seeking exceptional AI engineers to join our core document understanding team. You will work at the intersection of computer vision, natural language processing, and production ML systems to push the boundaries of what's possible in document parsing and understanding.

Our document understanding team builds the intelligence behind LlamaParse, LlamaExtract, and our other processing products. These systems are processing millions of complex documents including PDFs, PowerPoints, Word documents, and spreadsheets. Your work will directly impact thousands of developers building RAG applications and document agents, while also contributing to our open-source frameworks that shape how the industry approaches document processing.

Depending on your background and interests, you might focus more on data curation and evaluation, model fine-tuning and experimentation, or ML infrastructure and production systems. We're hiring multiple people and will work with you to find the best fit.

Responsibilities:
  • Develop, train, and optimize machine learning models for document structure understanding, table extraction, layout analysis, and multimodal content processing

  • Build robust data pipelines, evaluation frameworks, and experimentation infrastructure

  • Design and implement production ML systems that handle complex, real-world documents at scale

  • Stay current with latest advances in vision-language models, document AI, and multimodal learning

  • Collaborate with engineering teams to integrate ML innovations into production APIs

  • Contribute to both our open-source frameworks and enterprise offerings

  • Drive technical decisions while balancing research exploration with product delivery

Required Qualifications:
  • 3-7 years of experience in machine learning engineering or applied research

  • Strong software engineering fundamentals with production Python experience (modern tooling: uv, ruff, mypy, Pydantic)

  • Hands-on experience training, fine-tuning, or deploying ML models in production

  • Deep understanding of modern ML techniques, particularly in computer vision, NLP, or multimodal learning

  • Experience with at least one of: data pipeline development, model training/fine-tuning, or ML infrastructure

  • Ability to read and implement from research papers and technical specifications

  • Track record of executing with high intensity in fast-paced environments

  • Strong technical communication skills and comfort with open-source collaboration

Preferred Qualifications:
  • Experience with vision-language models, transformer architectures, or model fine-tuning (LoRA, QLoRA)

  • Experience building evaluation frameworks, benchmarks, or data quality pipelines

  • Experience with model serving frameworks (vLLM, TensorRT, ONNX) or MLOps tools

  • Experience specifically with document understanding, OCR, or layout analysis

  • Contributions to open-source ML projects or frameworks

  • Experience with LLM applications and RAG systems

  • Strong understanding of model optimization techniques (quantization, distillation, pruning)

  • Experience with Docker/Kubernetes and distributed systems

  • Active participation in ML research community

Location:

We offer a hybrid-friendly culture based out of our downtown San Francisco office. Remote candidates will be considered for exceptional fits.

Why Join Us?
  • Impactful Mission: Work on innovative AI products that redefine how knowledge is accessed and utilized. Your models will process millions of documents and directly impact thousands of developers.

  • Cutting-Edge Technology: Work with the latest vision-language models, contribute to open-source frameworks used industry-wide, and shape the future of document AI.

  • Collaborative Team: Join a focused team of passionate engineers and researchers committed to pushing the boundaries of what's possible in document understanding.

  • Technical Autonomy: Significant creative freedom to explore new approaches while maintaining focus on delivering high-quality, production-ready solutions.

  • Growth Opportunities: Be at the forefront of the AI revolution, with ample opportunities to grow alongside our scaling organization. Shape your role based on your interests and strengths.

Additional Benefits:
  • Competitive base salary and equity compensation

  • Comprehensive medical/dental/vision coverage for you and your family

  • Unlimited paid time off policy

  • Daily catered lunch and snacks in the San Francisco office

  • Budget for conferences, research materials, and professional development

  • Access to cutting-edge compute resources and research tools

Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.

LlamaIndex does not accept unsolicited agency resumes. Please do not forward resumes to our jobs alias, employees, or any other organization location. LlamaIndex is not responsible for any fees related to unsolicited resumes.

Top Skills

Docker
Kubernetes
Mypy
Onnx
Pydantic
Python
Ruff
Tensorrt
Uv
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, California
44 Employees

What We Do

The data framework for LLMs Python: Github: https://github.com/jerryjliu/llama_index Docs: https://docs.llamaindex.ai/ Typescript/Javascript: Github: https://github.com/run-llama/LlamaIndexTS Docs: https://ts.llamaindex.ai/ Other: Discord: discord.gg/dGcwcsnxhU LlamaHub: llamahub.ai Twitter: https://twitter.com/llama_index Blog: blog.llamaindex.ai #ai #llms #rag

Similar Jobs

Snap Inc. Logo Snap Inc.

Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Remote or Hybrid
6 Locations
5000 Employees
133K-235K Annually

2K Logo 2K

Senior Server Engineer - NBA 2K (REMOTE)

Gaming • Information Technology • Mobile • Software • Esports
Remote or Hybrid
Novato, CA, USA
3505 Employees
118K-174K Annually

Zapier Logo Zapier

Manager, Mid Market Sales

Artificial Intelligence • Productivity • Software • Automation
Remote
2 Locations
800 Employees
258K-335K Annually

Wipfli Logo Wipfli

Senior Manager, Accounting Advisory - Tribal Government Industry

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
United States
3000 Employees
142K-195K Annually

Similar Companies Hiring

Fairly Even Thumbnail
Software • Sales • Robotics • Other • Hospitality • Hardware
New York, NY
Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account