AI Developer

Posted Yesterday
Hiring Remotely in United States
Remote
Mid level
Blockchain • Software • Automation
The Role
Design, train, optimize, and deploy LLMs for offline and on-prem environments. Build end-to-end LLM pipelines: data preprocessing, SFT, LoRA/Q-LoRA, quantization, RAG, vector stores, MCP integrations, and inference systems. Implement document parsing, GPU optimization, CI/CD for local ML workflows, and maintain model registries and versioning in restricted networks.
Summary Generated by Built In

About Salvo Software

Salvo Software is a global firm that provides cost-effective software solutions to guide enterprises and startups through digital transformation. With distributed teams across the US, LATAM, and India, we partner with clients to build high-performance, scalable systems that solve complex technical challenges. Our culture values innovation, ownership, and engineering excellence.

Role Overview

We are seeking a highly skilled AI Developer with a strong backend and machine learning engineering background to design, train, optimize, and deploy LLM models in on-prem and offline environments. This role is deeply technical and hands-on.

You will work closely with our engineering and product teams to build end-to-end LLM pipelines — including data preprocessing, supervised fine-tuning, model quantization, evaluation, RAG pipeline design, and deployment using local or air-gapped infrastructure. If you enjoy working with cutting-edge open-source LLMs, building context-aware AI systems, and designing reliable backend pipelines, this role is for you.

Key Responsibilities

Core LLM Development

  • Train and fine-tune LLMs using supervised fine-tuning (SFT).
  • Work with open-source models such as LLaMA, Mistral, Qwen, and similar architectures.
  • Build LoRA / Q-LoRA pipelines for efficient fine-tuning.
  • Implement and optimize data preprocessing workflows, including tokenization and long-context handling.
  • Use and extend Hugging Face Transformers & Datasets for training and inference.
  • Parse and process structured and semi-structured data, including XML/XSD files.
  • Implement document parsing solutions for Office formats (python-docx, OpenXML).

RAG & Context-Aware Systems

  • Design and implement end-to-end Retrieval-Augmented Generation (RAG) pipelines for document-grounded question answering and knowledge retrieval.
  • Build and maintain vector stores and embedding pipelines using tools such as FAISS, Chroma, Weaviate, or pgvector.
  • Optimize retrieval strategies including hybrid search, re-ranking, and chunking approaches tailored for domain-specific corpora.
  • Develop and maintain MCP (Model Context Protocol) server integrations to enable LLMs to interact dynamically with tools, APIs, and external data sources.
  • Design agentic workflows that leverage MCP to give models structured access to internal systems and context in a controlled, auditable manner.

Offline / On-Prem Model Expertise

  • Deploy, run, and maintain models fully offline and in air-gapped environments.
  • Perform model optimization and quantization (GGUF, GPTQ, AWQ, bitsandbytes).
  • Build and maintain inference systems using frameworks like vLLM, TGI, and Ollama.
  • Optimize GPU usage (CUDA, cuDNN, VRAM-aware batching).
  • Maintain local CI/CD pipelines for ML models without cloud dependencies.
  • Manage local model registries, versioning, and artifacts.
  • Ensure RAG and MCP components are fully operational in offline and restricted network environments.

Backend & DevOps

  • Build backend services in Python for ML training and inference workflows.
  • Work with relational databases (Postgres/MySQL) and vector databases for RAG storage layers.
  • Use Docker and Git for reliable development and deployment pipelines.
  • Use Azure DevOps for CI/CD, including local runners when applicable.

Requirements

Technical Skills

  • Strong experience in Python for backend and machine learning development.
  • Expertise with ML frameworks such as PyTorch or TensorFlow, along with scikit-learn and pandas.
  • Solid knowledge of Postgres or MySQL for data storage.
  • Experience with Docker and Git.
  • Hands-on experience with LLM training, fine-tuning, and optimization.
  • Experience with Hugging Face Transformers & Datasets.
  • Familiarity with XML/XSD and Office document parsing tools.
  • Experience deploying models with vLLM, TGI, or Ollama.
  • Understanding of quantization techniques such as GGUF, GPTQ, or AWQ.
  • Experience with GPU optimization and the CUDA stack.
  • Experience building solutions for offline, on-prem, and air-gapped environments.
  • Hands-on experience designing and implementing RAG pipelines, including embedding models, vector stores, and retrieval optimization strategies.
  • Experience building or integrating MCP (Model Context Protocol) servers to connect LLMs with external tools, APIs, and structured data sources.
  • Experience with advanced RAG techniques such as HyDE or multi-hop retrieval.

Nice to Have

  • Experience building agentic systems using MCP in production or near-production environments.
  • Experience managing ML model registries in offline environments.
  • Familiarity with AWS for hybrid deployments.
  • Experience with secure environments, restricted networks, or enterprise compliance requirements.

Soft Skills

  • Experience discussing complex technical topics with both technical and non-technical stakeholders.

Skills Required

  • Strong experience in Python for backend and ML development
  • Experience with PyTorch or TensorFlow
  • Experience with scikit-learn and pandas
  • Hands-on LLM training, supervised fine-tuning (SFT), and LoRA/Q-LoRA pipelines
  • Experience with Hugging Face Transformers & Datasets
  • Experience with model quantization techniques (GGUF, GPTQ, AWQ) and bitsandbytes
  • Deployment and inference frameworks: vLLM, TGI, or Ollama
  • GPU optimization and CUDA stack knowledge (CUDA, cuDNN, VRAM-aware batching)
  • Design and implement RAG pipelines, embeddings, and vector stores
  • Experience with FAISS, Chroma, Weaviate, or pgvector
  • Backend service development in Python for training and inference workflows
  • Experience with relational databases: Postgres or MySQL
  • Experience building solutions for offline, on-prem, and air-gapped environments
  • Experience parsing structured and semi-structured data and Office formats (XML/XSD, python-docx, OpenXML)
  • Use of Docker and Git for development and deployment
  • Experience with Azure DevOps for CI/CD (including local runners)
  • Experience integrating or building MCP (Model Context Protocol) servers
  • Familiarity with advanced RAG techniques (HyDE, multi-hop retrieval) and retrieval optimization
  • Experience with model registries, versioning, and local CI/CD for ML artifacts
  • Familiarity with AWS for hybrid deployments
  • Experience with secure/restricted network environments and enterprise compliance
  • Ability to communicate complex technical topics to technical and non-technical stakeholders
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: VANCOUVER, WA
16 Employees
Year Founded: 2017

What We Do

We design custom-built solutions to help you transform, scale, and grow your business along with a team that cares about you. Salvo software is a global firm with near-shoring capabilities headquartered in Vancouver, WA. That provides cost-effective software solutions to guide enterprises and startups through digital transformation. We help our partners to improve their client’s customer experience and optimize their business process times by providing hand-selected teams of experts that meet their needs and help them to make smart decisions.

Similar Jobs

Optum Logo Optum

Artificial Intelligence Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office or Remote
Eden Prairie, MN, USA
160000 Employees
146K-250K Annually

Zscaler Logo Zscaler

Artificial Intelligence Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
USA
8697 Employees
186K-265K Annually

Capital One Logo Capital One

Artificial Intelligence Engineer

Fintech • Machine Learning • Payments • Software • Financial Services
Remote or Hybrid
4 Locations
55000 Employees
245K-335K Annually
Remote or Hybrid
3 Locations
289097 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account