The Role
Hands-on Principal Engineer to design, code, and ship production-grade Generative AI systems: build RAG pipelines, embeddings and vector DB integrations, LLMOps, multi-agent orchestration, and integrate AI workflows with backend microservices and cloud infrastructure. Mentor via code reviews and own end-to-end delivery.
Summary Generated by Built In
We're looking for a hands-on AI Engineer who combines strong backend engineering fundamentals with hands-on experience building production Generative AI systems. You'll design and ship RAG pipelines, integrate LLMs into real products, and build the backend services that support them, writing code daily, not just architecting on paper.
This is a purely technical IC role, not a managerial one. You’ll lead by example, mentor through code reviews, and own end-to-end technical delivery.
Key Responsibilities
- Design and build RAG systems, embeddings, vector search, chunking, and evaluation pipelines.
- Build and maintain multi-agent orchestration workflows (LangGraph, AutoGen, CrewAI, or similar).
- Develop backend services and APIs (Python — Flask/FastAPI) that expose AI workflows to production systems.
- Deploy and scale AI workloads in cloud-native environments, using serverless or containerized patterns.
- Implement LLMOps practices: prompt versioning, cost tracking, monitoring, and evaluation.
- Write clean, tested code, and use AI-assisted tools (Copilot, Cursor, Claude Code) to move faster without cutting corners.
- Work with data and platform engineers to ship GenAI features quickly, from prototype to production.
Skills, Knowledge and Expertise
Must-Have Skills
- 5+ years of backend experience, with strong Python coding skills.
- Proven experience shipping RAG systems (vector DBs, embeddings, chunking).
- Familiarity with orchestration frameworks (LangGraph, LangChain, AutoGen, or similar).
- Experience with APIs, microservices, and cloud-native development (AWS preferred).
- Familiarity with distributed systems concepts (async, message queues, caching).
- Experience with unstructured data (PDFs, tables, images).
- Builder mindset: thrives on writing, debugging, and improving production code.
- Collaborative, humble, and open to feedback.
- Strong communicator who explains design decisions clearly.
Influences through contribution, not hierarchy.
About
Emumba is a global engineering and consulting company with strengths in software development and an established AWS cloud practice focused on Data and GenAI. For 15 years, our teams across the US, the UAE, and Pakistan have earned trust through quality delivery and ownership of work. We look for people who value the culture they work in as much as the craft they bring to it.
Skills Required
- 5+ years of backend or ML engineering experience
- Strong Python coding skills
- Proven experience shipping RAG systems (vector DBs, embeddings, chunking)
- Familiarity with orchestration frameworks (LangGraph, LangChain, AutoGen, or similar)
- Understanding of LLM behavior, evaluation, and fine-tuning workflows
- Experience with APIs, microservices, and cloud-native development
- Experience with AWS
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Emumba is a global software services company specializing in enterprise-grade software, AI systems, cloud, and DevOps solutions, enabling innovation for Fortune 500 customers.








