Senior AI Engineer- Remote, India

Posted 6 Days Ago
Be an Early Applicant
Hiring Remotely in Gurugram, Haryana , IND
In-Office or Remote
Senior level
Artificial Intelligence • Blockchain • Information Technology • Internet of Things
The Role
Build and operate production Generative AI systems, including LLM applications, RAG pipelines, agentic workflows, evaluation harnesses, model serving, and inference optimization. Own AI components end to end, from retrieval and model integration through deployment, monitoring, and cost governance. Lead technical direction, conduct reviews, mentor engineers, prototype emerging techniques, and collaborate with product, architecture, and client stakeholders on enterprise AI delivery.
Summary Generated by Built In

This is a remote position.

We are looking for a hands-on Senior AI Engineer with 4-6 years of engineering experience and deep, current expertise in building Generative AI systems that run in production. This is a core AI engineering role, not a data analysis or reporting role: you will spend your time on LLM application architecture, retrieval design, agentic systems, fine-tuning, evaluation, inference optimization and model serving. We expect a strong grasp of how transformer models actually behave, including tokenization, context windows, attention costs, embeddings and decoding, and the ability to turn that understanding into systems that are accurate, fast and affordable at scale. You will own AI components end to end, from data and retrieval pipelines through model integration, evaluation harnesses, deployment and production monitoring, and will work directly with product managers, solution architects and client stakeholders on enterprise AI delivery. As a senior engineer you will also set the technical direction on the AI components you own, review the work of others, and raise the bar for what ships. The role suits an engineer who reads papers and model cards, prototypes quickly, and holds a high standard for production quality.



Requirements Responsibilities
  • Design and build production LLM applications: RAG pipelines, agentic and tool-calling systems, and multi-step reasoning workflows.
  • Engineer retrieval layers on vector databases, covering chunking strategy, embedding model selection, metadata filtering, hybrid search and re-ranking.
  • Design context and prompting strategies: system prompt architecture, structured output, function/tool schemas, context compression and memory management.
  • Build evaluation harnesses and guardrails to measure accuracy, groundedness, safety, regression and hallucination rates, and drive improvements from the results.
  • Optimize inference cost and latency through caching, batching, model routing, quantization, streaming and speculative decoding.
  • Deploy and serve models using vLLM, TGI, Triton or managed inference, including self-hosted open-weight deployments for data-sensitive clients.
  • Select and integrate foundation models (OpenAI, Anthropic, Google, Llama, Mistral, Qwen) against accuracy, latency, cost and data-residency requirements.
  • Expose AI capabilities as secure, well-documented, streaming-capable APIs and integrate them into enterprise applications.
  • Operate AI services in production with containers and CI/CD, including tracing, token and cost telemetry, prompt and model versioning, and rollback.
  • Lead code and design reviews on AI components, mentor engineers, and contribute reusable frameworks and standards to the AI practice.
  • Prototype against new model releases, agent frameworks and inference techniques, and bring what proves out into client delivery.
Essential Skills
Job
  • 4-6 years of engineering experience, with at least 2-3 years building and shipping production LLM or Generative AI systems.
  • Expert Python, covering asynchronous programming, typing, testing, streaming and API development with FastAPI.
  • Deep hands-on LLM application development: prompt and context engineering, structured output, tool calling, function schemas and failure handling.
  • Working understanding of transformer internals: tokenization, attention and context-length costs, embeddings, temperature and sampling, and decoding behaviour.
  • Proven experience designing and shipping RAG systems, including chunking strategies, embedding selection, hybrid retrieval, re-ranking and retrieval evaluation.
  • Hands-on experience with agent and orchestration frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI or the OpenAI/Anthropic agent SDKs.
  • Strong working knowledge of vector databases (e.g., Pinecone, Weaviate, Qdrant, FAISS, pgvector) and index and similarity trade-offs.
  • Practical experience with fine-tuning and model adaptation (LoRA/QLoRA, PEFT, instruction tuning), including when not to fine-tune.
  • Hands-on PyTorch, Hugging Face Transformers and the surrounding ecosystem (datasets, accelerate, PEFT).
  • Experience with LLM evaluation and observability tooling (RAGAS, LangSmith, DeepEval, Langfuse) and building custom eval sets.
  • Experience reducing inference cost and latency in production, with a clear account of what you measured and what you changed.
  • Hands-on experience with at least one major cloud platform (AWS, Azure or GCP) and its AI/ML services, plus Docker and CI/CD.
  • Strong database and API fundamentals across SQL (PostgreSQL) and NoSQL (MongoDB, Redis), with Git and disciplined code review practices.
Personal
  • Strong engineering judgement, with the ability to reason about accuracy, cost and latency trade-offs rather than defaulting to the largest model.
  • Genuine depth of interest in the field: reads papers, model cards and evaluations, and tests claims instead of taking them at face value.
  • Clear communication skills, including the ability to explain AI trade-offs and limitations to non-technical stakeholders.
  • Strong ownership and accountability, with the ability to drive work independently end to end.
  • Able to mentor engineers and lift the technical standard of the team.
  • Comfortable working in fast-paced, ambiguous and rapidly evolving problem spaces.

Preferred Skills
Job
  • Experience with multi-agent orchestration, planning loops and long-running autonomous workflows.
  • Experience with multimodal models covering vision, speech or document understanding, including OCR-heavy document pipelines.
  • Experience with distributed or accelerated training and GPU resource management.
  • Knowledge of graph-based retrieval (GraphRAG), knowledge graphs and hybrid symbolic approaches.
  • Experience with model serving infrastructure, autoscaling GPU workloads and cost governance.
  • Familiarity with responsible AI practices: bias evaluation, red-teaming, PII handling and governance frameworks (EU AI Act, NIST AI RMF).
  • Exposure to microservices architecture, event-driven systems and distributed systems fundamentals.
  • Open-source contributions, published work, or a portfolio of AI systems built outside of client work.
  • Personal
  • Builder's mindset, comfortable prototyping quickly and discarding what does not hold up.
  • Consulting orientation, balancing technical ideals against client timelines and budgets.
Other Relevant Information
  • Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field.
  • Relevant certifications in AI/ML or cloud platforms (AWS/Azure/GCP) are a plus.
  • A portfolio of production Generative AI systems, open-source contributions or published work is highly desirable.

Benefits
  • This role offers the flexibility of working remotely in India.

LeewayHertz is an equal opportunity employer and does not discriminate based on race, colour, religion, sex, age, disability, national origin, sexual orientation, gender identity, or any other protected status. We encourage a diverse range of applicants.


Skills Required

  • 4-6 years of engineering experience
  • At least 2-3 years building and shipping production LLM or Generative AI systems
  • Expert Python, including asynchronous programming, typing, testing, streaming, and FastAPI development
  • Deep hands-on experience with LLM application development, prompt engineering, context engineering, structured output, tool calling, and failure handling
  • Working understanding of transformer internals, tokenization, attention, context-length costs, embeddings, sampling, and decoding
  • Production experience designing and shipping RAG systems, including chunking, embeddings, hybrid retrieval, re-ranking, and retrieval evaluation
  • Experience with agent and orchestration frameworks such as LangChain, LangGraph, LlamaIndex, CrewAI, or OpenAI and Anthropic agent SDKs
  • Strong working knowledge of vector databases, including index and similarity trade-offs
  • Experience with fine-tuning and model adaptation, including LoRA, QLoRA, PEFT, and instruction tuning
  • Hands-on PyTorch and Hugging Face Transformers ecosystem experience
  • Experience with LLM evaluation and observability tools such as RAGAS, LangSmith, DeepEval, or Langfuse
  • Experience reducing inference cost and latency in production
  • Hands-on experience with AWS, Azure, or GCP and associated AI/ML services
  • Experience with Docker and CI/CD
  • Strong SQL and PostgreSQL fundamentals, plus NoSQL experience with MongoDB or Redis
  • Git and disciplined code review practices
  • Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field
  • Experience with multi-agent orchestration, planning loops, and long-running autonomous workflows
  • Experience with multimodal models, vision, speech, document understanding, or OCR-heavy pipelines
  • Experience with distributed or accelerated training and GPU resource management
  • Knowledge of GraphRAG, knowledge graphs, and hybrid symbolic approaches
  • Experience with model serving infrastructure, autoscaling GPU workloads, and cost governance
  • Familiarity with responsible AI practices and governance frameworks such as the EU AI Act and NIST AI RMF
  • Relevant AI/ML or cloud certifications
  • Portfolio of production Generative AI systems, open-source contributions, or published work
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Gurugram, Haryana
113 Employees
Year Founded: 2007

What We Do

Headquartered at San Francisco and founded in 2007, LeewayHertz is one of the first few companies to build and launch a commercial app on Apple's App Store. Our team of certified designers and developers has designed and developed more than 100 digital platforms on Mobile, Cloud, AI, IoT and Blockchain. At LeewayHertz, we have developed digital solutions for Fortune 500 companies and startups to ease their business functions with the latest technologies. Some of our reputed clients include ESPN, NASCAR, Hershey's, McKinsey, P&G, Siemens, 3M, Pearson and more. Being an award-winning software development company, we have also proven our expertise in blockchain development and worked on more than 20+ blockchain projects. We have created a workforce of blockchain developers who can build blockchain apps on different blockchain platforms such as Ethereum, Hyperledger Fabric, Hyperledger Sawtooth, Hyperledger Iroha, Hyperledger Indy, EOS, Stellar, Tron and Corda. We design, develop, deploy and maintain technology products. Uber and Twitter are using our inventions and patents. We work with tech geeks and passionate technologists who are trained by the experts at Apple and Google and always remains at the cutting edge of technology. If you meet this criterion, join us at www.leewayhertz.com

Similar Jobs

Atlassian Logo Atlassian

Marketing Manager

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
India
11000 Employees

Coursera + Udemy  Logo Coursera + Udemy

Content Marketing Manager

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
India
1500 Employees
106K-143K Annually

Atlassian Logo Atlassian

Account Executive

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
India
11000 Employees

Atlassian Logo Atlassian

Manager, Account Executives, Mid-Market

Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Remote
India
11000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account