Design and implement Retrieval-Augmented Generation (RAG) pipelines and architectures.
Develop document retrieval, contextual augmentation, and chunking strategies for large-scale unstructured data.
Optimize RAG indexing, retrieval accuracy, and context relevance using advanced evaluation metrics.
Implement fine-tuning and prompt engineering techniques to improve retrieval and generation quality.
Manage token limits, context windows, and retrieval latency for high-performance inference.
Integrate LLM frameworks like LangChain or LlamaIndex for pipeline orchestration.
Utilize APIs from OpenAI, Hugging Face Transformers, or other LLM providers for model integration.
Perform noise reduction, diversity sampling, and retrieval optimization to enhance output reliability.
Collaborate with cross-functional teams to deploy scalable RAG-based analytics solutions.
Requirements
Experience in MLOps for deploying and monitoring LLM/RAG-based solutions.
Understanding of semantic search algorithms and context ranking models.
Exposure to knowledge retrieval, contextual augmentation, or multi-document summarization.
Master’s degree in Computer Science, Artificial Intelligence, Data Science, or related field.
Skills Required
- 6+ years of total professional experience
- 2+ years of relevant experience building RAG or LLM-based systems
- Strong hands-on experience with Python
- In-depth understanding of RAG pipelines, RAG architecture, and retrieval optimization
- Practical experience with vector databases such as FAISS, Pinecone, Weaviate, ChromaDB, or Milvus
- Knowledge of generating and fine-tuning embeddings for semantic search and document retrieval
- Experience with LangChain, LlamaIndex, OpenAI API, and Hugging Face Transformers
- Strong understanding of token and context management, retrieval latency, and inference efficiency
- Familiarity with retrieval accuracy, context relevance, and answer faithfulness metrics
- Experience in MLOps for deploying and monitoring LLM/RAG-based solutions
- Understanding of semantic search algorithms and context ranking models
- Exposure to knowledge retrieval, contextual augmentation, or multi-document summarization
- Master's degree in Computer Science, Artificial Intelligence, Data Science, or a related field
What We Do
nHRMS is a strategic human-resources partner providing end-to-end solutions for organizations in the United States and India. Its services span executive search and talent acquisition, performance management, HR advisory, organization strategy, HR technology, leadership development, labor-code compliance, workforce productivity, and learning. The firm supports clients across the employee lifecycle, combining people-first consulting with technology-enabled systems to help organizations scale.
.png)

.jpg)




