This is a remote position.
RequirementsResponsibilities
- Own the end-to-end architecture of GenAI solutions across the retrieval, orchestration, model, integration, and deployment layers.
- Translate ambiguous business problems into AI solution designs with clear scope, feasibility assessment and success metrics.
- Define reference architectures, design patterns and reusable accelerators for RAG, agentic workflows and LLM integration.
- Lead model and platform selection, documenting the cost, latency, accuracy and data-residency trade-offs behind each decision.
- Design the non-functional envelope: scalability, latency budgets, availability, observability and inference cost control.
- Architect data and retrieval pipelines covering ingestion, chunking, embedding strategy, vector store selection and hybrid search.
- Define evaluation strategy and guardrails so accuracy, groundedness, safety and hallucination rates can be measured and governed.
- Embed security, privacy and compliance into the design: PII handling, tenancy isolation, access control and audit.
- Support pre-sales and discovery through solution workshops, effort estimation, technical proposals and client presentations.
- Guide delivery teams, run design reviews and mentor engineers, while staying hands-on in prototyping and unblocking hard problems.
- Maintain architecture documentation and decision records, and assess which advances in generative AI are ready for enterprise adoption.
- Own multiple client engagements simultaneously while maintaining delivery quality.
- Lead discovery workshops, challenge assumptions, and refine business requirements into technically sound solutions.
- Push back on unrealistic timelines, architectures, or requirements using engineering judgement and data.
- Build strong relationships with Team, product owners, and executive stakeholders.
- Mentor senior engineers and cultivate future architects and technical leaders.
- Lead architectural governance, design reviews, and technical decision records.
- Set engineering standards, coding guidelines, AI development best practices, and review critical code.
- Remain hands-on by building prototypes, solving difficult technical problems, and contributing production-quality code when needed.
- Drive cross-project reuse through internal frameworks, accelerators, and reference implementations.
- Present architecture, trade-offs, risks, and implementation strategy confidently to executive audiences.
Job
- 10+ years in software engineering, data or AI roles, including at least 3 years in an architect or technical lead capacity.
- Demonstrated experience architecting and delivering production Generative AI systems, not only prototypes or POCs.
- Strong hands-on Python, with the ability to prototype designs and review production code.
- Deep expertise in LLM application architecture: prompt and context engineering, structured output, tool calling and orchestration.
- Proven experience designing RAG systems end-to-end, including chunking, embedding selection, hybrid retrieval and re-ranking.
- Proven ability to manage multiple enterprise AI programs simultaneously.
- Strong client-facing consulting experience with executive communication.
- Excellent presentation, whiteboarding, and workshop facilitation skills.
- Demonstrated experience influencing technical decisions across multiple teams.
- Experience managing senior engineers and mentoring future technical leaders.
- Strong engineering judgement balancing quality, cost, delivery timelines, and business value.
- Comfortable making architectural decisions with incomplete information.
- Experience with agent and orchestration frameworks such as LangChain, LangGraph, LlamaIndex or CrewAI.
- Strong knowledge of vector databases (Pinecone, Weaviate, Qdrant, FAISS, pgvector) and their operational trade-offs.
- Solid machine learning and deep learning fundamentals, including fine-tuning and adaptation approaches (LoRA/QLoRA, PEFT).
- Strong cloud architecture skills on AWS, Azure or GCP, including their AI/ML and data services.
- Experience with microservices, API design, event-driven patterns and enterprise system integration.
- Working knowledge of MLOps and LLMOps: CI/CD, containerization, model versioning, monitoring and rollback.
- Experience defining LLM evaluation and observability approaches (RAGAS, LangSmith, DeepEval or equivalent).
- Personal
- Strong communication and stakeholder management skills, including with non-technical and client-side audiences.
- Sound engineering judgement, with the confidence to defend a design and the openness to revise it.
- Strong ownership across the full delivery lifecycle, not only the design phase.
- Ability to mentor engineers and lead through influence rather than authority.
- Comfortable operating with ambiguity in a fast-moving technology space.
- Strong communication and stakeholder management skills, including with non-technical and client-side audiences.
- Sound engineering judgement, with the confidence to defend a design and the openness to revise it.
- Strong ownership across the full delivery lifecycle, not only the design phase.
- Ability to mentor engineers and lead through influence rather than authority.
- Comfortable operating with ambiguity in a fast-moving technology space
Job
- Experience architecting multi-agent systems and complex autonomous workflows.
- Exposure to multimodal AI covering vision, speech or document understanding.
- Experience with inference optimization and serving at scale (vLLM, TensorRT-LLM, Triton, quantization).
- Experience deploying open-weight models on-premise or in-VPC for data-sensitive clients.
- Knowledge of graph-based retrieval (GraphRAG) and knowledge-graph modelling.
- Familiarity with AI governance and responsible AI frameworks (EU AI Act, NIST AI RMF, ISO/IEC 42001).
- Experience with data platform architecture and pipelines (Airflow, dbt, Spark, lakehouse patterns).
- Pre-sales, solutioning or client-facing consulting experience in a services organisation.
- Domain depth in one or more of BFSI, healthcare, retail, supply chain or manufacturing.
- Personal
- Proactive mindset with a genuine interest in tracking a fast-moving field.
- Consulting orientation, balancing technical ideals against client timelines and budgets.
- Proactive mindset with a genuine interest in tracking a fast-moving field.
- Consulting orientation, balancing technical ideals against client timelines and budgets.
- Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field.
- Relevant certifications in AI/ML, cloud architecture (AWS/Azure/GCP), or enterprise architecture are a plus.
- A portfolio of production Generative AI architectures, open-source contributions or published work is highly desirable.
Benefits
- This role offers the flexibility of working remotely in India.
Skills Required
- 10+ years of experience in software engineering, data, or AI roles
- At least 3 years of experience in an architect or technical lead capacity
- Experience architecting and delivering production Generative AI systems
- Strong hands-on Python development and production code review experience
- Expertise in LLM application architecture, including prompt engineering, context engineering, structured output, tool calling, and orchestration
- Experience designing end-to-end RAG systems, including chunking, embeddings, hybrid retrieval, and re-ranking
- Ability to manage multiple enterprise AI programs simultaneously
- Client-facing consulting experience and executive communication skills
- Presentation, whiteboarding, and workshop facilitation skills
- Experience influencing technical decisions across multiple teams
- Experience managing senior engineers and mentoring technical leaders
- Experience with agent and orchestration frameworks such as LangChain, LangGraph, LlamaIndex, or CrewAI
- Knowledge of vector databases such as Pinecone, Weaviate, Qdrant, FAISS, or pgvector
- Machine learning and deep learning fundamentals, including fine-tuning and adaptation methods such as LoRA, QLoRA, and PEFT
- Cloud architecture experience on AWS, Azure, or GCP, including AI/ML and data services
- Experience with microservices, API design, event-driven patterns, and enterprise system integration
- Working knowledge of MLOps and LLMOps, including CI/CD, containerization, model versioning, monitoring, and rollback
- Experience defining LLM evaluation and observability approaches using RAGAS, LangSmith, DeepEval, or equivalent tools
- Strong communication, stakeholder management, engineering judgment, ownership, mentoring, and ability to operate with ambiguity
- Experience architecting multi-agent systems and complex autonomous workflows
- Experience with multimodal AI involving vision, speech, or document understanding
- Experience with inference optimization and serving at scale using vLLM, TensorRT-LLM, Triton, or quantization
- Experience deploying open-weight models on-premises or in a VPC
- Knowledge of GraphRAG and knowledge-graph modeling
- Familiarity with AI governance and responsible AI frameworks, including EU AI Act, NIST AI RMF, or ISO/IEC 42001
- Experience with data platform architecture and pipelines using Airflow, dbt, Spark, or lakehouse patterns
- Pre-sales, solutioning, or client-facing consulting experience in a services organization
- Domain experience in BFSI, healthcare, retail, supply chain, or manufacturing
- Bachelor’s or Master’s degree in Computer Science, Data Science, Engineering, or a related field
- Relevant AI/ML, cloud architecture, or enterprise architecture certifications
- Portfolio of production Generative AI architectures, open-source contributions, or published work
What We Do
Headquartered at San Francisco and founded in 2007, LeewayHertz is one of the first few companies to build and launch a commercial app on Apple's App Store. Our team of certified designers and developers has designed and developed more than 100 digital platforms on Mobile, Cloud, AI, IoT and Blockchain. At LeewayHertz, we have developed digital solutions for Fortune 500 companies and startups to ease their business functions with the latest technologies. Some of our reputed clients include ESPN, NASCAR, Hershey's, McKinsey, P&G, Siemens, 3M, Pearson and more. Being an award-winning software development company, we have also proven our expertise in blockchain development and worked on more than 20+ blockchain projects. We have created a workforce of blockchain developers who can build blockchain apps on different blockchain platforms such as Ethereum, Hyperledger Fabric, Hyperledger Sawtooth, Hyperledger Iroha, Hyperledger Indy, EOS, Stellar, Tron and Corda. We design, develop, deploy and maintain technology products. Uber and Twitter are using our inventions and patents. We work with tech geeks and passionate technologists who are trained by the experts at Apple and Google and always remains at the cutting edge of technology. If you meet this criterion, join us at www.leewayhertz.com


.png)






