GenAI / AI-ML Engineer

Posted 9 Days Ago
Be an Early Applicant
Bengaluru, Karnataka, IND
In-Office
Senior level
Cloud
The Role
Design, develop, deploy, and operate production-grade machine learning, Generative AI, LLM, RAG, and agentic AI solutions on AWS. Build and evaluate models and pipelines, integrate enterprise systems, optimize quality, latency, reliability, security, and cost, and use Bedrock and SageMaker for production AI architectures. Monitor model performance and drift while collaborating with engineering, data, architecture, and product teams.
Summary Generated by Built In
Role summary
We are seeking an experienced GenAI / AI-ML Engineer to design, build, deploy and operate production-grade AI/ML and Generative AI solutions on AWS. The role requires strong Python engineering, solid understanding of conventional machine-learning algorithms, practical experience with LLMs, RAG and agentic AI, and strong hands-on expertise with AWS AI/ML and Generative AI services.
The engineer will be responsible for taking AI solutions from problem definition and solution design through development, evaluation, production deployment and optimization. The candidate should be capable of making architecture and technology decisions based on scalability, latency, reliability, security, model quality and cost.
The role also requires a strong understanding of LLM token consumption, model pricing and AWS service/infrastructure costing, with the ability to design solutions that balance business requirements, technical performance and operational cost.

Key responsibilities

• Translate business problems into measurable AI/ML and Generative AI objectives, solution approaches and evaluation criteria.
• Design and implement production-grade LLM, RAG, conversational AI and agentic AI solutions on AWS.
• Build RAG pipelines including document ingestion, chunking, embeddings, metadata filtering, vector retrieval, reranking and generation.
• Design and implement Amazon Bedrock Agents and Bedrock AgentCore-based agentic solutions, including tool use, memory/state, runtime execution and orchestration as applicable.
• Work with Amazon Bedrock foundation models and evaluate models based on quality, latency, token consumption and cost.
• Build integrations between LLMs/agents and enterprise APIs, databases, applications, tools and knowledge sources.
• Develop and evaluate supervised, unsupervised and deep-learning models based on business requirements.
• Apply appropriate machine-learning algorithms for classification, regression, clustering, anomaly detection and other relevant use cases.
• Perform data preprocessing, feature engineering, model training, validation, hyperparameter tuning and error analysis.
• Use Amazon SageMaker for appropriate ML development, training, experimentation, deployment and model lifecycle requirements.
• Establish evaluation mechanisms for LLM response quality, retrieval quality, groundedness, hallucination, latency and cost.
• Analyze and optimize LLM input/output token consumption, context size and model selection to control GenAI costs.
• Estimate and optimize AWS infrastructure and service costs across model inference, compute, storage, APIs, vector search and other components.
• Design end-to-end AWS architectures using appropriate managed AI/ML, compute, data, integration and security services.
• Implement security, access control, logging, monitoring, observability and operational controls for production AI systems.
• Deploy and operate AI/ML solutions on AWS and troubleshoot latency, scalability, reliability, model quality, infrastructure and cost issues.
• Monitor deployed ML models for performance, drift, reliability and business-aligned metrics.
• Collaborate with data engineers, software engineers, architects and product teams to build reliable production solutions.
• Produce reusable components, technical documentation, architecture decisions and implementation guidance for delivery teams.


Required qualifications and experience

• 5+ years of relevant experience in AI/ML engineering, machine learning, Generative AI, applied data science or software engineering.
• Strong hands-on Python development experience with good understanding of object-oriented design, API integration, debugging and production engineering.
• Demonstrable experience taking AI/ML or GenAI solutions from POC/experimentation through production deployment.
• Mandatory hands-on AWS experience implementing and deploying AI/ML or GenAI solutions.
• Strong hands-on experience with Amazon Bedrock and foundation models.
• Hands-on experience with Amazon Bedrock AgentCore and/or Bedrock Agents for agentic AI solutions.
• Hands-on experience with Amazon SageMaker for ML model development, training, deployment or lifecycle management.
• Experience with AWS services supporting GenAI architectures, including Amazon S3, Amazon OpenSearch Service/Serverless, AWS Lambda, API Gateway, Step Functions, CloudWatch, IAM, KMS and Secrets Manager.
• Strong understanding of LLMs, embeddings, vector search, prompt/context engineering and RAG.
• Working experience with agentic AI, tool/function calling, agent orchestration and state management.
• Strong understanding of conventional machine-learning algorithms, model evaluation and statistical reasoning.
• Ability to select appropriate algorithms, models and AWS services based on the business problem.
• Strong understanding of trade-offs between model quality, accuracy, latency, scalability, reliability and cost.
• Understanding of LLM token economics, including input/output tokens, context consumption and model pricing.
• Ability to understand and estimate AWS service/infrastructure costs for an AI/ML solution.
• Ability to translate business requirements into secure, scalable, maintainable and cost-effective AWS AI/ML architectures.

Mandatory Skills
  • Python, Pandas, NumPy, Scikit-learn
  • Conventional ML – Classification, Regression, Clustering, Anomaly Detection
  • Feature Engineering, Model Training & Evaluation
  • Cross-validation, Hyperparameter Tuning, Error Analysis
  • AWS – Bedrock, SageMaker, Knowledge Bases, OpenSearch/OpenSearch Serverless
  • RAG, Embeddings, Vector Retrieval
  • Agentic Workflows, Tool/Function Calling, State Management
  • LLM Evaluation, Groundedness, Hallucination Mitigation
  • LLM Token Consumption & Cost Estimation
  • AWS Service/Infrastructure Cost Estimation & Optimization

Good-to-Have Skills
  • Docker, CI/CD, Kubernetes/EKS
  • MLflow / SageMaker MLflow
  • PyTorch/TensorFlow
  • Advanced MLOps & Model Monitoring
  • Fine-tuning, PEFT, LoRA, QLoRA
  • Advanced NLP, Computer Vision, Forecasting
  • Bedrock Agents / AgentCore
  • Multimodal GenAI
  • Voice AI / Conversational AI
  • GraphRAG / Knowledge Graphs
  • Bedrock Guardrails / Evaluation
  • MCP / A2A & Agent Interoperability
  • Advanced LLM Observability & Evaluation
  • Advanced Multi-Agent Architectures


Skills Required

  • 5+ years of relevant experience in AI/ML engineering, machine learning, Generative AI, applied data science, or software engineering
  • Strong hands-on Python development experience, including object-oriented design, API integration, debugging, and production engineering
  • Experience taking AI/ML or Generative AI solutions from experimentation or proof of concept through production deployment
  • Hands-on AWS experience implementing and deploying AI/ML or Generative AI solutions
  • Strong hands-on experience with Amazon Bedrock and foundation models
  • Experience with Amazon Bedrock AgentCore and/or Bedrock Agents
  • Hands-on experience with Amazon SageMaker for model development, training, deployment, or lifecycle management
  • Experience with AWS services including S3, OpenSearch, Lambda, API Gateway, Step Functions, CloudWatch, IAM, KMS, and Secrets Manager
  • Strong understanding of LLMs, embeddings, vector search, prompt and context engineering, and RAG
  • Experience with agentic AI, tool or function calling, agent orchestration, and state management
  • Strong understanding of conventional machine-learning algorithms, model evaluation, and statistical reasoning
  • Experience with classification, regression, clustering, anomaly detection, feature engineering, model training, validation, hyperparameter tuning, and error analysis
  • Understanding of LLM token economics, model pricing, and AWS service or infrastructure cost estimation
  • Ability to design secure, scalable, maintainable, and cost-effective AWS AI/ML architectures
  • Experience with Docker, CI/CD, and Kubernetes or EKS
  • Experience with MLflow, PyTorch, TensorFlow, advanced MLOps, model monitoring, fine-tuning, PEFT, LoRA, or QLoRA
  • Experience with advanced NLP, computer vision, forecasting, multimodal GenAI, voice AI, GraphRAG, knowledge graphs, Bedrock Guardrails, MCP, A2A, or multi-agent architectures
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tampa, FL
375 Employees
Year Founded: 2012

What We Do

CloudThat is the first company in India to offer Cloud Training & Cloud Consultancy services for mid-market & enterprise clients from across the globe. Founded by Bhavesh Goswami, an ex-Amazonian with more than 17 years of experience in the Cloud Computing space, CloudThat has been empowering tech professionals with the best cloud training and cloud consulting services. Our journey in the Cloud space has won us many accolades, such as the prestigious Finalist of Microsoft Learning Partner of the Year Awards 2022 and 2020. CloudThat is also a proud AWS Advanced Consulting Partner, VMware Authorized Training Reseller, Google Cloud Platform Partner, Microsoft Gold Partner and Databricks Partner. Since our inception in 2012, we have trained over 500K IT professionals from fortune 500 companies in different types of cloud computing technologies such as Microsoft Azure, Amazon Web Services, Artificial Intelligence, Machine Learning, Google Cloud Platform, IoT, OpenStack, and OpenShift, DevOps, MongoDB, Big Data and more. We have also delivered over 200 AWS consulting services to top companies in need of cloud solutions. We have a global presence with offices in Bengaluru (India), the UK, and the USA. Our on-site and pre-scheduled public batches are available in different IT-centric cities in India and Overseas.

Similar Jobs

Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

BlackRock Logo BlackRock

Managing Director, Global Head of Derived Data

Fintech • Information Technology • Financial Services
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
25000 Employees

The Aerospace Corporation Logo The Aerospace Corporation

Integrity &Compl Specialist II

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Remote or Hybrid
India
4600 Employees

DigitalOcean Logo DigitalOcean

Engineering Manager

Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
1400 Employees

Similar Companies Hiring

Rundoo Thumbnail
Cloud • Information Technology • Internet of Things • Software
Redwood City, California
70 Employees
NetBox Labs Thumbnail
Cloud • Software
US
125 Employees
Toro TMS Thumbnail
Cloud • Enterprise Web • Sales • Software • Transportation
Chicago, IL
80 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account