Agentic AI Data Engineer
Role Overview
Total Experience required : 5-10 Years
We are seeking a highly skilled Agentic AI Data Engineer to design, build, and optimize intelligent, autonomous data systems that power next-generation AI applications. This role blends data engineering, machine learning infrastructure, and emerging agent-based AI frameworks to enable scalable, self-orchestrating pipelines and decision-making systems.
You will work at the intersection of data platforms, large language models (LLMs), and cloud-native architectures—building systems that can reason, act, and adapt autonomously.
Key Responsibilities
- Design and implement agentic AI systems that autonomously orchestrate data workflows and decision pipelines
- Build scalable data pipelines for structured and unstructured data (batch + real-time)
- Develop and manage LLM-powered applications using retrieval-augmented generation (RAG), tool use, and multi-agent frameworks
- Integrate AWS AI/ML services into production-grade architectures
- Develop and optimize data lakes, warehouses, and lakehouse architectures
- Build APIs and microservices to expose AI/ML capabilities
- Ensure data quality, governance, and security across pipelines
- Collaborate with data scientists, ML engineers, and product teams to deploy AI solutions
- Implement monitoring, logging, and observability for AI agents and pipelines
- Optimize cost and performance of cloud-based AI workloads
Required Technical Skills
Cloud & AWS Ecosystem
- Strong experience with AWS services, including:
- Amazon S3, Glue, Lambda, Step Functions
- Amazon Redshift / Athena
- Amazon SageMaker (training, deployment, pipelines)
- Amazon Bedrock (foundation models, agents, knowledge bases)
AI/ML & Agentic Systems
- Experience with LLMs and generative AI systems
- Hands-on with agent frameworks (e.g., multi-agent orchestration, tool calling, planning systems)
- Familiarity with AgentCore / agent orchestration platforms
- Understanding of RAG architectures, embeddings, and vector databases
- Experience with model deployment, inference optimization, and prompt engineering
Data Engineering
- Strong proficiency in Python and SQL
- Experience with ETL/ELT tools and frameworks
- Distributed data processing (Spark, PySpark, or similar)
- Streaming technologies (Kafka, Kinesis, or similar)
- Data modeling and schema design
Data & AI Infrastructure
- Experience with vector databases (e.g., Pinecone, FAISS, OpenSearch)
- Knowledge of data lakehouse architectures (Delta Lake, Iceberg, Hudi)
- Containerization (Docker) and orchestration (Kubernetes)
- CI/CD for ML and data pipelines
Preferred Qualifications
- Experience building autonomous AI agents for enterprise use cases
- Knowledge of multi-agent collaboration systems and planning algorithms
- Familiarity with LangChain, LlamaIndex, or similar frameworks
- Experience with MLOps and LLMOps practices
- Understanding of graph-based workflows and knowledge graphs
- Exposure to real-time AI systems and event-driven architectures
Soft Skills
- Strong problem-solving and system design skills
- Ability to work in fast-paced, evolving AI environments
- Effective communication and cross-functional collaboration
- Curiosity and adaptability to emerging AI technologies
Education & Experience
- Bachelor’s or Master’s degree in Computer Science, Engineering, or related field
- 4+ years of experience in data engineering or ML engineering
- Hands-on experience with production-grade AI/ML systems
Nice-to-Have
- Experience with reinforcement learning or planning systems
- Background in distributed systems design
- Contributions to open-source AI/data projects
- Certifications in AWS (e.g., Solutions Architect, Machine Learning Specialty)
What You’ll Build
- Autonomous data pipelines that self-heal and optimize
- AI agents capable of reasoning over enterprise data
- Scalable LLM-powered applications integrated with business workflows
- Intelligent systems that move beyond automation into decision-making
Key Responsibilities
- Design and implement agentic AI systems that autonomously orchestrate data workflows and decision pipelines
- Build scalable data pipelines for structured and unstructured data (batch + real-time)
- Develop and manage LLM-powered applications using retrieval-augmented generation (RAG), tool use, and multi-agent frameworks
- Integrate AWS AI/ML services into production-grade architectures
- Develop and optimize data lakes, warehouses, and lakehouse architectures
- Build APIs and microservices to expose AI/ML capabilities
- Ensure data quality, governance, and security across pipelines
- Collaborate with data scientists, ML engineers, and product teams to deploy AI solutions
- Implement monitoring, logging, and observability for AI agents and pipelines
- Optimize cost and performance of cloud-based AI workloads
Required Technical Skills
Cloud & AWS Ecosystem
- Strong experience with AWS services, including:
- Amazon S3, Glue, Lambda, Step Functions
- Amazon Redshift / Athena
- Amazon SageMaker (training, deployment, pipelines)
- Amazon Bedrock (foundation models, agents, knowledge bases)
AI/ML & Agentic Systems
- Experience with LLMs and generative AI systems
- Hands-on with agent frameworks (e.g., multi-agent orchestration, tool calling, planning systems)
- Familiarity with AgentCore / agent orchestration platforms
- Understanding of RAG architectures, embeddings, and vector databases
- Experience with model deployment, inference optimization, and prompt engineering
Data Engineering
- Strong proficiency in Python and SQL
- Experience with ETL/ELT tools and frameworks
- Distributed data processing (Spark, PySpark, or similar)
- Streaming technologies (Kafka, Kinesis, or similar)
- Data modeling and schema design
Data & AI Infrastructure
- Experience with vector databases (e.g., Pinecone, FAISS, OpenSearch)
- Knowledge of data lakehouse architectures (Delta Lake, Iceberg, Hudi)
- Containerization (Docker) and orchestration (Kubernetes)
- CI/CD for ML and data pipelines
Preferred Qualifications
- Experience building autonomous AI agents for enterprise use cases
- Knowledge of multi-agent collaboration systems and planning algorithms
- Familiarity with LangChain, LlamaIndex, or similar frameworks
- Experience with MLOps and LLMOps practices
- Understanding of graph-based workflows and knowledge graphs
- Exposure to real-time AI systems and event-driven architectures
Skills Required
- 5-10 years experience in data engineering or ML engineering
- Bachelor's or Master's degree in Computer Science, Engineering, or related field
- Strong experience with AWS services: S3, Glue, Lambda, Step Functions, Redshift, Athena, SageMaker, Bedrock
- Proficiency in Python
- Proficiency in SQL
- Experience with ETL/ELT tools and frameworks
- Distributed data processing (Spark, PySpark or similar)
- Streaming technologies (Kafka, Kinesis or similar)
- Experience with LLMs and generative AI systems
- Hands-on experience with agent frameworks and AgentCore
- Understanding of RAG architectures, embeddings, and vector databases (Pinecone, FAISS, OpenSearch)
- Experience with model deployment, inference optimization, and prompt engineering
- Experience with data lakehouse architectures (Delta Lake, Iceberg, Hudi) and data modeling
- Containerization (Docker) and orchestration (Kubernetes)
- CI/CD for ML and data pipelines (MLOps/LLMOps)
- Experience building autonomous AI agents for enterprise use cases
- Familiarity with LangChain, LlamaIndex, or similar frameworks
- Experience with reinforcement learning or planning systems
- Background in distributed systems design and contributions to open-source AI/data projects
- AWS certifications (Solutions Architect, Machine Learning Specialty)
What We Do
Choosing a digital partner is about more than capabilities — it’s about collaboration and character. Unrealistic overhauls and off-the-shelf products ignore what matters most — your unique needs, culture, goals, and your legacy data and technology environments. At EXL, our collaboration is built on ongoing listening and learning to adapt our methodologies. We’re your business evolution partner—tailoring solutions that make the most of data to make better business decisions and drive more intelligence into your increasingly digital operations. Whether your goals are scaling the use of AI and digital, redesign operating models, or driving better and faster decisions, we’re here to partner with you to help you gain—and maintain—competitive advantage with efficient, sustainable models at scale. Our expertise in transformation, data science, and change management helps make your business more efficient and effective, improve customer relationships and enhance revenue growth. Instead of focusing on multi-year, resource- and time-intensive platform designs or migrations, we look deeper at your entire value chain to integrate strategies with impact. We use our specialization in analytics, digital interventions, and operations management—alongside deep industry expertise — to deliver solutions that help you outperform the competition. At EXL, it’s all about outcomes—your outcomes—and delivering success on your terms. Share your goals with us and together, we’ll optimize how you leverage data to drive your business forward. For more information, visit www.exlservice.com.






