The Role
Architect and build enterprise-scale MLOps and AI platforms supporting thousands of production ML models and high-volume distributed workloads. Define architectures for feature stores, MLflow, model registries, distributed training, deployment, inference, monitoring, governance, lineage, explainability, and automated retraining. Establish secure, scalable Databricks Lakehouse standards using Unity Catalog, Delta Lake, DLT, and Mosaic AI. Partner with data teams, optimize reliability and cost, and mentor engineering teams.
Summary Generated by Built In
Job Role: MLOps Architect
Location: Hyderabad (Hybrid)
Experience: 12–18 Years
Role Objective
We are looking for an experienced MLOps Architect to design and build enterprise-scale AI/ML platforms from the ground up. This role will define the end-to-end ML Operations architecture, enabling scalable, secure, governed, and production-ready ML platforms supporting 10,000+ ML models and 750+ trillion records.
The ideal candidate will have deep expertise in Machine Learning Operations Architecture, Databricks, MLflow, Feature Stores, Model Governance, and Enterprise AI Platforms, with proven experience operationalizing large-scale ML workloads in production.
Key Responsibilities
- Architect and build enterprise-scale MLOps platforms from scratch supporting the complete Machine Learning lifecycle.
- Define architecture for Feature Engineering, Feature Stores, Experiment Tracking, MLflow, Model Registry, Distributed Training, Hyperparameter Optimization, Model Deployment, Batch & Real-time Inference, Monitoring, Drift Detection, Explainability, AI Governance, Lineage, and Automated Retraining.
- Design scalable AI platforms capable of supporting thousands of production ML models and large-scale distributed AI workloads.
- Architect enterprise Lakehouse & AI platforms using Databricks, Unity Catalog, Delta Lake, Delta Live Tables (DLT), MLflow, and Mosaic AI.
- Define standards for Model Lifecycle Management, AI Governance, Responsible AI, Observability, Lineage, Security, and Compliance.
- Build scalable Feature Store and Model Serving architectures for batch, streaming, and real-time inference.
- Partner with Data Science, Data Engineering, and Enterprise Architecture teams to operationalize ML models into production.
- Drive platform scalability, reliability, performance, availability, and cost optimization.
- Mentor MLOps, ML, and Data Engineering teams while establishing enterprise architecture standards and best practices.
Required Skills & Experience
- 12–18 years of experience in MLOps, Machine Learning Platform Engineering, AI Platform Architecture, or Data Engineering.
- Proven experience architecting and implementing enterprise MLOps platforms from scratch.
- Deep expertise in architecting the end-to-end ML lifecycle, including Feature Engineering, Feature Stores, Experiment Tracking, MLflow, Model Registry, Distributed Training, Hyperparameter Optimization, Model Deployment, Batch & Real-time Inference, Monitoring, Drift Detection, Explainability, AI Governance, Lineage, and Automated Retraining for enterprise-scale AI platforms.
- Strong expertise with Databricks, MLflow, Unity Catalog, Delta Lake, Delta Live Tables (DLT), Mosaic AI, Apache Spark, PySpark, Structured Streaming, and Lakehouse Architecture.
- Experience building large-scale AI platforms supporting thousands of production ML models and high-volume distributed data workloads.
- Strong programming skills in Python, PySpark, SQL, and distributed data processing.
- Experience working on Azure (Preferred), AWS, or GCP.
- Working knowledge of Kubernetes, Docker, Terraform, Linux, and cloud-native platforms to support scalable ML workloads.
Good to Have
- Experience with GenAI, LLMOps, RAG, LangChain, LangGraph, Vector Databases, NVIDIA AI Stack, AI Agents, and modern AI frameworks.
- Experience building AI platforms supporting 10,000+ ML models, petabyte-scale data, or hyperscale enterprise workloads.
- Databricks, AWS, Azure, GCP, Kubernetes, or Terraform certifications are a plus.
Skills Required
- 12–18 years of experience in MLOps, machine learning platform engineering, AI platform architecture, or data engineering
- Experience architecting and implementing enterprise MLOps platforms from scratch
- Expertise across the end-to-end machine learning lifecycle, including feature stores, experiment tracking, MLflow, model deployment, inference, monitoring, governance, lineage, and automated retraining
- Strong expertise with Databricks, MLflow, Unity Catalog, Delta Lake, Delta Live Tables, Mosaic AI, Apache Spark, PySpark, Structured Streaming, and Lakehouse architecture
- Experience building large-scale AI platforms supporting thousands of production ML models and high-volume distributed data workloads
- Strong programming skills in Python, PySpark, SQL, and distributed data processing
- Experience working on Azure, AWS, or GCP
- Working knowledge of Kubernetes, Docker, Terraform, Linux, and cloud-native platforms
- Experience with GenAI, LLMOps, RAG, LangChain, LangGraph, vector databases, NVIDIA AI Stack, AI agents, and modern AI frameworks
- Experience supporting 10,000+ ML models, petabyte-scale data, or hyperscale enterprise workloads
- Databricks, AWS, Azure, GCP, Kubernetes, or Terraform certifications
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Anblicks is a Cloud Data Analytics Company based out of Dallas, TX, with offices in USA, India, and Australia. Since 2004, Anblicks has been helping customers by bringing value to their data and implementing modern data architecture and advanced analytics solutions in the cloud.






