𝗧𝗵𝗶𝘀 𝗿𝗼𝗹𝗲 𝗶𝘀 𝗳𝗼𝗿 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗪𝗲𝗲𝗸𝗱𝗮𝘆'𝘀 𝗰𝗹𝗶𝗲𝗻𝘁𝘀
𝗦𝗮𝗹𝗮𝗿𝘆 𝗿𝗮𝗻𝗴𝗲: 𝗥𝘀 𝟭𝟰𝟬𝟬𝟬𝟬𝟬 - 𝗥𝘀 𝟯𝟮𝟬𝟬𝟬𝟬𝟬 (𝗶𝗲 𝗜𝗡𝗥 𝟭𝟰-𝟯𝟮 𝗟𝗣𝗔)
Experience: 2+ yrs
Location: Bengaluru, Karnataka, India
Job Type: Full-time
We are looking for a technically strong DataOps / MLOps Engineer to build, deploy, operate, and scale modern data, machine learning, and Generative AI infrastructure in enterprise cloud environments.
The role focuses on creating reliable and automated engineering pipelines across DataOps, MLOps, LLMOps, and AI platforms, with strong emphasis on Databricks and AWS. The ideal candidate will have hands-on experience with cloud data platforms, CI/CD automation, model deployment, observability, infrastructure management, and production operations.
RequirementsKEY RESPONSIBILITIES
- Design, build, and maintain scalable DataOps, MLOps, and LLMOps pipelines for data, ML, and Generative AI workloads.
- Automate data provisioning, model deployment, model evaluation, monitoring, and enterprise workflow execution.
- Implement and maintain robust CI/CD pipelines for data pipelines, ML models, containerised services, and AI applications.
- Manage, configure, optimise, and scale Databricks workspaces, clusters, jobs, and enterprise data-processing environments.
- Collaborate with Data Engineers, ML Engineers, Software Engineers, and Product teams to deploy data pipelines, feature stores, ML models, and serving systems into production.
- Implement proactive monitoring, event instrumentation, alerting, and self-healing mechanisms for data and model quality issues.
- Support incident response, production troubleshooting, infrastructure upgrades, capacity planning, and cloud resource optimisation.
- Work with DevOps, SRE, IT, and Security teams to implement governance, data lineage, compliance, security, and enterprise deployment standards.
- Monitor and troubleshoot distributed data pipelines, ETL workflows, model-serving systems, and production ML infrastructure.
- Build and maintain observability and model-monitoring capabilities using tools such as MLflow, Weights & Biases, and LangSmith.
- Develop hands-on Proofs of Concept (POCs) for modern data platforms, feature stores, streaming technologies, and ML infrastructure.
- Implement infrastructure provisioning and configuration management using tools such as Terraform, CloudFormation, or Ansible.
- Support containerised applications and ML workloads using Docker and Kubernetes.
- Participate in code reviews, engineering design discussions, on-call support, and knowledge-sharing initiatives.
- Identify opportunities to improve reliability, scalability, automation, cost efficiency, and engineering productivity.
- Stay current with evolving DataOps, MLOps, LLMOps, cloud, AI, and data-platform technologies.
- Contribute to the continuous modernisation of enterprise data and ML infrastructure.
- 2–5 years of experience across DataOps, MLOps, ML Engineering, or Data Engineering in enterprise cloud environments.
- Strong hands-on expertise in Databricks administration, workspace management, cluster optimisation, jobs, and enterprise data workloads.
- Strong experience with AWS data and ML services, including services such as SageMaker, Glue, EMR, Athena, and S3.
- Solid understanding of DevOps, DataOps, MLOps, and LLMOps methodologies and practices.
- Proven experience building CI/CD automation pipelines for containerised Python, Java, or Scala applications, microservices, and ML-serving systems.
- Strong understanding of model lifecycle management, evaluation, monitoring, governance, and observability.
- Experience with tools such as MLflow, Weights & Biases, LangSmith, or similar platforms.
- Strong understanding of databases, replication, relational and NoSQL databases, and vector databases such as Pinecone, FAISS, Milvus, or Weaviate.
- Practical experience deploying, monitoring, debugging, and supporting distributed data pipelines and ETL workflows.
- Strong Git knowledge and familiarity with standard branching and collaborative development workflows.
- Experience with Terraform, CloudFormation, Ansible, or similar infrastructure-as-code and configuration-management tools.
- Proficiency in at least one scripting/programming language such as Python, Bash, or JavaScript.
- Hands-on experience with Docker and exposure to Kubernetes orchestration.
- Good understanding of the machine learning lifecycle, feature engineering, and ML deployment workflows.
- Familiarity with PyTorch, TensorFlow, NLP, computer vision, and Generative AI concepts is desirable.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent cross-functional communication and collaboration skills.
- Demonstrated ability to promote a collaborative DevOps/DataOps/MLOps culture.
- Strong ownership mindset with the ability to support production systems and participate in on-call responsibilities.
- Ability to adapt quickly to evolving technology stacks and contribute to modernisation initiatives.
Skills Required
- 2-5 years of experience in DataOps, MLOps, ML Engineering, or Data Engineering in enterprise cloud environments
- Hands-on expertise with Databricks administration, workspace management, cluster optimization, jobs, and enterprise data workloads
- Experience with AWS data and ML services, including SageMaker, Glue, EMR, Athena, and S3
- Understanding of DevOps, DataOps, MLOps, and LLMOps methodologies
- Experience building CI/CD automation pipelines for Python, Java, or Scala applications, microservices, and ML-serving systems
- Understanding of model lifecycle management, evaluation, monitoring, governance, and observability
- Experience with MLflow, Weights & Biases, LangSmith, or similar platforms
- Understanding of relational, NoSQL, and vector databases, including Pinecone, FAISS, Milvus, or Weaviate
- Experience deploying, monitoring, debugging, and supporting distributed data pipelines and ETL workflows
- Strong Git knowledge and familiarity with branching and collaborative development workflows
- Experience with Terraform, CloudFormation, Ansible, or similar infrastructure-as-code tools
- Proficiency in at least one of Python, Bash, or JavaScript
- Hands-on Docker experience and exposure to Kubernetes orchestration
- Understanding of the machine learning lifecycle, feature engineering, and ML deployment workflows
- Familiarity with PyTorch, TensorFlow, NLP, computer vision, and Generative AI concepts
- Strong troubleshooting, analytical, problem-solving, communication, and collaboration skills
- Ability to support production systems and participate in on-call responsibilities
- Ability to adapt to evolving technology stacks and contribute to modernization initiatives
What We Do
Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.









