The Role
Design, deploy, and maintain scalable machine learning systems on GCP. Build training and inference pipelines, automate CI/CD and model deployment with Vertex AI, manage model versioning and rollouts, monitor drift and performance, and support production troubleshooting. Collaborate with data scientists, engineers, and stakeholders to operationalize models, improve platform reliability, and document production-ready workflows.
Summary Generated by Built In
Job Summary
We are looking for an experienced Machine Learning Engineer / MLOps Engineer to design, deploy, and maintain scalable machine learning solutions on Google Cloud Platform (GCP). The ideal candidate will have strong expertise in production ML systems, CI/CD automation, Vertex AI, and model lifecycle management while collaborating closely with cross-functional teams to operationalize machine learning models.
Key Responsibilities:
- Design, build, and maintain training and inference pipelines for storm outage prediction workflows.
- Implement CI/CD, orchestration, and automation for machine learning workflows using Vertex AI and related GCP services.
- Deploy machine learning models into production environments and manage model lifecycle processes, including versioning and rollout support.
- Set up and maintain baseline monitoring for model drift, performance, reliability, and alerting.
- Create scalable, production-ready ML workflows and supporting technical documentation.
- Collaborate with data scientists, engineers, and project stakeholders to operationalize models and align deployment architecture with project needs.
- Support troubleshooting, performance tuning, and continuous improvement of ML platform components.
- Contribute to engineering best practices across code quality, release processes, and environment stability.
Required Qualifications:
- 5–10 years of experience in machine learning engineering, MLOps, or related production ML engineering roles.
- Strong proficiency in Python and experience developing scalable data and ML workflows.
- Hands-on experience with Google Cloud Platform, including Vertex AI, BigQuery, Cloud Storage, Cloud Run, Cloud Functions, or comparable services.
- Experience with GitHub Actions or similar CI/CD tooling.
- Demonstrated experience building CI/CD pipelines and automating ML model deployment and orchestration.
- Experience with model monitoring, observability, and production support for machine learning systems.
- Strong understanding of software engineering best practices, including version control, testing, and documentation.
- Ability to work effectively across distributed teams and communicate clearly with technical and non-technical stakeholders.
Preferred Qualifications:
- Experience supporting forecasting, outage prediction, or other data-intensive operational use cases.
- Familiarity with utility, energy, weather, or geospatial data domains.
- Exposure to model registry, retraining automation, and ML lifecycle governance practices.
- Prior experience working in offshore or globally distributed delivery models.
Skills Required
- 5-10 years of experience in machine learning engineering, MLOps, or related production ML engineering roles
- Strong proficiency in Python and experience developing scalable data and ML workflows
- Hands-on experience with Google Cloud Platform, including Vertex AI, BigQuery, Cloud Storage, Cloud Run, Cloud Functions, or comparable services
- Experience with GitHub Actions or similar CI/CD tooling
- Experience building CI/CD pipelines and automating ML model deployment and orchestration
- Experience with model monitoring, observability, and production support for machine learning systems
- Strong understanding of software engineering best practices, including version control, testing, and documentation
- Ability to work effectively across distributed teams and communicate clearly with technical and non-technical stakeholders
- Experience supporting forecasting, outage prediction, or other data-intensive operational use cases
- Familiarity with utility, energy, weather, or geospatial data domains
- Exposure to model registry, retraining automation, and ML lifecycle governance practices
- Prior experience working in offshore or globally distributed delivery models
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Cognine Technologies is an IT services and technology solutions company with offices across the United States, India, and Europe. It provides application and product development, API and systems integration, cloud and DevOps services, quality assurance, and AI, machine learning, and automation solutions. The company’s work helps organizations build, modernize, integrate, and improve technology platforms and digital products for businesses worldwide.








