Senior MLOps / ML Platform Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Kraków, Małopolskie, POL
In-Office or Remote
Senior level
Software
The Role
Build and maintain production-grade ML platform infrastructure, including training orchestration, model lifecycle automation, multi-tenant environments, deployment workflows, observability, drift monitoring, reproducibility, cost tracking, and CI/CD automation. Collaborate with DevOps and SRE teams while documenting operational processes and supporting platform handover.
Summary Generated by Built In
Company Description

We are looking for a Senior MLOps Engineer to join Sigma Software and help build a production-grade ML platform for a large-scale AdTech ecosystem. You will work on infrastructure powering predictive decision-making systems that process hundreds of millions of auction requests daily.

As part of a dedicated engineering team, you will contribute to scalable ML orchestration, model lifecycle automation, observability, and real-time optimization workflows. This role is ideal for engineers with strong production experience who enjoy solving complex platform and operational challenges.

We at Sigma Software offer the opportunity to work on cutting-edge ML infrastructure projects, collaborate with experienced engineers, and influence architecture decisions in a long-term strategic engagement.

CUSTOMER

Our Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a high-load ad exchange platform handling hundreds of millions of auction requests every day and is investing in advanced predictive decision-making capabilities to improve advertiser outcomes and real-time optimization processes.

PROJECT

Sigma Software is building a predictive modeling and optimization platform integrated with a live ad exchange environment. The solution enables real-time supply scoring and filtering, audience look-alike generation, contextual performance estimation, and multi-objective optimization under operational constraints.

The project combines large-scale ML infrastructure, automated model lifecycle management, multi-tenant architecture, and advanced observability practices. The team focuses on delivering reliable, reproducible, and scalable ML systems ready for long-term Customer ownership.

Key Technologies: Python, Kubernetes, Docker, GCP, Vertex AI, MLflow, Airflow, Kubeflow, Argo Workflows, Terraform

Job Description

  • Build and maintain ML training orchestration pipelines across hourly, daily, and weekly schedules
  • Implement retries, backfills, and idempotent execution mechanisms
  • Design and support model registry workflows including versioning, lineage, evaluation gates, and promotion processes
  • Develop isolated per-advertiser model environments with namespace and configuration separation
  • Build scalable refresh pipelines and publishing workflows for serving infrastructure
  • Implement shadow mode and champion/challenger deployment strategies
  • Develop monitoring and alerting for ML-specific metrics including feature drift, prediction drift, train/serve skew, and calibration decay
  • Ensure reproducibility of ML workflows using containerized environments, pinned dependencies, and data snapshots
  • Monitor training and scoring costs across tenants
  • Collaborate with DevOps and SRE engineers on CI/CD and infrastructure automation
  • Prepare operational documentation and platform handover materials

Qualifications

  • 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems
  • Strong Python skills and experience building platform-level tooling and automation
  • Hands-on experience with Kubernetes and Docker
  • Experience building CI/CD pipelines for ML workloads
  • Hands-on production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar orchestration and ML lifecycle platforms
  • Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow
  • Strong understanding of ML observability including drift detection, train/serve skew monitoring, and incident response
  • Experience designing or supporting multi-tenant ML systems and isolated model environments
  • Experience working with cloud platforms, preferably GCP
  • Experience with infrastructure-as-code tools such as Terraform
  • Experience with Linux environments
  • Understanding of the ML lifecycle and productionization processes
  • Upper-Intermediate English level or higher

WILL BE A PLUS

  • Experience with feature stores and feature consistency management
  • Experience with large-scale batch scoring systems operating under freshness SLAs
  • Familiarity with experiment tracking platforms and evaluation gates
  • Experience with on-premises Kubernetes or bare-metal Linux infrastructure
  • Knowledge of DVC, lakeFS, or other data versioning tools
  • Experience with Bigtable, Redis, Aerospike, or similar low-latency serving databases
  • GPU scheduling and training cost optimization experience
  • Familiarity with SOC 2, ISO 27001, or GDPR-related compliance requirements

Additional Information

PERSONAL PROFILE

  • Strong ownership mindset and focus on operational reliability
  • Ability to work independently in complex distributed systems environments
  • Strong collaboration and communication skills
  • Analytical thinking with attention to scalability and maintainability
  • Comfortable working in fast-paced product-oriented environments

Skills Required

  • 5+ years of experience in MLOps, ML platform engineering, or infrastructure engineering supporting production ML systems
  • Strong Python skills and experience building platform-level tooling and automation
  • Hands-on experience with Kubernetes and Docker
  • Experience building CI/CD pipelines for ML workloads
  • Production experience with MLflow, Kubeflow, Airflow, Argo Workflows, Vertex Pipelines, or similar platforms
  • Experience with ML platforms and model lifecycle tools such as Vertex AI, MLflow, or Kubeflow
  • Strong understanding of ML observability, drift detection, train/serve skew monitoring, and incident response
  • Experience designing or supporting multi-tenant ML systems and isolated model environments
  • Experience working with cloud platforms, preferably GCP
  • Experience with infrastructure-as-code tools such as Terraform
  • Experience with Linux environments
  • Understanding of the ML lifecycle and productionization processes
  • Upper-Intermediate English level or higher
  • Experience with feature stores and feature consistency management
  • Experience with large-scale batch scoring systems operating under freshness SLAs
  • Familiarity with experiment tracking platforms and evaluation gates
  • Experience with on-premises Kubernetes or bare-metal Linux infrastructure
  • Knowledge of DVC, lakeFS, or other data versioning tools
  • Experience with Bigtable, Redis, Aerospike, or similar low-latency serving databases
  • GPU scheduling and training cost optimization experience
  • Familiarity with SOC 2, ISO 27001, or GDPR-related compliance requirements
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
1,516 Employees

What We Do

Sigma Software Group, an award-winning and trusted IT partner, has been serving customers for over 21 years, providing comprehensive IT solutions to various businesses, ranging from startups to established software product houses. As one of Europe's substantial IT consultancies, it brings together a dedicated workforce of over 2,100 professionals in 40 offices across 19 countries. With a diverse client base, including more than 300 enterprises, including Fortune 500 stalwarts, Sigma Software Group is a preferred choice for developing solutions that help businesses create cutting-edge products while meeting their unique needs. Sigma Software Group operates as a dynamic ecosystem of tech companies, offering 25 ready-to-implement innovative products and 40+ value-added services. Furthermore, Sigma Software Group is committed to fostering innovation through initiatives such as the Sigma Software Labs business incubator, Sigma Software University, the SID Venture Partners VC Fund, UA Tech Network, Techosystem, the European Business Association, and other collaborative efforts. Since 2015, Sigma Software Group has consistently earned recognition on the IAOP's prestigious World's Top 100 Outsourcing list. The company's accomplishments have also been acknowledged by prominent global media outlets such as Forbes, CNBC, The Times, and Reuters

Similar Jobs

Air Space Intelligence Logo Air Space Intelligence

IT Support Analyst

Aerospace • Artificial Intelligence • Logistics • Machine Learning • Software • Transportation • Defense
Remote or Hybrid
Poland
150 Employees

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Poland
4000 Employees

Capco Logo Capco

Integration Engineer

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
Poland
6000 Employees

Pfizer Logo Pfizer

Research Associate

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote
Poland
121990 Employees
218K-218K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account