Senior MLOps Engineer

Posted 7 Hours Ago
Be an Early Applicant
Hiring Remotely in Latvia
Remote
Senior level
Software • Cybersecurity
The Role
Architects and operates production ML infrastructure on GCP, including model deployment, serving, training pipelines, CI/CD, observability, drift detection, and scalable inference systems. Collaborates with AI researchers, data engineers, and backend teams to transition prototypes into secure, resilient microservices. Supports feature engineering, dataset versioning, retraining automation, and optimized AI workloads.
Summary Generated by Built In

Point Wild helps customers monitor, manage, and protect against the risks associated with their identities and personal information in a digital world. Backed by WndrCo, Warburg Pincus and General Catalyst, Point Wild is dedicated to creating the world’s most comprehensive portfolio of industry-leading cybersecurity solutions. Our vision is to become THE go-to resource for every cyber protection need individuals may face - today and in the future. 

Join us for the ride!

About the Role:

As a Senior MLOps Engineer, you will play a critical role in architecting, building, and maintaining the infrastructure, pipelines, and tooling that enable complex AI models to be deployed, scaled, and monitored in production on Google Cloud Platform (GCP). You’ll collaborate closely with AI Researchers, Data Engineers, and Backend teams to bridge the gap between experimentation and high-performance, enterprise-grade production systems.

Your Day to Day:

  • GCP ML Infrastructure: Architect and manage scalable GCP-based ML infrastructure using Vertex AI, Google Kubernetes Engine (GKE), Google Cloud Storage (GCS), Cloud Run, and GPU/TPU compute instances.
  • Model Deployment & Serving: Own the end-to-end deployment lifecycle for machine learning models. Build high-throughput, low-latency inference services using containerization and specialized serving frameworks (e.g., Triton Inference Server, vLLM, MLflow).
  • Continuous Integration & Training (CI/CD/CT): Build automated, reproducible pipelines for model training, testing, evaluation, and deployment using tools like Airflow, Vertex AI Pipelines, and GitHub Actions.
  • Production Observability & Monitoring: Implement robust monitoring systems for both system health (latency, throughput, uptime) and ML-specific metrics (feature drift, prediction accuracy, and data distribution shifts) to enable automated retraining triggers.
  • Supporting AI & Research Engineers: Provide scalable training environments, optimized runtime infrastructure, and standardized deployment templates that allow AI engineers to move fast without compromising reliability.
  • Data & Feature Engineering Support: Collaborate with Data Engineers to integrate model pipelines with feature stores, dataset versioning, and stream/batch data processing workflows.
  • Scaling PoCs to Production: Lead the technical transition of raw AI prototypes and notebooks into resilient, secure, and auto-scaling microservices.

What You Bring to the Table:

  • Senior MLOps Experience: At least 5 years of hands-on experience designing, deploying, and maintaining production ML workloads in cloud environments.
  • GCP Ecosystem Mastery: Deep, practical experience with Google Cloud Platform (GCP), including Vertex AI, Cloud Storage, GKE, Cloud Run, and IAM/VPC configurations.
  • Model Serving & Tooling: Expertise with containerization (Docker, Kubernetes/GKE) and specialized serving tools (Triton, vLLM, MLflow).
  • Orchestration & CI/CD: Proven track record with workflow orchestrators (Airflow, Vertex AI Pipelines) and modern CI/CD tools (GitHub Actions, ArgoCD).
  • Infrastructure as Code (IaC): Solid experience managing cloud resources using Terraform.
  • Software Development Skills: Proficiency in Python and SQL for scripting, automation, API development, and data manipulation.
  • ML Observability: Hands-on experience with logging, telemetry, and drift detection tools (Grafana, Prometheus, GCP Cloud Monitoring, or specialized ML observability frameworks).

Nice to Have:

  • Experience running large-scale LLM or Deep Learning inference/training workloads.
  • GCP Professional Machine Learning Engineer or GCP Professional Cloud Architect certifications.
  • Familiarity with feature stores (e.g., Feast, Vertex AI Feature Store).

Why This Role Matters:

  • Operationalizing AI: AI models only create value when they operate reliably at scale. You are the architect making production AI possible.
  • Infrastructure Backbone: You provide the core practices, automation, and tooling that empower AI teams to innovate rapidly while maintaining system stability.
  • Cross-Functional Bridge: You connect the worlds of data science, cloud operations, and software engineering to maintain production reliability.

As part of Point Wild, you will:

Solve real customer problems. Point Wild’s point solutions allow consumers to address their immediate cyber protection needs. Our mandate is to continuously anticipate our customers’ evolving digital security needs to create best-in-class solutions aimed at keeping them safe.

See your impact. We are a scrappy, nimble organization where individual contributions are needed and valued. You will see your impact every day.

Accelerate your career.  As we expand, you will have the opportunity to learn new technologies, products, and markets in a fast-paced, growth-oriented environment.

Most importantly, you’ll get to work with other talented people at a company where people matter. If you want to put your fingerprint on an organization and leapfrog your growth, this is the place for you.

In keeping with our beliefs and goals, no employee or applicant will face discrimination or harassment based on race, color, ancestry, national origin, religion, age, gender, marital domestic partner status, sexual orientation, gender identity, disability status, or veteran status. Above and beyond discrimination or harassment based on “protected categories,” Point Wild is committed to being an inclusive community where all feel welcome. Whether blatant or hidden, barriers to success have no place at Point Wild.

Important privacy information for United States based job applicants can be found here.


Skills Required

  • At least 5 years of experience designing, deploying, and maintaining production machine learning workloads in cloud environments.
  • Practical expertise with Google Cloud Platform, including Vertex AI, Cloud Storage, GKE, Cloud Run, IAM, and VPC configurations.
  • Experience with Docker, Kubernetes or GKE, Triton Inference Server, vLLM, and MLflow.
  • Experience with Airflow, Vertex AI Pipelines, GitHub Actions, and modern CI/CD workflows.
  • Experience managing cloud infrastructure with Terraform.
  • Proficiency in Python and SQL.
  • Experience with logging, telemetry, ML monitoring, and drift detection using tools such as Grafana, Prometheus, or GCP Cloud Monitoring.
  • Experience with large-scale LLM or deep learning inference or training workloads.
  • GCP Professional Machine Learning Engineer or Professional Cloud Architect certification.
  • Familiarity with feature stores such as Feast or Vertex AI Feature Store.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boston, Massachusetts
105 Employees

What We Do

Point Wild is a global leader in online protection, operating a portfolio of best-in-class device security, online privacy and identity theft protection brands. Serving consumers, partners and enterprises, Point Wild provides protection to more than 25 million users worldwide. To learn more, visit www.pointwild.com.

Similar Jobs

Remote
25 Locations
393 Employees
179K-179K Annually

MWDN Logo MWDN

AI Video Creator

Information Technology • Consulting
Remote
Latvia
143 Employees
In-Office or Remote
5 Locations
257 Employees

Ruby Labs Logo Ruby Labs

Growth Marketing Lead (B2B SaaS & Payments)

Information Technology • Software
In-Office or Remote
26 Locations
28 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account