Platform Architect (AI/ML Infrastructure, GCP-focused)

Posted 2 Days Ago
Be an Early Applicant
5 Locations
In-Office or Remote
Senior level
HR Tech • Information Technology • Professional Services • Software
The Role
Own the architecture, deployment, reliability, and cost optimization of production AI/ML infrastructure on GCP. Build model serving systems, reproducible ML pipelines, GitOps workflows, and multi-tenant Kubernetes platforms. Manage LLM and agentic workloads, Terraform infrastructure, observability, SLOs, drift monitoring, accelerator capacity, and inference costs. Define operational standards, improve automation, and guide platform evolution for scalable enterprise-grade ML systems.
Summary Generated by Built In

We're looking for a Platform Architect who can set the standard for how we build, ship, and operate ML and AI systems at scale. You sit at the intersection of ML infrastructure and SRE. You'll own the path from model and pipeline to reliable production service, and you'll bring DevOps rigor to systems that are historically under-engineered. The immediate focus is AI/ML infrastructure on Google Cloud.

This is not a ticket-processing role, and it's not a research role. You'll tackle hard problems: model serving reliability, inference cost and latency, reproducible pipelines, and agentic workload operations. You'll have the scope to solve them properly. Senior professionals here identify problems before they're asked and raise the ceiling on what the platform can do.

WHAT YOU'LL WORK ON

  • Build and operate model and inference serving infrastructure, managing latency, throughput, autoscaling, and reliability for real-time and batch inference across multiple tenants.

  • Own the ML deployment lifecycle: model registry, versioning, promotion workflows, rollout strategies (canary, shadow, A/B), and safe rollback.

  • Operate agentic and LLM workloads in production, managing inference providers and gateways, quota and throttling behavior (TPS/TUPS limits), guardrails, prompt/version management, and graceful degradation under load.

  • Build reproducible, automated ML pipelines: training, evaluation, and deployment pipelines as code, with lineage and reproducibility built in.

  • Extend infrastructure-as-code to ML systems, using Terraform patterns and multi-project design that bring ML infrastructure under the same standards as the rest of the platform.

  • Operate GitOps for ML workloads, owning ArgoCD configuration and promotion workflows across environments and tenants.

  • Run ML and AI workloads on multi-tenant Kubernetes (GKE), managing GPU/accelerator scheduling, workload placement, tenant isolation, and cost-aware capacity.

  • Own ML reliability and observability: SLOs for inference services, model and data drift detection, performance regression monitoring, alert quality, on-call ergonomics, and runbook culture.

  • Drive ML cost efficiency by right-sizing accelerators, managing committed-use and Spot VM capacity, and attributing inference cost across tenants and workloads.

  • Use agentic coding tools for infrastructure and pipeline work: scaffolding environments, generating and reviewing IaC and pipeline code, and accelerating automation.

WHAT YOU WON'T FIND HERE

A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.

MUST HAVE

  • 5+ years in platform engineering, SRE, MLOps, or infrastructure, with meaningful time operating production systems at scale.

  • Hands-on experience deploying and operating ML or AI workloads in production: serving, inference, or training infrastructure that real users depended on.

  • Strong SRE/DevOps foundation. You've owned reliability for production services, defined and measured SLOs, run post-mortems, and driven measurable improvements.

  • Deep Terraform expertise. You actively manage complex Terraform state, reusable modules, and multi-project configurations in production, with CI-driven plan/apply workflows.

  • Strong GitOps background (ArgoCD or Flux in production). You understand declarative infrastructure management at depth and have opinions on how to do it well.

  • Deep Kubernetes knowledge. You've operated clusters in production, dealt with real failure modes, and understand the system at the control plane level. Production GKE experience (Standard and/or Autopilot) is strongly preferred.

  • Strong GCP background: VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project/organization design.

  • Hands-on experience with GCP data services, especially BigQuery in production: partitioning and clustering, query cost and performance tuning, and dataset-level IAM. Familiarity with at least one of Dataflow, Pub/Sub, or Dataproc.

  • Hands-on experience building and operating CI/CD pipelines (GitHub Actions, Cloud Build, GitLab CI, or equivalent), plus an understanding of how ML pipelines differ from standard application CI/CD.

  • Automation-first thinking at a senior level. You implement systems that eliminate entire categories of manual work.

  • Active user of agentic coding tools. You know how to direct them effectively, review their output critically, and use them to multiply your output.

  • Strong communicator. You can articulate operational decisions, model performance trade-offs, and incident summaries clearly to engineers and leadership alike.

NICE TO HAVE

  • Experience with GPU/accelerator scheduling and node lifecycle management in production (e.g., GKE node auto-provisioning, GPU time-sharing, or equivalent).

  • Experience operating LLM inference at scale, managing provider quotas/throttling (TPS/TUPS), gateways, caching, and guardrails (e.g., Vertex AI, Gemini API, or equivalent).

  • Experience with ML pipeline and orchestration tooling such as Argo Workflows, Kubeflow, Cloud Composer/Airflow, Vertex AI Pipelines, or equivalent.

  • Experience with model registries, feature stores, and experiment tracking (e.g., MLflow, Feast, or equivalent).

  • Familiarity with model and data drift monitoring and ML-specific observability.

  • Background in FinOps: inference cost attribution, committed use discount (CUD) and reservation planning, and accelerator capacity forecasting.

  • Familiarity with data infrastructure such as object storage, CDC pipelines, or lakehouse patterns.

  • Experience with multi-tenant infrastructure: isolation patterns, noisy neighbor mitigation, and tenant lifecycle management.

  • Prior experience scaling ML or platform infrastructure at a startup moving toward enterprise-grade requirements.

Location: Remote in LATAM

Payment in USD

Working hours: EST time zone

Skills Required

  • 5+ years of experience in platform engineering, SRE, MLOps, or infrastructure, including operating production systems at scale
  • Hands-on experience deploying and operating production ML or AI workloads, including serving, inference, or training infrastructure
  • Strong SRE and DevOps foundation, including production reliability ownership, SLOs, post-mortems, and measurable improvements
  • Deep production Terraform expertise, including complex state, reusable modules, multi-project configurations, and CI-driven workflows
  • Strong production GitOps experience with ArgoCD, Flux, or equivalent
  • Deep Kubernetes production experience, including cluster failure modes and control-plane knowledge
  • Strong GCP experience with VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project or organization design
  • Production experience with GCP data services, especially BigQuery, including partitioning, clustering, query tuning, and dataset-level IAM
  • Familiarity with at least one of Dataflow, Pub/Sub, or Dataproc
  • Hands-on experience building and operating CI/CD pipelines using GitHub Actions, Cloud Build, GitLab CI, or equivalent
  • Understanding of how ML pipelines differ from standard application CI/CD
  • Senior-level automation-first mindset focused on eliminating manual work
  • Active experience using and critically reviewing agentic coding tools
  • Strong written and verbal communication skills for operational decisions, model trade-offs, and incident summaries
  • Production experience with GPU or accelerator scheduling and node lifecycle management
  • Experience operating LLM inference at scale, including quotas, throttling, gateways, caching, and guardrails
  • Experience with ML orchestration tools such as Argo Workflows, Kubeflow, Airflow, Cloud Composer, or Vertex AI Pipelines
  • Experience with model registries, feature stores, or experiment tracking tools such as MLflow or Feast
  • Familiarity with model and data drift monitoring and ML-specific observability
  • FinOps experience with inference cost attribution, committed-use discounts, reservations, and accelerator forecasting
  • Familiarity with object storage, CDC pipelines, or lakehouse data infrastructure
  • Experience with multi-tenant infrastructure, isolation, noisy-neighbor mitigation, and tenant lifecycle management
  • Experience scaling ML or platform infrastructure at a startup transitioning to enterprise requirements
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
30 Employees
Year Founded: 2021

What We Do

Wizdaa is an IT recruitment services company that specializes in sourcing and placing top-tier remote developers from Latin America with startups, primarily in U.S. time zones. From its headquarters in Miami, it sources and vets elite-level developers to ensure seamless, real-time collaboration for clients. The company provides end-to-end solutions including onboarding, payroll, and tax management, helping startups build world-class development teams efficiently.

Similar Jobs

Samsara Logo Samsara

Supervisor, Sales and Contract Administration

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
México
4000 Employees

Mastercard Logo Mastercard

Senior Analyst, Deal Management

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Mexico City, Ciudad De México, MEX
38800 Employees

Pfizer Logo Pfizer

Medical Science Liaison Oncology

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote
México
121990 Employees

Pfizer Logo Pfizer

Medical Manager - Genitourinary Oncology

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote
México
121990 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account