Senior Platform Engineer / SRE - Remote (LATAM)

Posted 28 Days Ago
Hiring Remotely in Brazil
Remote
Senior level
Artificial Intelligence • HR Tech • Software
The Role
Own the architecture, deployment, reliability, observability, and cost efficiency of production AI/ML infrastructure on GCP. Build model serving systems, reproducible ML pipelines, GitOps promotion workflows, and multi-tenant Kubernetes platforms. Operate LLM and agentic workloads, manage GPUs and autoscaling, define SLOs, monitor drift, and improve inference performance. Apply Terraform, CI/CD, and SRE practices to automate infrastructure and support reliable real-time and batch inference at enterprise scale.
Summary Generated by Built In

Senior Platform Engineer / SRE

Level: Senior
Location: Remote in LATAM Only (Brazil preferred)
Working hours: EST
Type: Full-Time (Payment in USD)
Reports to: Engineering Manager, Platform


Company Overview

ReadyOn.AI is an AI native Labor Operating System redefining how the world’s largest enterprises manage frontline labor. Born out of a Stanford AI Lab, ReadyOn applies advanced AI and market design principles to one of the world’s hardest optimization problems: matching 2.7 billion frontline workers to the right shifts, in real time.

Frontline workers increasingly expect the flexibility and autonomy of gig platforms, while large employers face relentless pressure to control labor costs. ReadyOn bridges that gap with a system of action that predicts workforce demand, dynamically matches it to available employees, and automates thousands of staffing decisions across complex, multi site operations.

The platform is already proven at global scale, powering labor operations for some of the world’s largest enterprises, including F250 and F500 organizations spanning hundreds of thousands of employees and billions of dollars in annual labor spend. ReadyOn has demonstrated that scheduling was never the real problem. It was a symptom. The real challenge is dynamically matching people and work at scale.

Headquartered in San Francisco with over 100 employees, ReadyOn grew revenue 8x year over year in 2025, driven by multiple seven figure Fortune 250 deployments and a rapidly expanding pipeline.

Transform How Frontline Work Runs

Frontline labor can represent up to 40% of a company’s P&L, yet the systems managing this multi trillion dollar market were built around static schedules and manual processes.

ReadyOn is rejecting that paradigm. Staffing is not a scheduling problem. It is a real time supply and demand orchestration problem. Our AI native Labor Operating System is built from the ground up for AI agents to optimize labor in real time, much like ridesharing platforms match drivers and riders, but applied to frontline labor instead of fixed, one-size-fits-all schedules.

The Role

We're looking for a Senior Platform Engineer / SRE who can set the standard for how we automate and operate systems at scale. You'll tackle hard problems — multi-tenant isolation, self-service infrastructure, reliability engineering — and have the scope to solve them properly.

This is not a ticket-processing role. Seniors here identify problems before they're asked, and raise the ceiling on what the platform can do.

What You'll Work On
  • Shape how we do infrastructure-as-code — Terraform patterns, multi-account design, and the standards that hold it together across teams.

  • Operate GitOps at scale — ArgoCD configuration, managing promotion workflows, and ensuring deployment reliability across multiple environments and tenants.

  • Operate multi-tenant Kubernetes infrastructure on AWS EKS — managing tenant isolation, workload placement, cluster topology, and maintaining scalability.

  • Maintain self-service infrastructure automation — managing provisioning pipelines and configuration management.

  • Use agentic coding tools for infrastructure work — scaffolding new environments, generating and reviewing IaC, and accelerating automation.

  • Own reliability — managing SLOs, monitoring error budgets, ensuring incident response quality, and driving the feedback loop that turns incidents into platform improvements.

  • Maintain observability standards — ensuring trace coverage, alert quality, on-call ergonomics, and runbook culture.

  • Maintain security posture — focusing on secrets management at scale and infrastructure hardening.

Must Have
  • 5+ years in platform engineering, SRE, or infrastructure — with meaningful time operating production systems at scale.

  • Deep IaC expertise — you actively manage complex Terraform state and multi-account configurations in production.

  • Strong GitOps background — you understand declarative infrastructure management at depth and have opinions on how to do it well.

  • Deep Kubernetes knowledge — you've operated clusters in production, dealt with real failure modes, and understand the system at the control plane level.

  • Strong AWS background — networking, compute, IAM, storage, multi-account design

  • Automation-first thinking at a senior level — you implement systems that eliminate entire categories of manual work.

  • Hands-on experience building and operating CI pipelines — GitHub Actions, CircleCI, GitLab CI, or equivalent.

  • Active user of agentic coding tools — you know how to direct them effectively, review their output critically, and use them to multiply your output.

  • Reliability engineering track record — SLOs defined and measured, post-mortems run, measurable improvements driven.

  • Strong communicator — you can articulate operational decisions and incident summaries clearly to engineers and leadership alike.

Nice to Have
  • Experience with Keycloak or other IdP

  • Experience with Argo: ArgoCD, Argo Workflows, Rollout

  • Experience with Karpenter and node lifecycle management in production

  • Background in FinOps — cost attribution, reserved capacity planning, workload right-sizing

  • Familiarity with data infrastructure — object storage, CDC pipelines, or lakehouse patterns

  • Experience with multi-tenant infrastructure — isolation patterns, noisy neighbor mitigation, and tenant lifecycle management.

  • Experience supporting AI/ML inference workloads or GPU-based compute in production

  • Prior experience scaling platform infrastructure at a startup moving toward enterprise-grade requirements

What You Won't Find Here

A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.

Skills Required

  • 5+ years of experience in platform engineering, SRE, MLOps, or infrastructure, including production systems at scale
  • Hands-on experience deploying and operating ML or AI workloads in production, including serving, inference, or training infrastructure
  • Strong SRE and DevOps foundation, including ownership of reliability, SLOs, post-mortems, and measurable improvements
  • Deep production Terraform expertise, including complex state, reusable modules, multi-project configurations, and CI-driven workflows
  • Strong production GitOps experience with ArgoCD, Flux, or equivalent
  • Deep Kubernetes production experience, including cluster failure modes and control-plane operations
  • Strong GCP experience with VPC networking, Compute Engine, IAM, Cloud Storage, and multi-project or organization design
  • Production experience with GCP data services, especially BigQuery; familiarity with Dataflow, Pub/Sub, or Dataproc
  • Hands-on experience building and operating CI/CD pipelines using GitHub Actions, Cloud Build, GitLab CI, or equivalent
  • Senior-level automation-first approach to eliminating manual work
  • Active use of agentic coding tools for infrastructure and pipeline development
  • Strong written and verbal communication skills
  • Production experience with GPU or accelerator scheduling and node lifecycle management
  • Experience operating LLM inference at scale, including quotas, throttling, gateways, caching, and guardrails
  • Experience with ML pipeline and orchestration tools such as Argo Workflows, Kubeflow, Airflow, or Vertex AI Pipelines
  • Experience with model registries, feature stores, or experiment tracking tools such as MLflow or Feast
  • Familiarity with model and data drift monitoring and ML-specific observability
  • FinOps experience with inference cost attribution, committed-use discounts, reservations, and accelerator forecasting
  • Familiarity with object storage, CDC pipelines, or lakehouse patterns
  • Experience with multi-tenant infrastructure, isolation, noisy-neighbor mitigation, and tenant lifecycle management
  • Experience scaling ML or platform infrastructure at an enterprise-focused startup
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
47 Employees
Year Founded: 2022

What We Do

ReadyOn is an AI-native labor operating system designed for hiring, scheduling, and managing flexible frontline workforces. The platform optimizes workforce deployment for enterprises by utilizing agentic AI to replace legacy fixed scheduling with a dynamic labor marketplace where employees choose their own shifts. This approach improves staffing reliability and significantly reduces hiring time for large-scale frontline operations.

Similar Jobs

NinjaOne Logo NinjaOne

Architect

Information Technology • Productivity • Software • Infrastructure as a Service (IaaS)
Remote or Hybrid
Brazil
2000 Employees

Circle (circle.so) Logo Circle (circle.so)

Senior Quality Engineer

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
2 Locations
250 Employees
120K-130K Annually

JumpCloud Logo JumpCloud

Consultant

Cloud • Information Technology • Security • Software
Easy Apply
In-Office or Remote
São Paulo, BRA
800 Employees

Mastercard Logo Mastercard

Director, Strategic Pricing

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
São Paulo, BRA
38800 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account