Senior DevOps / Platform Engineer

Posted Yesterday
Be an Early Applicant
Hiring Remotely in México
Remote
Senior level
Computer Vision • Software
The Role
Own the reliability, security, and operability of a multi-tenant AWS AI platform. Manage infrastructure as code, cloud services, CI/CD pipelines, observability, secrets, deployments, tenant metering, SLOs, and compliance controls. Operate systems including Langfuse and stateful analytics or database platforms, support development teams with environment and performance issues, and enforce security, data governance, and operational standards across production workloads.
Summary Generated by Built In
About the role

Helpware is building a multi-tenant B2B AI platform for customer support operations — a product portfolio spanning Agent Assist, AI-powered QA, document processing, support automation, workforce management, and others. We're hiring a Sr. PlatformOps Engineer to own the reliability, security, and operability of this infrastructure, and to carry classic DevOps responsibilities: CI/CD, environments, and developer enablement.

This is a platform-ownership role. Your job is to help the development team make architectural designs operationally real, keep them healthy as tenant count grows, and enforce them structurally (in pipelines and shared tooling).

What you'll do

Platform infrastructure (AWS)

  • Contribute with cloud architectural designs, focusing on Well-Architected Framework pillars: Security, Operational Excellence, Performance Efficiency, Reliability and Cost Optimization.
  • Own the AWS footprint: ECS Fargate services, Lambda and Batch workloads, RDS PostgreSQL, ElastiCache Redis, S3, CloudWatch, networking/IAM, Bedrock, etc.
  • Manage everything as code (Terraform preferred), with reviewable, repeatable environment provisioning for dev/test/staging/prod.
  • Operate our Observability and self-hosted Langfuse stack: deployment, upgrades, backup/restore, and access control (SSO-backed, role-scoped, itself audit-logged).

DevOps & developer enablement

  • Build and maintain CI/CD pipelines, including the platform's conformance gates: contract tests for the logging standard and the logging conformance linter.
  • Own secrets management, artifact/image pipelines, and deployment strategies (blue/green or rolling) across services.

Monitoring and Support

  • Own our OpenTelemetry pipeline (SDK → Collector → CloudWatch for app/infra; Langfuse for LLM traces), keeping the three-layer separation between infra observability, LLM observability, and the business audit log intact.
  • Maintain per-tenant usage and cost infrastructure, platform-level dashboards, alerting, and SLOs.
  • Keep trace retention and masking aligned with the data governance standard (masked content only in LLM traces, defined retention windows per layer).
  • Support product teams day-to-day: environment issues, deployment questions, performance debugging.

Security & compliance

  • Maintain the security posture underpinning SOC 2 / HIPAA / GDPR commitments: least-privilege IAM, network isolation, encryption at rest and in transit, audit trails.
  • Ensure prompt/completion data, support the PII masking architecture at the infrastructure level.
What we're looking for
  • 5+ years in platform, SRE, or DevOps roles running production workloads on AWS (and other Cloud Platforms, preferred)
  • Deep hands-on experience with ECS (or EKS), RDS, ElastiCache, S3, CloudWatch, IAM, and VPC networking.
  • Strong infrastructure-as-code practice (Terraform or CDK) and CI/CD ownership (GitHub Actions, GitLab CI, or similar).
  • Experience with Claude Code and AI Spec Driven Design
  • Experience operating stateful open-source systems in production — running ClickHouse, or comparable columnar/analytics stores (Druid, Pinot) or demonstrably transferable database operations depth (Postgres at scale, Elasticsearch, Kafka).
  • Working knowledge of OpenTelemetry: collectors, exporters, sampling, and instrumentation conventions.
  • A security-first habit: you treat IAM policies, secrets, and data boundaries as design work, not afterthoughts.
Nice to have
  • Experience with multiple Cloud providers (AWS, Azure, GCP).
  • Experience self-hosting LLM/GenAI tooling (Langfuse, LiteLLM, vLLM) or operating GenAI workloads (Bedrock, token cost management).
  • Scripting/automation fluency (Python or Go preferred; the platform's services are Python/FastAPI).
  • Prior work in a compliance-driven environment (SOC 2 audits, HIPAA, GDPR data residency).
  • Experience with multi-tenant SaaS operations: tenant isolation, per-tenant metering, noisy-neighbor management.
  • FinOps experience — AWS cost allocation, tagging strategies, per-tenant cost attribution.

Skills Required

  • 5+ years of experience in platform, SRE, or DevOps roles running production workloads on AWS
  • Deep hands-on experience with ECS or EKS, RDS, ElastiCache, S3, CloudWatch, IAM, and VPC networking
  • Strong infrastructure-as-code experience with Terraform or CDK
  • Experience owning CI/CD pipelines using GitHub Actions, GitLab CI, or similar
  • Experience with Claude Code and AI Spec Driven Design
  • Experience operating stateful open-source systems in production, such as ClickHouse, Druid, Pinot, Postgres at scale, Elasticsearch, or Kafka
  • Working knowledge of OpenTelemetry collectors, exporters, sampling, and instrumentation conventions
  • Security-first experience with IAM policies, secrets, encryption, network isolation, and data boundaries
  • Experience with multiple cloud providers, including AWS, Azure, or GCP
  • Experience self-hosting LLM or GenAI tooling such as Langfuse, LiteLLM, or vLLM, or operating GenAI workloads with Bedrock
  • Scripting and automation fluency in Python or Go
  • Experience with Python or FastAPI platform services
  • Prior experience in compliance-driven environments involving SOC 2, HIPAA, or GDPR data residency
  • Experience with multi-tenant SaaS operations, tenant isolation, metering, and noisy-neighbor management
  • FinOps experience with AWS cost allocation, tagging strategies, and per-tenant cost attribution
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Lexington, KY
1,061 Employees
Year Founded: 2015

What We Do

Founded in 2015, Helpware is a company taking a modern approach to the outsourcing industry. We created the company to change perceptions of what outsourcing is and can be, and we did that by building amazing cultures in each of our locations, and by simply treating our employees better. With Helpware, we are all a team and family, and you'll see that true difference when partnering with us. Helpware builds customized teams in Customer Service and Back Office for industry-leading startups and modern companies. With offices in California, Colorado, Kentucky, Ukraine, Philippines, Germany, and Mexico, we have the global scale to tailor custom teams and processes for success to our many powerhouse clients. Helpware has grown over the years, initially catering to startup client partners, and has now evolved into creating client partnerships with large enterprises as well.

Similar Jobs

Mondelēz International Logo Mondelēz International

Controller

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
Toluca, México, MEX
90000 Employees

Crunchyroll Logo Crunchyroll

Senior Specialist, Distribution Operations

Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Remote or Hybrid
Mexico City, Ciudad De México, MEX
1300 Employees

Mastercard Logo Mastercard

VP - Services Business Development

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Mexico City, Ciudad De México, MEX
38800 Employees

Mastercard Logo Mastercard

Specialist, Implementation

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Mexico City, Ciudad De México, MEX
38800 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account