Infrastructure Leader

Posted Yesterday
Be an Early Applicant
2 Locations
Remote or Hybrid
Senior level
Artificial Intelligence • Productivity • Software
The Role
Leads the architecture and operations of Kubernetes and Azure cloud infrastructure for an agentic platform. Defines environment topology, autoscaling, networking, workload isolation, infrastructure-as-code, GitOps, identity, secrets, observability, security, disaster recovery, and cost controls. Supports reliable LLM-serving workloads, including GPU/CPU scheduling and endpoint reliability, while establishing capacity standards, operational dashboards, incident-response practices, and recovery runbooks.
Summary Generated by Built In
Job Summary

CONVO is seeking a Senior Infrastructure Leader to own the Kubernetes and cloud foundation for its agentic platform. You’ll lead the architecture and operations of scalable, secure, and resilient cloud infrastructure across environments, with a strong focus on Kubernetes, Azure, Terraform, GitOps, observability, identity, workload isolation, and cost optimization. You’ll also support reliable LLM-serving workloads and establish infrastructure standards, capacity controls, and operational practices for the platform.

Technical mission

Own the Kubernetes and cloud foundation for the agentic platform, including scalability, isolation, observability, identity integration and cost-per-request control.

Key responsibilities
  • Define the target Kubernetes and cloud topology for development, test and production environments.
  • Own autoscaling, workload placement, network boundaries and runtime isolation for agent, tool and LLM-serving workloads.
  • Set infrastructure-as-code, GitOps, identity, secrets, gateway and certificate-management standards.
  • Define service-level indicators, operational dashboards, alerting and incident-response expectations for the platform.
  • Establish capacity and FinOps controls, including cost-per-request budgets and resource-consumption attribution.
  • Review platform changes for security, resilience, recoverability and vendor decoupling.
Required technical capabilities
  • Minimum experience: 8+ years in cloud/platform infrastructure, including 3+ years leading production Kubernetes architecture or operations.
  • Production architecture and operations experience with Kubernetes, including autoscaling, networking, storage and workload isolation.
  • Strong Terraform or equivalent infrastructure-as-code capability and practical GitOps delivery experience.
  • Hands-on cloud platform experience, preferably Azure, covering compute, networking, managed identity and secure service exposure.
  • Observability engineering across metrics, logs and traces, with actionable service and cost dashboards.
  • Experience integrating identity providers, API gateways, secrets management and policy controls.
  • Understanding of LLM inference or model-serving workloads, including GPU/CPU scheduling and endpoint reliability.
Preferred experience
  • AKS, Azure Monitor/Application Insights, managed identity and private networking.
  • Policy-as-code, service mesh, multi-cluster operations and disaster-recovery design.
  • FinOps practices for shared AI platforms and usage-based cost allocation.
Expected deliverables / acceptance evidence
  • Approved cloud/Kubernetes reference architecture and environment topology.
  • Version-controlled IaC and GitOps baseline with security and rollback controls.
  • Capacity, availability and cost-per-request dashboards with alert thresholds.
  • Operational runbooks for deployment, failure recovery and platform incidents.
Primary interfaces
  • Works with the Agent Runtime team, DevOps, Architecture Office, security/identity owners and the data-platform infrastructure lead.
Why Join CONVO?

At CONVO, we’re reimagining how sales organizations in FMCG run smarter, faster, and more profitably. You’ll be joining a team of thinkers, builders, and consultants who thrive at the intersection of business and technology.

We offer:
  • Global exposure working with Tier-1 FMCG clients
  • A high-impact role where your input shapes real-world commercial outcomes
  • A collaborative culture with mentorship, continuous learning, and career growth opportunities

Skills Required

  • 8+ years of experience in cloud or platform infrastructure
  • 3+ years leading production Kubernetes architecture or operations
  • Production Kubernetes architecture and operations experience, including autoscaling, networking, storage, and workload isolation
  • Strong Terraform or equivalent infrastructure-as-code capability
  • Practical GitOps delivery experience
  • Hands-on cloud platform experience, preferably Azure, including compute, networking, managed identity, and secure service exposure
  • Observability engineering across metrics, logs, and traces
  • Experience integrating identity providers, API gateways, secrets management, and policy controls
  • Understanding of LLM inference or model-serving workloads, including GPU/CPU scheduling and endpoint reliability
  • Experience with AKS, Azure Monitor or Application Insights, managed identity, and private networking
  • Experience with policy-as-code, service mesh, multi-cluster operations, and disaster-recovery design
  • FinOps experience for shared AI platforms and usage-based cost allocation
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Los Altos, CA
138 Employees
Year Founded: 2012

What We Do

Convo provides an enterprise social collaboration platform that enables easy, secure conversations between desk / non-desk workers to accelerate company productivity and engagement. Unlike existing email-focused or chat-centric collaboration platforms, only Convo combines the ease of social networks with rich collaboration capabilities to simplify and optimize work interactions for all employees -- no email required.

Gallery

Gallery

Similar Jobs

GitLab Logo GitLab

Back-end Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
India
2500 Employees

GitLab Logo GitLab

Back-end Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
India
2500 Employees

Rubrik Logo Rubrik

Manager, Resilience Gaurdiance

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Remote
India
3000 Employees

Pfizer Logo Pfizer

Manager, Medical Communications and Content Solutions, Vaccines

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote
India
121990 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account