Senior Cloud Infrastructure Engineer

Posted 2 Days Ago
Be an Early Applicant
San Jose, CA, USA
In-Office
Senior level
Artificial Intelligence • Robotics • Automation • Manufacturing
The Role
Design, implement, and operate multi-cloud cloud foundations across GCP and AWS: account/project topology, networking, IaC modules, state strategy, Kubernetes cluster foundations, IAM/workload identity, cost visibility, and controls to support SOC 2. Drive reusable Terraform/OpenTofu patterns, networking, and infrastructure standards for growth and reliability.
Summary Generated by Built In

About Trener

Trener is building the software foundation for a new generation of industrial robotics.

We combine advanced AI, intuitive programming, and pre-trained skill models to make robots easier to deploy,

operate, and adapt to real-world industrial work. With teams in San Jose and Trondheim, we build software

that connects cloud infrastructure, developer platforms, and deployed robotic systems.

About the Role

We run a production cloud platform on Google Cloud today and are expanding our AWS footprint. We need a senior infrastructure engineer to help establish a deliberate multi-cloud foundation and define how GCP and AWS should coexist as the platform grows.

You will own the Cloud Foundations layer: cloud account and project structure, infrastructure-as-code, state strategy, networking, Kubernetes foundations, and implementation of cloud access controls. You will help decide where Trener should standardize across providers and where provider-specific architecture is the right boundary.

This is a senior individual contributor role. You will not be managing people, and you will not be inheriting a greenfield.

What You Will Own

  • Cloud foundations across GCP and AWS - project and account topology, landing zones, organization policies, baseline guardrails, and enrollment of new environments.

  • Cloud and cross-cloud networking - VPCs, routing, peering, DNS, load balancing, firewall policy, IP addressing, and private connectivity between cloud environments.

  • Infrastructure as code - reusable Terraform/OpenTofu modules, environment composition, lifecycle management, and patterns that keep infrastructure understandable as the estate grows.

  • State strategy and blast-radius boundaries - how infrastructure state is partitioned, how dependencies are expressed, and how those patterns evolve across providers.

  • Cloud identity implementation - implement cloud IAM, workload identity, federation, and access controls in partnership with Security.

  • Kubernetes foundations - cluster lifecycle, baseline configuration, upgrades, and shared infrastructure services beneath application workloads.

  • Infrastructure cost visibility - implement the tagging, labeling, budgets, alerts, and reporting needed to make cloud spend understandable and surface obvious infrastructure waste.

  • Infrastructure controls supporting our SOC 2 program - resource labeling, access boundaries, configuration standards, and evidence-producing infrastructure practices.

Requirements

  • Proven experience operating production infrastructure as a cloud, infrastructure, platform, or site reliability engineer.

  • Terraform or OpenTofu at module-authoring depth - writing and versioning reusable modules, managing state across environments, and handling the lifecycle of real infrastructure over time.

  • Cloud foundation design experience - multi-account, multi-project, landing-zone, guardrail, IAM, or network architecture beyond a handful of isolated workloads.

  • Experience with infrastructure composition or orchestration - Atmos, Terragrunt, Terraspace, or an equivalent approach to managing reusable infrastructure across environments.

  • Strong networking depth - VPCs, routing, peering, DNS, TLS, load balancing, firewall policy, IP addressing, and connectivity troubleshooting.

  • Kubernetes working knowledge - cluster operations, RBAC, networking, shared services, and troubleshooting workloads that will not start or communicate correctly.

  • Helm chart authoring, not just chart installation.

  • Production cloud experience across at least two major providers, or deep experience with one plus substantial hands- on involvement extending an organization into another.

  • Python and Bash for automation, CLIs, diagnostics, and glue.

  • Linux fluency.

  • Experience with controlled infrastructure environments - SOC 2, HIPAA, HITRUST, FedRAMP, PCI, or similar security/compliance expectations.

  • Fluency in English and strong written communication. Architectural decisions here are expected to be documented and reviewed in writing

Experience That Will Help You Succeed

  • GitOps workflows across multiple environments and comfort operating infrastructure through reviewed, auditable changes.

  • Externalized secrets management and Kubernetes-native secret delivery patterns.

  • Operating Prometheus/Grafana/Loki or comparable observability tooling as an infrastructure consumer and operator.

  • Experience introducing or formalizing a second cloud, including a clear account of what you would repeat and what you would change.

  • Cloud identity federation and workload identity across providers.

  • Disaster-recovery, regional resilience, or high-availability infrastructure work.

  • Private cloud networking or overlay connectivity.

What We Offer

The opportunity to work at the intersection of robotics, AI, and cloud infrastructure. A senior technical role with meaningful influence over the foundations on which the company builds.

A growing international engineering organization with teams in the United States and Norway

 

Skills Required

  • Proven experience operating production infrastructure as a cloud, infrastructure, platform, or site reliability engineer.
  • Terraform or OpenTofu at module-authoring depth, including state management and lifecycle.
  • Cloud foundation design experience (multi-account, landing zones, guardrails, IAM, network architecture).
  • Experience with infrastructure composition/orchestration (Atmos, Terragrunt, Terraspace, or equivalent).
  • Strong networking depth (VPCs, routing, peering, DNS, TLS, load balancing, firewall policy, IP addressing, connectivity troubleshooting).
  • Kubernetes working knowledge: cluster operations, RBAC, networking, shared services, and troubleshooting.
  • Helm chart authoring experience.
  • Production cloud experience across at least two major providers, or deep experience with one plus hands-on multi-cloud expansion.
  • Python and Bash for automation, CLIs, diagnostics, and glue.
  • Linux fluency.
  • Experience with controlled infrastructure environments (SOC 2, HIPAA, HITRUST, FedRAMP, PCI, or similar).
  • Fluency in English and strong written communication for architectural documentation and reviews.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
45 Employees
Year Founded: 2024

What We Do

Trener Robotics is a Physical AI company that builds the intelligence layer for industrial robots. Its platform, Acteris, leverages artificial intelligence to enable natural language programming, allowing industrial robots to operate autonomously, adapt to variability, and perform complex manufacturing tasks. The company aims to transform traditional, scripted machines into intelligent, self-learning systems to accelerate the deployment of advanced industrial automation.

Similar Jobs

CrowdStrike Logo CrowdStrike

Development Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
11000 Employees
125K-180K Annually

Oscar Health Logo Oscar Health

Senior Software Engineer

Healthtech • Insurance
In-Office
San Francisco, CA, USA
2200 Employees
181K-237K Annually

Oscar Health Logo Oscar Health

Senior Software Engineer

Healthtech • Insurance
In-Office
Los Angeles, CA, USA
2200 Employees
181K-237K Annually

NVIDIA Logo NVIDIA

Senior Devops Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office
Santa Clara, CA, USA
21960 Employees
184K-288K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account