SRE (Infrastructure) Lead Engineer

Sorry, this job was removed at 10:34 a.m. (UTC) on Friday, Sep 04, 2026
Be an Early Applicant
Tokyo, JPN
Hybrid
Senior level
Artificial Intelligence • Retail • Robotics • Automation
The Role
Leads hands-on reliability and platform engineering for Azure and Kubernetes systems supporting retail robotics. Responsibilities include incident response, SLOs, observability, disaster recovery, infrastructure as code, GitOps, CI/CD, cloud networking, identity, security, cost optimization, and operational tooling. The role mentors a small Infra/SRE team, sets technical direction, improves on-call practices, and collaborates with engineering, robotics, security, product, and operations teams to build a resilient, automated platform.
Summary Generated by Built In
Role Overview

The SRE (Infrastructure) Lead Engineer owns the reliability, operability, and continuous improvement of the platform that powers our retail robotics products. This is a hands-on technical leadership role for someone who can operate production systems directly, improve infrastructure through code, and guide a small Infra/SRE team toward stronger reliability, security, automation, and cost discipline.

 

Our current platform is primarily Azure-based, but this role does not require Azure-only experience. We value strong hands-on cloud infrastructure experience on any major cloud platform, with the ability to learn the specifics of Azure, AKS, and our tooling quickly.

 

You will work closely with Backend, Frontend, Robotics, Security, Product, and Operations teams to keep our cloud and Kubernetes environments dependable for live store operations, robot and smart shelf workflows, telemetry processing, internal tooling, and customer-facing services.

Current Platform State

    We operate a production platform supporting:

     

    • Retail robotics SaaS services and internal operations tools.
    • Kubernetes-based application workloads across development, staging, infrastructure, and production environments.
    • Robot, smart shelf, and store operations systems that depend on reliable cloud-to-edge communication.
    • GitOps-style Kubernetes manifests and environment overlays.
    • Infrastructure as Code using Azure Bicep, Terraform, and Terragrunt.
    • CI/CD automation with GitHub Actions, including cloud authentication through OIDC.
    • Observability through Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, and alerting workflows.
    •  

      You will inherit systems that are already running and help evolve them into a more standardized, automated, and resilient platform.

Tech Stack

    • Cloud: Microsoft Azure today; AWS or GCP experience is also welcome if paired with strong cloud fundamentals
    • Compute & Runtime: AKS, Kubernetes, Docker, App Service, Functions
    • Infrastructure as Code: Azure Bicep, Terraform, Terragrunt
    • GitOps & Manifests: Argo CD, Kustomize, Helm, Kubernetes YAML
    • Networking & Ingress: VNet, subnets, NSG, load balancers, Traefik, ingress-nginx, cert-manager, TLS
    • Identity & Security: Microsoft Entra ID, Azure RBAC, workload identity, managed identities, Key Vault, GitHub Actions OIDC
    • Data & Storage: Azure Cosmos DB, PostgreSQL, Redis / Redis Enterprise, Blob Storage, Storage Accounts
    • Observability: Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, Prometheus-compatible metrics, scheduled query alerts
    • Languages & Scripting: Bash, Python, TypeScript, C# or similar production scripting/application languages
    • Development Tools: Git/GitHub, GitHub Actions, Azure CLI, kubectl, Helm, Argo CD CLI

Why This Role Matters

    • Production reliability: our store operations, robot workflows, and customer systems depend on stable infrastructure.
    • Hands-on platform evolution: this role improves running systems through code, automation, standards, and direct operation.
    • Cloud-edge complexity: our platform connects cloud services, Kubernetes workloads, store systems, robots, smart shelves, and telemetry flows.
    • High business impact: reliability, deployment quality, and cost control directly affect operational efficiency and customer trust.
    • Team leadership: you will shape the practices, rituals, and technical judgment of a small Infra/SRE team.
    • Standardization opportunity: we are actively improving IaC ownership, environment separation, observability, incident response, and deployment guardrails.

Key Responsibilities

    Reliability & Operations
    • Own reliability practices for production cloud and Kubernetes systems, including SLOs, SLIs, error budgets, alert quality, and operational readiness.
    • Lead incident response for infrastructure-related outages, act as Incident Commander when needed, and drive blameless post-incident reviews.
    • Build and maintain runbooks, dashboards, alerts, and operational tooling that help engineers diagnose and recover systems quickly.
    • Improve backup, restore, disaster recovery, and business continuity practices for critical compute, data, storage, and configuration systems.
    • Identify recurring operational pain and remove it through automation, better design, or clearer ownership.
    • Infrastructure & Platform Engineering
      • Design, operate, and improve Kubernetes environments, including AKS clusters, workload scheduling, autoscaling, ingress, TLS, RBAC, and cluster upgrades.
      • Maintain and evolve Infrastructure as Code using Bicep, Terraform, and Terragrunt across development, staging, infrastructure, and production environments.
      • Own GitOps-style application deployment patterns using Argo CD, Kustomize, Helm, and environment overlays.
      • Improve CI/CD pipelines with safe promotion flows, validation, drift detection, security checks, and rollback strategies.
      • Manage cloud networking, identity, secrets, certificates, and access controls in partnership with Security Engineering.
      • Support application teams with platform guidance for resource requests, scaling, dependency management, release safety, and production readiness.
      • Observability, Security, and Cost
        • Standardize logs, metrics, traces, dashboards, and alerts across services and infrastructure.
        • Operate and improve observability tooling such as Azure Monitor, Application Insights, Log Analytics, Grafana, Loki, Alloy, and scheduled query alerts.
        • Partner with Security Engineering on RBAC, workload identity, managed identities, Key Vault usage, secret rotation, cloud access boundaries, and auditability.
        • Drive cloud cost visibility and optimization through tagging, right-sizing, autoscaling, log retention controls, budget review, and FinOps practices.
        • Keep reliability, security, and cost decisions practical for a growing robotics business.
        • Team Leadership
          • Lead and mentor 3-6 Infra/SRE/Platform engineers while staying deeply hands-on.
          • Set technical direction for infrastructure reliability, deployment automation, observability, and cloud operations.
          • Run design reviews, code reviews, operational reviews, and regular improvement planning.
          • Build healthy on-call practices, escalation paths, incident roles, and postmortem follow-through.
          • Communicate priorities, risks, and tradeoffs clearly to engineering leadership and cross-functional partners.
          • Cross-Team Collaboration
            • Work with Backend and Frontend teams to make services easier to deploy, observe, scale, and operate.
            • Work with Robotics and Operations teams to understand how infrastructure behavior affects live robot and store workflows.
            • Work with Product and Engineering Management to translate non-functional requirements into roadmap items and engineering guardrails.
            • Help teams adopt standard platform patterns without blocking delivery.

Qualifications

    Must Have
    • 5+ years of professional infrastructure, platform, SRE, DevOps, backend infrastructure, or cloud engineering experience.
    • 2+ years in a senior, lead, or technical leadership role with responsibility for production reliability or infrastructure direction.
    • Strong hands-on experience operating production systems on at least one major cloud platform such as Azure, AWS, or GCP.
    • Production Kubernetes experience, including deployments, services, ingress, autoscaling, resource management, RBAC, troubleshooting, and upgrades.
    • Hands-on Infrastructure as Code experience with Terraform, Bicep, CloudFormation, Pulumi, CDK, or similar tools.
    • Experience designing or operating CI/CD pipelines and deployment workflows for production services.
    • Strong Linux, networking, DNS, TLS, and cloud identity fundamentals.
    • Practical observability experience with metrics, logs, traces, dashboards, alerting, and incident response.
    • Experience leading incidents, writing postmortems, and driving corrective actions to completion.
    • Ability to write scripts or small tools in Bash, Python, TypeScript, Go, C#, or a comparable language.
    • Clear communication skills and the ability to work across engineering, operations, security, and product teams.
    • Nice to Have
      • Azure production experience, especially AKS, Azure Monitor, Application Insights, Log Analytics, Key Vault, Entra ID, managed identities, and Azure RBAC.
      • Experience with Argo CD, Kustomize, Helm, cert-manager, Traefik, ingress-nginx, Grafana, Loki, Alloy, or Prometheus-style monitoring.
      • Experience with GitHub Actions OIDC, workload identity, or federated cloud authentication patterns.
      • Experience operating data services such as PostgreSQL, Redis, Cosmos DB, MongoDB-compatible databases, or cloud storage systems.
      • Experience with drift detection, policy-as-code, cloud guardrails, or compliance-oriented infrastructure workflows.
      • Background in robotics, IoT, retail operations, edge computing, telemetry systems, or other cyber-physical production environments.
      • Experience building or improving an on-call program for a growing engineering organization.
      • Japanese language skills are helpful but not required.

Success Profile

    • Hands-on operator: comfortable debugging real incidents, reading manifests, reviewing Terraform/Bicep, and using cloud/Kubernetes CLIs directly.
    • Reliability-minded: thinks in SLOs, failure modes, blast radius, recovery paths, and operational feedback loops.
    • Pragmatic leader: balances engineering quality with business urgency and team capacity.
    • Automation-oriented: turns repeated manual work into durable tools, workflows, or platform patterns.
    • Security-aware: treats access, secrets, identity, network boundaries, and auditability as part of daily infrastructure work.
    • Cost-conscious: understands that cloud architecture must be reliable and financially sustainable.
    • Collaborative teacher: raises the operational maturity of surrounding teams through guidance, reviews, and shared standards.

Vision & Growth

    • Build a stronger SRE and platform engineering practice for retail robotics.
    • Standardize infrastructure ownership across cloud resources, Kubernetes clusters, manifests, CI/CD, observability, and incident response.
    • Help scale the platform from today’s operational needs toward enterprise-grade reliability across more stores, robots, smart shelves, and customer environments.
    • Shape the long-term infrastructure roadmap while remaining close to the systems that keep the business running.

Skills Required

  • 5+ years of professional infrastructure, platform, SRE, DevOps, backend infrastructure, or cloud engineering experience
  • 2+ years in a senior, lead, or technical leadership role responsible for production reliability or infrastructure direction
  • Hands-on production systems experience on at least one major cloud platform such as Azure, AWS, or GCP
  • Production Kubernetes experience, including deployments, services, ingress, autoscaling, resource management, RBAC, troubleshooting, and upgrades
  • Hands-on Infrastructure as Code experience with Terraform, Bicep, CloudFormation, Pulumi, CDK, or similar tools
  • Experience designing or operating CI/CD pipelines and production deployment workflows
  • Strong Linux, networking, DNS, TLS, and cloud identity fundamentals
  • Practical observability experience with metrics, logs, traces, dashboards, alerting, and incident response
  • Experience leading incidents, writing postmortems, and driving corrective actions to completion
  • Ability to write scripts or small tools in Bash, Python, TypeScript, Go, C#, or a comparable language
  • Clear communication skills and ability to work across engineering, operations, security, and product teams
  • Azure production experience, especially AKS, Azure Monitor, Application Insights, Log Analytics, Key Vault, Entra ID, managed identities, and Azure RBAC
  • Experience with Argo CD, Kustomize, Helm, cert-manager, Traefik, ingress-nginx, Grafana, Loki, Alloy, or Prometheus-style monitoring
  • Experience with GitHub Actions OIDC, workload identity, or federated cloud authentication patterns
  • Experience operating PostgreSQL, Redis, Cosmos DB, MongoDB-compatible databases, or cloud storage systems
  • Experience with drift detection, policy-as-code, cloud guardrails, or compliance-oriented infrastructure workflows
  • Background in robotics, IoT, retail operations, edge computing, telemetry systems, or other cyber-physical production environments
  • Experience building or improving an on-call program for a growing engineering organization
  • Japanese language skills

Similar Jobs

Braze Logo Braze

Senior Delivery Manager

Marketing Tech • Mobile • Software
Easy Apply
Hybrid
Tokyo, JPN
2000 Employees

Braze Logo Braze

Senior Customer Success Manager

Marketing Tech • Mobile • Software
Easy Apply
Hybrid
Tokyo, JPN
2000 Employees

Datadog Logo Datadog

Executive Assistant

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
Tokyo, JPN
6500 Employees

Mastercard Logo Mastercard

Director, Specialist Sales, Consumer Solutions

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Tokyo, JPN
38800 Employees
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
124 Employees
Year Founded: 2017

What We Do

Telexistence Inc. is a robotics company founded in 2017 that designs, manufactures, and operates remote-controlled and AI-powered robots. Their mission is to transform robotics and automate tasks, particularly within the retail sector, such as beverage shelf-stocking, to reduce labor burdens and improve operational efficiency. They focus on pushing the state of the art in robotics to create meaningful change in the world.

Similar Companies Hiring

LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account