Tech Infra Engineer

Posted 3 Days Ago
Be an Early Applicant
2 Locations
In-Office
Expert/Leader
eCommerce • Fintech • Logistics • Retail
The Role
Design and build a unified Kubernetes-based application platform for hybrid, multi-cluster, multi-region, and multi-cloud environments. Develop controllers, operators, node daemons, automation, control planes, scaling systems, and failover capabilities that support millions of pods. Improve reliability, performance, cost, latency, security, governance, and observability through resilient multi-tenant systems, automated remediation, and SLO-driven controls. Partner with product leaders and customers to create reusable platform primitives and developer-friendly experiences.
Summary Generated by Built In

Please complete the attached Internal Transfer Request Form and submit.  

Please make sure to apply with your Coupang e-mail address.  

   

As a Staff Systems Engineer in Developer Platform, you will partner with leaders of multiple platform teams. You will work closely with product to define and implement simple solutions to complex orchestration problems, building a highly scalable, reliable, and efficient platform for our customers. You will engineer and develop Kubernetes controllers, operators, and node-level daemons for the application runtime; drive performance tuning and scaling; and design multi-cluster control-plane capabilities that scale to millions of pods across thousands of clusters.

What You Will Do

  • Engineer and develop a unified application platform for hybrid (multi-cluster, multi-region, multi-cloud) application management using Kubernetes controllers and feedback-driven control systems to meet SLOs.
  • Deliver end-to-end automation for application lifecycle (deployments, rollouts, failovers, policy enforcement) to minimize manual work for users.
  • Drive fleet-wide optimization for cost, performance, and latency through data-informed controls and capacity management, improving $/RPS and tail latency.
  • Build resilient, multi-tenant control planes and workflows that safely scale to millions of pods across thousands of clusters.
  • Ensure reliability, security, and governance with clear guardrails, safe defaults, and automated remediation.
  • Partner with product and customers to turn complex orchestration problems into simple, reusable platform primitives and great developer experiences.
  • Champion observability and continuous improvement with measurable, outcome-focused metrics.

Basic Qualifications

  • Bachelor’s degree in Computer Science, Electrical Engineering, Math, or a closely related field (or equivalent experience)
  • 10+ years in backend software development and operations
  • Recent experience designing and operating large-scale distributed systems (last 3 years) • Fluency in one or more among Go, C/C++, Python, or Java
  • Proven track record of delivering mission-critical systems
  • Experience with cloud computing using AWS or Azure or GCP

 

Preferred Qualifications

  • Kubernetes API machinery and semantics: SSA, SMP, server-side dry-run, watches/informers/listers, rate-limited workqueues, finalizers, owner references, leader election, API Priority and Fairness
  • Controllers/operators and node daemons in Go: client-go/controller-runtime, reconciliation patterns, backoff and retry, idempotency, partitioned/sharded controllers, HA and failover
  • CRDs and webhooks: versioning, conversion functions/webhooks, validating/mutating admission webhooks, policy frameworks and best practices
  • Pod/runtime semantics: sidecars, init/ephemeral containers, probes (readiness/liveness/startup), lifecycle hooks, termination behavior, PDBs, QoS classes, ResourceQuota/LimitRange, topology spread, affinity/anti-affinity
  • Scaling systems: HPA (resource/custom/external metrics), VPA, cluster autoscaler; multi-dimensional scaling, health-aware/autopilot-style policies; external metrics adapters and SLO-driven scaling
  • Federated and multi-cluster: placement/propagation, failover, drift detection, reconciliation strategies; consistent hashing and partitioning for scale
  • Distributed systems: CRDTs and eventual consistency paradigms; Raft/memberlist/gossip; deep familiarity with etcd, Kafka, Redis and their operational characteristics (compaction, backpressure, retention, failover)
  • Observability and data: Prometheus (cardinality control, recording rules), tracing; experience with vector databases for search and diagnostics; strong time-series forecasting (classical + ML) and statistical modeling for proactive optimization
  • Languages and interfaces: Go (primary), Java/Python as needed; gRPC/protobuf; JSON/YAML/Jsonnet
  • Leadership: ability to handle multiple competing priorities in a fast-paced environment and lead the delivery of large-scale services for complex business offerings

 

Recruitment Process  

  • Application Review - Phone Interview - Onsite (or Virtual Onsite) Interview – Offer
  • The exact nature of the recruitment process may vary according to the specific job and may be changed due to scheduling or other circumstances.
  • Interview schedules and the results will be informed to the applicant via the e-mail address submitted at the application stage. 

 

Details to Consider 

  • This job posting may be closed prior to the stated end date for application if all openings are filled.
  • Coupang has the right to rescind an offer of employment if a candidate is found to have submitted false information as part of the application process.
  • Those eligible for employment protection (recipients of veteran’s benefits, the disabled, etc.) may receive preferential treatment for employment in accordance with applicable laws. 

 

Privacy Notice​ 

  • Your personal information will be collected and managed by Coupang as stated in the Application Privacy Notice located below: https://www.coupang.jobs/privacy-policy/ 

 

 

Please complete the attached Internal Transfer Request Form and submit.  

Please make sure to apply with your Coupang e-mail address. 


Skills Required

  • Bachelor's degree in Computer Science, Electrical Engineering, Mathematics, a closely related field, or equivalent experience
  • 10+ years of backend software development and operations experience
  • Recent experience designing and operating large-scale distributed systems within the last three years
  • Fluency in one or more of Go, C, C++, Python, or Java
  • Proven track record of delivering mission-critical systems
  • Experience with cloud computing using AWS, Azure, or GCP
  • Experience with Kubernetes API machinery and semantics
  • Experience building Kubernetes controllers, operators, and node daemons in Go
  • Experience with CRDs, versioning, conversion functions, and admission webhooks
  • Knowledge of pod and runtime semantics, scaling systems, and SLO-driven scaling
  • Experience with federated and multi-cluster systems, failover, drift detection, and partitioning
  • Deep familiarity with distributed systems and etcd, Kafka, Redis, Raft, memberlist, and gossip operations
  • Experience with Prometheus, tracing, vector databases, time-series forecasting, and statistical modeling
  • Experience with Go, Java, Python, gRPC, Protocol Buffers, JSON, YAML, and Jsonnet
  • Ability to manage competing priorities and lead delivery of large-scale services
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
108,000 Employees
Year Founded: 2010

What We Do

Coupang is a U.S. technology and global commerce company founded in 2010, operating online retail and cross-border commerce, restaurant delivery, video streaming, and fintech/payment services. Through brands including Coupang, Eats, Play, Rocket Now, and Farfetch, it uses technology, logistics, and fulfillment infrastructure to serve millions of customers in Korea, Taiwan, the United States, and more than 190 countries and territories worldwide.

Similar Jobs

Dynatrace Logo Dynatrace

Senior Data Analyst

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
United States
5600 Employees
90K-110K Annually

Agero Logo Agero

Remote Response Associate, Roadside Assistance (CSR)

Automotive • Big Data • Insurance • Software • Transportation
Easy Apply
Remote or Hybrid
Tennessee, USA
1600 Employees
16-16 Hourly

Agero Logo Agero

Customer Service Representative

Automotive • Big Data • Insurance • Software • Transportation
Easy Apply
Hybrid
Clarksville, TN, USA
1600 Employees
17-17 Hourly

Enverus Logo Enverus

Consulting Geologist, Anadarko Basin - Contract - 26330

Big Data • Information Technology • Software • Analytics • Energy
In-Office or Remote
2 Locations
1800 Employees
60-75 Hourly

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account