Senior Infrastructure Engineer

Posted 9 Days Ago
Be an Early Applicant
Auckland, NZL
In-Office
Senior level
Artificial Intelligence • Edtech • Gaming • Software
The Role
Architect, build, and operate Level’s AWS cloud infrastructure and developer platform. Own Kubernetes/EKS, infrastructure as code, GitOps, CI/CD, networking, DNS, observability, security tooling, disaster recovery, and cost optimization. Lead incident response and post-mortems, establish infrastructure standards, enable developer self-service, and mentor engineers. The role requires strong production experience with AWS, Terraform/OpenTofu, Kubernetes, automation, networking, security, CI/CD, and observability.
Summary Generated by Built In

About Level
Level is a learning technology company dedicated to helping students build real academic and life skills with confidence and joy. We combine proven curriculum principles with world class interactive design to make meaningful practice something students want to come back to, not something they struggle through. We support what teachers, schools, and parents are already doing by increasing student engagement with high quality, standards aligned practice that reinforces classroom learning.

As an Senior Infrastructure Engineer on the Platform team, you will architect, build, and operate the cloud infrastructure and developer platform that every Level product runs on. You will own critical infrastructure end-to-end — the Kubernetes platform, infrastructure-as-code, CI/CD and GitOps delivery, networking (ingress and egress), DNS, observability, and cloud security posture — and provide the reliable, self-service foundations the rest of engineering builds on. You will work on a small, senior-leaning team where infrastructure decisions have direct, visible impact on reliability, performance, cost, and developer velocity.

 

What You'll Do

  • Cloud Infrastructure & IaC — Design, build, and operate secure, highly available AWS infrastructure using Terraform/OpenTofu with a GitOps workflow (Atlantis). Own capacity planning, DR, and cost optimization for the systems you run.

  • Kubernetes & Platform Operations — Operate and evolve EKS: autoscaling (Karpenter), upgrades, core add-ons, and Helm-based delivery (ArgoCD).

  • CI/CD & Developer Enablement — Build and maintain GitHub Actions pipelines that let platform and product teams ship fast and safely, with self-service tooling where it makes sense.

  • Networking, Ingress & DNS — Own ingress/egress (Traefik), service mesh and mTLS (Linkerd/Envoy), load balancing, edge TLS, and DNS (Route 53, Terraform-managed).

  • Observability & Reliability — Build observability with OpenTelemetry and SigNoz; use telemetry to drive reliability, performance, and cost decisions. Serve as an escalation point for complex incidents, leading troubleshooting and post-mortems.

  • Security & Compliance — Apply cloud security best practices across identity, secrets, and network boundaries, with particular care for student data and K-12 privacy. Operate posture/vulnerability tooling (Security Hub, GuardDuty, Inspector, Snyk) and org guardrails (Control Tower, SCPs).

  • Ownership & Mentorship — Set standards, mentor engineers, and leave the platform better than you found it.


What You Need

  • 5+ years operating large-scale cloud infrastructure (AWS strongly preferred)

  • Deep IaC experience (Terraform/OpenTofu; CloudFormation/Pulumi/CDK also relevant)

  • Strong Docker/Kubernetes (EKS) production experience

  • Scripting/automation proficiency (Python, Go, or Bash)

  • Solid cloud networking fundamentals (VPC, DNS, load balancing, ingress, firewalls/WAF, VPNs) and security best practices

  • Proven CI/CD and GitOps experience (GitHub Actions or similar)

  • Observability experience (metrics/logs/traces) used to drive real decisions

  • Track record leading infrastructure projects independently, end to end

  • Strong communication across technical and non-technical audiences

     

Nice to Have

  • ArgoCD, Atlantis, Linkerd/Envoy, Traefik, Karpenter, Helm

  • OpenTelemetry, SigNoz (our stack), Grafana, Datadog, or Prometheus

  • Backstage or other internal developer platform experience

  • AI/ML infra experience (GPU scheduling, model/agent hosting, inference gateways)

  • Rust service CI/CD, CloudFront/CDN experience

  • AWS Solutions Architect / DevOps Engineer – Professional certification

  • Distributed-systems background, OSS infrastructure contributions

 

Skills Required

  • 5+ years operating large-scale cloud infrastructure
  • AWS cloud infrastructure experience
  • Deep infrastructure-as-code experience with Terraform or OpenTofu
  • Production experience with Docker and Kubernetes, preferably Amazon EKS
  • Scripting or automation proficiency in Python, Go, or Bash
  • Strong cloud networking fundamentals, including VPC, DNS, load balancing, ingress, firewalls/WAF, and VPNs
  • Cloud security best-practices experience
  • Proven CI/CD and GitOps experience, including GitHub Actions or similar
  • Observability experience using metrics, logs, and traces to drive operational decisions
  • Track record leading infrastructure projects independently from end to end
  • Strong communication with technical and non-technical audiences
  • Experience with ArgoCD, Atlantis, Linkerd or Envoy, Traefik, Karpenter, or Helm
  • Experience with OpenTelemetry, SigNoz, Grafana, Datadog, or Prometheus
  • Internal developer platform experience with Backstage or similar tools
  • AI/ML infrastructure experience, including GPU scheduling, model hosting, or inference gateways
  • Rust service CI/CD or CloudFront/CDN experience
  • AWS Solutions Architect or DevOps Engineer Professional certification
  • Distributed-systems background or open-source infrastructure contributions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
218 Employees
Year Founded: 2018

What We Do

Level is an educational technology company that develops learning science products and academic content, focusing on ELA and Math. The company integrates game systems engineering, animation, and AI into its tools, employing a multidisciplinary team of designers, engineers, and content specialists to create its educational platform.

Similar Jobs

Partly Logo Partly

Product Manager

Artificial Intelligence • Machine Learning • Software
In-Office
Auckland, NZL
140 Employees

SearchApi Logo SearchApi

Infrastructure Engineer

Information Technology • Software
In-Office or Remote
59 Locations
5 Employees

Halter Logo Halter

Workforce Planner, Customer Experience

Greentech • Hardware • Internet of Things • Machine Learning • Software • Business Intelligence • Agriculture
In-Office
Auckland, NZL
350 Employees

Pfizer Logo Pfizer

Feasibility Strategic Analytics Lead (FSAL) - Senior Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
33 Locations
121990 Employees
116K-193K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account