Platform Infrastructure Engineer (SRE Core)

Reposted 8 Days Ago
27 Locations
Remote
Mid level
Cloud • Security • Cybersecurity
The Role
Build, operate, and automate Menlo Security's global cloud-native platform across GCP and AWS. Manage VM fleets and Kubernetes clusters, author Terraform IaC and Spacelift workflows, implement observability with Grafana/Prometheus/OTel, handle networking, DNS, certificates, and service mesh, and reduce operational toil through automation while participating in 24x7 on-call rotation and cross-team architecture and DR planning.
Summary Generated by Built In

Menlo Security is the leader in Browser Security for human and agentic workforces. Our mission is to enable humans and agents to connect, communicate, and collaborate securely, without compromise. The Menlo Browser Security Platform protects organizations from cyberattacks by stopping threats across the web, documents, and email before they reach the user. With Menlo Agent Runtime Security (MARS), that protection now extends to the AI agents working alongside every employee. Menlo Security is trusted by major global businesses, including Fortune 500 companies and government agencies, to protect their most valuable asset, their data, and is backed by top-tier investors.

About the Role

Platform Infrastructure Engineering is responsible for building and operating Menlo Security's Infrastructure Platform. Together with the rest of our engineering teams, we enable our customers to connect to the Internet without compromise. Our environment provides services globally. We expect failure, build security in by design, create evolvable systems, and enable multi-tenancy across the infrastructure. Automation is an absolute for us.

We are committed to getting it done properly, the first time.

As a Platform Infrastructure Engineer, you'll join a group of experienced engineers who are part of a globally distributed team responsible for building and managing the company's core infrastructure services and maintaining our constantly growing platform. The team operates a sophisticated cloud-native infrastructure built on Google Kubernetes Engine and VMs spanning multiple environments globally from development to production. We manage infrastructure as code with Terraform and Spacelift orchestration, and deploy services using Helm charts. Our platform emphasizes security-first design, comprehensive observability, and multi-region resilience. Success in this role requires working with a vast VM fleet in AWS and GCP as well as Kubernetes, writing Infrastructure as Code, and a passion for automation and reliability engineering.

Responsibilities

  • Design, deploy, and maintain VM and Kubernetes infrastructure on GCP and AWS across dozens of clusters spanning development, staging, and production environments in multiple regions.

  • Coordinate with your peers in your direct team as well as across teams to ensure that the tasks you’re working on are going to solve the problems that we need them to solve.

  • Build and maintain Infrastructure as Code (IaC) using Terraform modules, managing resources through Spacelift or equivalent Terraform Automation and Collaboration Software (TACOS). Provision cloud infrastructure including networking, compute, storage, and security components primarily on GCP, with secondary AWS support.

  • Implement and manage workflows with sophisticated multi-layer configuration management.

  • Build and maintain comprehensive observability solutions using Grafana Cloud, Prometheus/Mimir, and OTel collectors. Design Grafana dashboards, configure alerting rules, and ensure visibility across all platform components.

  • Manage certificate lifecycle, DNS automation, ingress controllers, and service mesh networking with Cilium.

  • Partner with Engineering, Product, Compliance, and Security teams to design resilient, scalable systems. Consult on capacity planning, disaster recovery, and architectural decisions for cloud-native applications.

  • Identify and eliminate toil through automation. Write scripts, develop tools, and build CI/CD pipelines to improve operational efficiency and reduce manual work.

  • Participate in a 24x7 on-call rotation as part of a globally distributed team, responding to incidents and driving post-incident reviews.

Requirements

  • Bachelor's degree in Computer Science, similar technical field of study, or equivalent practical experience.

  • Proficiency in common programming & scripting languages. We use a lot of python, bash and go.

  • Understanding of network topologies, communication protocols (ie. TCP/IP, HTTP/S, UDP, TLS) and enterprise grade connectivity solutions.

  • Kubernetes expertise including cluster administration, RBAC, networking, workload management, and troubleshooting across production environments.

  • Proven experience with Terraform for infrastructure provisioning and management.

  • Knowledge of Google Cloud Platform services including GKE, VPC networking, Cloud DNS, Artifact Registry, Secret Manager, IAM, Gemini Code Assist, and Workload Identity.

  • Experience with GitOps methodologies and tools.

  • Clear understanding of how to use LLM code assist tools to effectively build software.

MSGL-I4

Follow us on LinkedIn!

Why Menlo?

At Menlo, we don't settle for the status quo — in our technology or our culture. How we think and act is just as important as what we build. Our culture is defined by five core mindsets: Proactive Leadership, Straight Talk, United Impact, Elevated Talent, and Customer-Compelled. We take ownership and drive outcomes without waiting to be told. We communicate directly and seek hard truths. We break down silos and win together. We hold a high bar for ourselves and the people around us. And we treat every customer interaction as mission-critical. If you're someone who sees it, owns it, solves it, and does it — you'll thrive here.

All qualified applicants will receive consideration for employment without regard to race, sex, color, religion, sexual orientation, gender identity, national origin, protected veteran status, or on the basis of disability.

TO ALL AGENCIES: Please, no phone calls or emails to any employee of Menlo Security outside of the Talent organization. Menlo Security’s policy is to only accept resumes from agencies via Ashby (ATS). Agencies must have a valid services agreement executed and must have been assigned by the Talent team to a specific requisition. Any resume submitted outside of this process will be deemed the sole property of Menlo Security. In the event a candidate submitted outside of this policy is hired, no fee or payment will be paid.

Skills Required

  • Bachelor's degree in Computer Science or equivalent practical experience
  • Proficiency in Python, Bash, and Go (programming and scripting)
  • Kubernetes expertise: cluster administration, RBAC, networking, workload management, production troubleshooting
  • Proven experience with Terraform for infrastructure provisioning and management
  • Experience managing cloud infrastructure on GCP (GKE, VPC, Cloud DNS, Artifact Registry, Secret Manager, IAM, Workload Identity) and AWS VMs
  • Experience with Terraform automation/orchestration (Spacelift or equivalent TACOS) and GitOps methodologies/tools
  • Observability implementation experience using Grafana Cloud, Prometheus/Mimir, and OpenTelemetry collectors; dashboarding and alerting
  • Experience with service mesh and networking tooling (Cilium), ingress controllers, DNS automation, and certificate lifecycle management
  • Strong networking knowledge: TCP/IP, HTTP/HTTPS, UDP, TLS and enterprise connectivity solutions
  • Experience building automation, scripts, tools, and CI/CD pipelines to reduce operational toil
  • Familiarity with LLM code assist tools (e.g., Gemini Code Assist) and ability to use them effectively
  • Ability to participate in a 24x7 on-call rotation and perform incident response and post-incident reviews
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mountain View, CA
312 Employees
Year Founded: 2013

What We Do

Menlo Security enables organizations to outsmart threats, completely eliminating attacks and fully protecting productivity with a one-of-a-kind, isolation-powered cloud security platform. It’s the only solution to deliver on the promise of cloud security—by providing the most secure Zero Trust approach to preventing malicious attacks; by making security invisible to end users while they work online; and by removing the operational burden for security teams. Now organizations can offer a safe online experience, empowering users to work without worry while they keep the business moving forward.

Similar Jobs

Mondelēz International Logo Mondelēz International

Program Manager

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
9 Locations
90000 Employees
4K-4K Annually

Pfizer Logo Pfizer

Quality Assurance Manager

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Remote or Hybrid
28 Locations
121990 Employees

Mondelēz International Logo Mondelēz International

Change Manager o9 MEU, Demand Planning

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
9 Locations
90000 Employees

Pfizer Logo Pfizer

Manager, ServiceNow ITOM Engineer

Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
In-Office or Remote
2 Locations
121990 Employees

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account