Staff Infrastructure Engineer

Posted Yesterday
2 Locations
Remote
145K-260K Annually
Senior level
Security • Database • Cybersecurity
The Role
Lead the architecture, scaling, reliability, and security of multi-region AWS infrastructure supporting high-throughput payment systems. Build immutable, self-healing environments with Terraform, Kubernetes, GitOps, CI/CD, and automated resiliency. Develop observability using Prometheus, Grafana, and OpenTelemetry; manage incidents and post-mortems; improve engineering velocity through Golden Paths and platform enablement; and mentor engineering teams on reliable infrastructure practices.
Summary Generated by Built In
About VGS

VGS is building the trust infrastructure for a new era of commerce. As commerce becomes agentic, payments are increasing dramatically in volume, speed, and complexity, and the companies driving that shift need a foundation they can fully trust.

That's where VGS comes in. We're the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies, merchants, banks, and fintechs to power modern payments and agentic commerce. We store more than 9 billion tokens and process more than 10 billion monthly interactions, touching a third of all e-commerce. That scale isn't incidental. It's why leading companies embed our universal token vault and credential management platform directly into their stack, taming the complexity of payment data so they can move faster.

Come build the platform that enables modern payments for the world's largest companies. We're helping businesses unlock new possibilities in an industry that never stops moving. And that takes exceptional people.

 

 

Trusted Engineering That Matters

Engineering at VGS

VGS is building the platform powering agentic commerce and global payments infrastructure. VGS Engineering is the pillar that supports that vision. We create a platform that is an indispensable, ubiquitous component of our partners' payments and security infrastructure. We build and operate high-scale, reliable, and resilient systems that manage mission-critical payment data across the globe. We don't just participate in the payments ecosystem; we power it.

 

Our Environment & Culture

  • Modern Tooling as Leverage: We value speed and impact. To ensure nothing slows you down, we arm you with a highly consistent modern stack (e.g. AWS, Kubernetes, Java, TypeScript, Python) and world-class AI-augmented SDLC tooling. We strip away the friction so you can focus on the problems that demand creativity and judgment.

  • Ownership of Outcomes: Scaling global payments infrastructure requires individuals who take personal ownership of outcomes from day one. We look for builders who combine high agency with deep strategic alignment; channeling their urgency and relentless iteration into what matters most. 

  • Mission-Critical Accountability: We power payments infrastructure for global market leaders, making system trust and integrity our core deliverable. We don't treat compliance (PCI, SOC2, ISO27001), security patching, and platform maintenance as administrative overhead—we engineer them with first-class product discipline.  In payments infrastructure, reliability is not background work, it’s part of what customers buy.

Learn More: See how our engineering practices drive trust and innovation in our official blog and our reliability practices.

 

 

About the Role

As a Staff Infrastructure Engineer, you will serve as a technical leader on our Platform Engineering team. You will design, scale, and fortify global cloud infrastructure designed to handle mission-critical, high-throughput payments applications with zero downtime.

You will take ownership of key platform foundations powering our core payments infrastructure. Rather than executing against a rigid task list, you will drive the execution, architecture, and reliability standards for your designated domains within our multi-region AWS environment. We are looking for high-agency engineers who want to own complex distributed system challenges end-to-end and elevate how the engineering team operates. If you thrive on solving complex distributed systems problems, building automated resiliency, and driving modern SRE practices, we want to build the future with you.

 

 What You'll Do
  • Build Immutable, Self-Healing Systems: Design, build, and optimize multi-region, high-availability AWS infrastructure. You will execute our evolution from hand-crafted environments to a standardized, globally scalable fleet managed entirely through code.

  • Drive Resiliency & Automation: Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and IaC.

  • Deep Observability & Resiliency: Implement end-to-end telemetry (Prometheus, Grafana, OpenTelemetry) to proactively spot bottlenecks. You will drive incident management and conduct blameless post-mortems to continuously harden our reliability baseline.

  • Force-Multiply Engineering Velocity: Partner closely with Product, Security, and Core Engineering teams to implement "Golden Paths" that strip away friction for feature teams. You will influence engineering practices within your domain and mentor engineers on how to move fast with high alignment.

  • Customer Impact: Implement and operate high-performance, low-latency private connectivity to optimize the experience for external customers. Partner with internal engineering teams at the design level for platform enablement and adoption.

 

 

What You Bring
  • Ownership at Scale: 7+ years of experience taking personal ownership of outcomes in complex, large-scale distributed systems within mission-critical environments.

  • AWS & Infrastructure-as-Code: Strong proficiency in AWS ecosystems leveraging Terraform to build reproducible environments.

  • Containerization & Orchestration: Strong, hands-on experience with Kubernetes (EKS), Docker, and GitOps workflows (Flux, Argo, GitHub Actions).

  • Automation & Scripting: Strong coding skills in Python, Go, or Bash to automate infrastructure and build operational tools.

  • Observability Expertise: Experience implementing and maintaining Prometheus, Grafana, or OpenTelemetry.

  • Security & Networking Foundations: Solid understanding of cloud security, API Gateways, load balancing, and network isolation, viewing security as a fundamental engineering constraint, not an afterthought.

Nice to Have:

  • Experience with tokenization, payment processing, or security products
  • BA/BS degree

  • A knack for out-of-the-box thinking that thrives in a fast-paced startup environment

  • Experience managing distributed data streaming platforms like Kafka (MSK).

  • Database performance tuning and query optimization skills.

  • Familiarity with Java / Spring Framework services.

 

Skills Required

  • 8+ years of experience owning outcomes in complex, large-scale distributed systems within mission-critical environments
  • Advanced proficiency with AWS and Terraform for infrastructure as code
  • Hands-on experience with Kubernetes, including EKS, Docker, and GitOps workflows
  • Experience with Flux, Argo, and GitHub Actions
  • Strong coding skills in Python, Go, or Bash for infrastructure automation and operational tools
  • Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale
  • Understanding of cloud security, API gateways, load balancing, and network isolation
  • Experience with tokenization, payment processing, cryptology, or security products
  • BA or BS degree
  • Experience managing distributed data streaming platforms such as Kafka or Amazon MSK
  • Database performance tuning and query optimization skills
  • Familiarity with Java and the Spring Framework
  • Ability to thrive in a fast-paced startup environment with out-of-the-box thinking
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
223 Employees
Year Founded: 2015

What We Do

Providing essential security and compliance infrastructure, Very Good Security (VGS) enables startups and enterprises to focus on their core business instead of compliance and regulatory overhead. With one single integration, VGS customers unlock the value of sensitive data without the cost and liability of securing it themselves, while also accelerating compliances like PCI, SOC 2 and more.

Similar Jobs

Affirm Logo Affirm

Staff Software Engineer

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
Canada
2200 Employees
181K-241K Annually

Very Good Security Logo Very Good Security

Infrastructure Engineer

Security • Database • Cybersecurity
Remote
2 Locations
223 Employees
185K-290K Annually

Nango Logo Nango

Staff Engineer

Artificial Intelligence • Software • Generative AI • Infrastructure as a Service (IaaS)
Remote
35 Locations
17 Employees
140K-220K Annually

Scribd, Inc. Logo Scribd, Inc.

Staff Software Engineer

Artificial Intelligence • Consumer Web • Digital Media • Software
In-Office or Remote
24 Locations
294 Employees
145K-275K Annually

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account