Sr. Staff Infrastructure Engineer

Posted 9 Hours Ago
2 Locations
Remote
185K-290K Annually
Senior level
Security • Database • Cybersecurity
The Role
Lead platform engineering for high-scale, multi-region AWS payment infrastructure. Architect immutable, highly available, self-healing systems; automate infrastructure through Terraform, GitOps, CI/CD, and Kubernetes; establish observability with Prometheus, Grafana, and OpenTelemetry; own incident management and resiliency improvements; design private connectivity; and mentor engineers while driving organization-wide platform standards.
Summary Generated by Built In
About VGS

VGS is the world's leader in payment tokenization, trusted by the most innovative AI and Fortune 500 companies, merchants, banks, and fintechs to power modern payments and agentic commerce.

We tokenize over 9 billion tokens and process more than 10 billion monthly interactions, touching a third of all e-commerce. That scale isn't incidental. It's why leading companies embed our universal token vault and credential management platform directly into their stack, taming the complexity of payment data so they can move faster.

We’re helping businesses unlock new possibilities in an industry that never stops moving. And that takes exceptional people.

 


Trusted Engineering That Matters

Engineering at VGS

Our vision is clear: VGS is building the platform powering agentic commerce and global payments infrastructure. VGS Engineering is the pillar that supports that vision. We create a platform that is an indispensable, ubiquitous component of our partners' payments and security infrastructure. To achieve this, we build and operate high-scale, reliable, and resilient systems that manage mission-critical payment data across the globe. We don't just participate in the payments ecosystem; we power it.

 

Our Environment & Culture

  • Tooling as Leverage for Impact: We value speed and impact. To ensure nothing slows you down, we arm you with a highly consistent modern stack (e.g. AWS, Kubernetes, Java, TypeScript, Python) and world-class AI-augmented SDLC tooling. We strip away the friction so you can focus on building simple things that work, pushing boundaries, and delivering value faster.

  • Impact Over Input: Scaling global payments infrastructure requires individuals who take personal ownership of outcomes from day one. We look for builders who combine high agency with deep strategic alignment; channeling their urgency and relentless iteration into what matters most. Leading from the front means multiplying the impact of those around you so we win collectively.

  • Mission-Critical Accountability: We power payments infrastructure for global market leaders, making system trust and integrity our core deliverable. We don't treat compliance (PCI, SOC2, ISO27001), security patching, and platform maintenance as administrative overhead—we engineer them with first-class product discipline. We apply the same rigor, roadmap thinking, and speed to system health as we do to new features. In high-scale payments, reliability isn't background work; it is our product.

 

Learn More: See how our engineering practices drive trust and innovation in our official blog and our reliability practices.

 


About the Role

As a Senior Staff Infrastructure Engineer, you will serve as a technical leader on our Platform Engineering team. You will architect, scale, and fortify global cloud infrastructure designed to handle mission-critical, high-throughput payments applications with zero downtime.

You will take ownership of key platform foundations powering our core payments infrastructure. Rather than executing against a rigid task list or holding blanket ownership over the entire platform, you will drive the technical strategy, architecture, and reliability standards for your designated domains within our multi-region AWS environment. We are looking for high-agency engineers who want to own complex distributed system challenges end-to-end and elevate how the entire engineering team operates.

If you thrive on solving complex distributed systems problems, building automated resiliency, and driving modern SRE practices, we want to build the future with you.

 


What You'll Do
  • Build Immutable, Self-Healing Systems: Design, build, and optimize multi-region, high-availability AWS infrastructure. You will drive our evolution from hand-crafted environments to a standardized, globally scalable fleet managed entirely through code.

  • Drive Resiliency & Automation: Replace manual toil with self-healing, automated infrastructure using GitOps, modern CI/CD pipelines, and IaC.

  • Deep Observability & Resiliency: Build end-to-end telemetry (Prometheus, Grafana, OpenTelemetry) to proactively spot bottlenecks. You will own incident management and conduct blameless post-mortems to continuously harden our reliability baseline.

  • Force-Multiply Engineering Velocity: Partner closely with Product, Security, and Core Engineering teams and lead from the front by designing "Golden Paths" that strip away friction for feature teams. You will influence company-wide engineering practices and mentor the organization on how to move fast with high alignment.

  • Customer Impact: Architect and operate high-performance, low-latency private connectivity to optimize the experience for external customers. Partner strategically with internal engineering teams at the design and architectural level for platform enablement and adoption.

 


What You Bring
  • Ownership at Scale: 10+ years of experience taking personal ownership of outcomes in complex, large-scale distributed systems within mission-critical environments.

  • AWS & Infrastructure-as-Code: Advanced proficiency in AWS ecosystems leveraging Terraform to build reproducible environments.

  • Containerization & Orchestration: Strong, hands-on experience with Kubernetes (EKS), Docker, and GitOps workflows (Flux, Argo, GitHub Actions).

  • Automation & Scripting: Strong coding skills in Python, Go, or Bash to automate infrastructure and build operational tools.

  • Observability Expertise: Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale.

  • Security & Networking Foundations: Solid understanding of cloud security, API Gateways, load balancing, and network isolation, viewing security as a fundamental engineering constraint, not an afterthought.

Nice to Have:

  • Experience with tokenization, payment processing, cryptology, or security products

  • BA/BS degree

  • A knack for out-of-the-box thinking that thrives in a fast-paced startup environment

  • Experience managing distributed data streaming platforms like Kafka (MSK).

  • Database performance tuning and query optimization skills.

  • Familiarity with Java / Spring Framework services.

 

Skills Required

  • 10+ years of experience with complex, large-scale distributed systems in mission-critical environments
  • Advanced proficiency with AWS and Terraform
  • Hands-on experience with Kubernetes, including EKS, Docker, and GitOps workflows
  • Experience with Flux, Argo, and GitHub Actions
  • Strong coding skills in Python, Go, or Bash
  • Deep experience implementing Prometheus, Grafana, or OpenTelemetry at scale
  • Understanding of cloud security, API gateways, load balancing, and network isolation
  • Experience with tokenization, payment processing, cryptology, or security products
  • BA or BS degree
  • Experience managing distributed data streaming platforms such as Kafka or MSK
  • Database performance tuning and query optimization skills
  • Familiarity with Java and Spring Framework services
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
223 Employees
Year Founded: 2015

What We Do

Providing essential security and compliance infrastructure, Very Good Security (VGS) enables startups and enterprises to focus on their core business instead of compliance and regulatory overhead. With one single integration, VGS customers unlock the value of sensitive data without the cost and liability of securing it themselves, while also accelerating compliances like PCI, SOC 2 and more.

Similar Jobs

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
2 Locations
2500 Employees
126K-314K Annually
Remote or Hybrid
6 Locations
200 Employees
157K-234K Annually

GitLab Logo GitLab

Staff Commercial Pricing Strategist

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
2 Locations
2500 Employees
139K-235K Annually
Remote
2 Locations
529 Employees
145K-183K Annually

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account