Site Reliability Engineering (SRE), The Core Engineering, Vice President, Dallas

Posted 2 Days Ago
Be an Early Applicant
Dallas, TX, USA
In-Office
Expert/Leader
Fintech • Financial Services
The Role
Lead site reliability engineering initiatives for highly available, observable, and resilient platforms. Establish SLOs, SLIs, and error budgets; architect fault-tolerant systems; automate operational processes; improve production readiness through testing, tuning, forecasting, and reviews; and lead complex incident response and blameless post-mortems. Partner with engineering teams and leadership to reduce toil, improve on-call practices, and drive adoption of SRE principles across the organization.
Summary Generated by Built In
WHAT WE DO

Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime.

This role is for software engineers who enjoy solving complex distributed system problems, building tools and platforms that make teams more effective, and championing SRE principles (such as SLOs, error budgets, and blameless post-mortems) across a large engineering organization.

Key Responsibilities
  • Partner with engineering leadership to establish service level objectives (SLOs), service level indicators (SLIs), and error budgets.
  • Collaborate with product developers to architect highly available, fault-tolerant, and self-healing systems. Conduct architectural reviews and introduce patterns like circuit breakers, graceful degradation, and rate limiting.
  • Reduce operational toil by building automation, tooling, and self-service capabilities that remove repetitive manual work.
  • Improve production readiness through load testing, performance tuning, capacity forecasting, and reliability reviews.
  • Lead the response to complex, multi-system production incidents. Facilitate blameless post-mortems to identify root causes and drive long-term preventative actions.
  • Promote sustainable operations by helping design healthy on-call models, clear escalation paths, and balanced pager responsibilities.

WHAT WE ARE LOOKING FOR Core Technical Skills
  • Strong proficiency in at least one major programming language (e.g., Java, Python, or Node.js) with a focus on writing clean, maintainable code for tooling and automation.
  • Hands-on experience with Infrastructure as Code (IaC) frameworks such as Terraform, Ansible, or CloudFormation.
  • Deep understanding of containerization and orchestration technologies, specifically Docker and Kubernetes (K8s), including service meshes and ingress controllers.
  • Advanced experience with major cloud providers (AWS, GCP, or Azure), specifically building and operating highly resilient cloud-native architectures.
  • Proficiency with Observability stacks, including distributed tracing, logging, and metrics (e.g., Prometheus, Grafana, Splunk, Datadog, OpenTelemetry, ELK, or CloudWatch)
  • Experience with automated testing and SDLC concepts, developing applications in a Linux environment, and sound knowledge of algorithms, data structures and software design.
  • Knowledge of networking protocols and load balancing strategies in a distributed systems environment.
Core Competencies & Soft Skills
  • Ability to analyze complex, distributed systems holistically and understand how individual components interact under load.
  • Strong interpersonal skills to collaborate with product developers, influence architectural decisions, prioritize toil reduction, and drive SRE adoption without direct authority.
  • Ability to translate complex technical issues into clear, actionable insights for both technical and non-technical stakeholders.
  • Highly motivated, pro-active and capable of multi-tasking under pressure in a fast-paced environment without compromising quality.
  • Commitment to fostering a blameless culture where failures are treated as opportunities to learn and improve systems.
  • Interest in financial markets and technology.
Preferred Qualifications
  • Bachelor’s degree in Computer Science, System Engineering, or a related technical field that involves programming.
  • 7 to 10 years of experience 

ABOUT GOLDMAN SACHS

The Goldman Sachs Group, Inc. is a leading global investment banking, securities and investment management firm that provides a wide range of financial services to a substantial and diversified client base that includes corporations, financial institutions, governments and individuals. Founded in 1869, the firm is headquartered in New York and maintains offices in all major financial centers around the world.

Skills Required

  • Strong proficiency in at least one major programming language, such as Java, Python, or Node.js
  • Experience with Infrastructure as Code frameworks such as Terraform, Ansible, or CloudFormation
  • Deep understanding of Docker and Kubernetes, including service meshes and ingress controllers
  • Advanced experience with AWS, GCP, or Azure and resilient cloud-native architectures
  • Proficiency with observability tools and practices, including tracing, logging, and metrics
  • Experience with automated testing and SDLC concepts
  • Experience developing applications in a Linux environment
  • Knowledge of algorithms, data structures, and software design
  • Knowledge of networking protocols and load balancing strategies in distributed systems
  • Bachelor’s degree in Computer Science, Systems Engineering, or a related technical field involving programming
  • 7 to 10 years of experience
  • Interest in financial markets and technology

Goldman Sachs Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Goldman Sachs and has not been reviewed or approved by Goldman Sachs.

  • Healthcare Strength Coverage includes medical, dental, vision, disability, life and accident insurance, with multiple plan options and most premiums subsidized; coverage often starts on day one. Wellness resources, on-site health centers in some locations, and EAP access reinforce the depth of health support.
  • Parental & Family Support Family care includes on-site childcare in some offices, expectant parent resources, and transitional programs for returning parents. Feedback suggests parental leave is very generous, with reports of around 20 weeks paid leave and stipends for adoption, surrogacy, and fertility-related services.
  • Retirement Support The firm provides a 401(k) plan with employer matching contributions and broad financial education to help employees plan for retirement. Resources also support saving for education and preparing for unexpected events.

Goldman Sachs Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
67,118 Employees

What We Do

At Goldman Sachs, we believe progress is everyone’s business. That’s why we commit our people, capital and ideas to help our clients, shareholders and the communities we serve to grow. Founded in 1869, Goldman Sachs is a leading global investment banking, securities and investment management firm. Headquartered in New York, we maintain offices in all major financial centers around the world. More about our company can be found at www.goldmansachs.com

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
8 Locations
40000 Employees
45K-85K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
11 Locations
40000 Employees
45K-85K Annually
Hybrid
Plano, TX, USA
289097 Employees

Cox Enterprises Logo Cox Enterprises

Client Solutions Manager - Major Accounts (Autotrader)

Artificial Intelligence • Automotive • Greentech • Information Technology • Machine Learning • Software • Cybersecurity
Hybrid
Fort Worth, TX, USA
30000 Employees
55K-153K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account