Principal, Staff Site Reliability

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
210K-230K Annually
Expert/Leader
Financial Services
The Role
Owns reliability, scalability, and developer experience for hybrid cloud and on-premises platforms in a regulated financial-services environment. Responsibilities include infrastructure architecture, Terraform and Ansible automation, CI/CD, observability, disaster recovery, incident command, production hardening, and AI-assisted incident response. The role mentors engineers, establishes reliability standards, leads postmortems, and drives improvements in uptime, MTTR, deployment velocity, and developer productivity.
Summary Generated by Built In

We are hiring a Staff SRE / Infrastructure Engineer to own the reliability, scalability, and developer experience of the platforms that power our investment, portfolio, and enterprise operations. You will set the technical direction for our hybrid cloud footprint (AWS + Azure +

on-prem), harden production for a regulated buy-side environment, and lead an emerging body of work applying AI to incident response and root-cause analysis. This is a hands-on senior IC role with material influence over architecture, tooling standards, and how engineers ship software.


What you'll do
  • Own end-to-end reliability for business-critical services — SLOs, error budgets, capacity planning, DR, and incident command — with a five-nines mindset appropriate to financial services workloads.
  • Design and evolve our multi-cloud and on-prem infrastructure across AWS, Azure, and colocated environments; drive workload placement, cost, and resilience trade-offs.
  • Build and maintain the Terraform, Ansible, and CI/CD backbone that lets product and data teams ship safely and quickly; codify golden paths and paved roads.
  • Advance observability (metrics, logs, traces, profiling) so on-call engineers can localize failures in minutes, not hours; instrument reliability as a first-class product surface.
  • Lead the buildout of AI-assisted SRE capabilities — LLM-driven triage, incident summarization, runbook synthesis, and automated RCA — with human-in-the-loop guardrails.
  • Partner with security, data platform, and application teams on hardening, patching, secrets, network segmentation, and change management appropriate to a regulated environment.
  • Participate in a leader-level on-call rotation; run blameless postmortems and drive systemic fixes to closure.
  • Mentor senior engineers; set the bar for infrastructure code review, production readiness reviews, and reliability practice across the org.

Required experience
  • 10+ years building and operating production infrastructure at scale, including hybrid cloud

+ on-prem.

  • Deep expertise in AWS and Azure compute, networking, IAM, and managed data services; comfortable with account/project topology, landing zones, and org-level guardrails.
  • Expert-level Linux, Terraform, Ansible, and modern CI/CD (GitHub Actions, GitLab CI, Argo, or equivalent).
  • Track record of measurable improvements in platform reliability, MTTR, deployment velocity, and developer productivity.
  • Strong scripting/software skills in Python and/or Go; can read and refactor application code well enough to debug across the stack.
  • Production experience with Kubernetes, service meshes, and container security posture.
  • Fluency with observability stacks (Datadog, Prometheus/Grafana, OpenTelemetry, ELK/Splunk).
  • Experience running incident command and driving durable postmortem outcomes.

Nice to have
  • Prior experience in financial services, buy-side, or another regulated environment (SOC 2, SOX, GLBA).
  • Hands-on work with AI/LLM tooling for SRE — agentic incident response, RAG over runbooks/telemetry, or automated RCA.
  • Experience with data-center automation, colocation, or modular infrastructure.
  • FinOps depth: unit economics, cost attribution, and budget governance across AWS/Azure.

The compensation range for this position is for a full-time employee in New York. The base salary offered will depend on qualifications, market data and internal equity.

Base Salary Range
$210,000—$230,000 USD

At DigitalBridge, we strive to create an inclusive environment where diverse employees want to work and where they can flourish professionally. In furtherance of our culture, all qualified applicants will receive consideration for employment without regard to race, national origin, gender, age, religion, disability, sexual orientation, veteran status, marital status or any other characteristics protected by law.


Skills Required

  • 10+ years building and operating production infrastructure at scale, including hybrid cloud and on-premises environments
  • Deep expertise in AWS and Azure compute, networking, IAM, managed data services, account or project topology, landing zones, and organizational guardrails
  • Expert-level Linux, Terraform, Ansible, and modern CI/CD
  • Track record of measurable improvements in platform reliability, MTTR, deployment velocity, and developer productivity
  • Strong scripting or software development skills in Python and/or Go, including the ability to debug application code across the stack
  • Production experience with Kubernetes, service meshes, and container security posture
  • Fluency with observability stacks such as Datadog, Prometheus/Grafana, OpenTelemetry, ELK, or Splunk
  • Experience running incident command and driving durable postmortem outcomes
  • Prior experience in financial services, buy-side, or another regulated environment
  • Hands-on experience with AI or LLM tooling for SRE, including agentic incident response, RAG, or automated RCA
  • Experience with data-center automation, colocation, or modular infrastructure
  • FinOps expertise in unit economics, cost attribution, and budget governance across AWS and Azure
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boca Raton, FL
237 Employees
Year Founded: 2013

What We Do

DigitalBridge (NYSE: DBRG) is a leading global digital infrastructure firm. With a heritage of over 25 years investing in and operating businesses across the digital ecosystem including cell towers, data centers, fiber, small cells, and edge infrastructure, the DigitalBridge team manages a $48 billion portfolio of digital infrastructure assets on behalf of its limited partners and shareholders. Headquartered in Boca Raton, DigitalBridge has key offices in New York, Los Angeles, London, and Singapore. For more information, visit: www.digitalbridge.com.

Similar Jobs

Alaffia Health Logo Alaffia Health

Operations Associate

Artificial Intelligence • Healthtech • Insurance • Machine Learning • Payments
Remote or Hybrid
United States
80 Employees
70K-80K Annually
In-Office or Remote
2 Locations
175633 Employees
100K-193K Annually

CDW Logo CDW

User Interface Designer

Information Technology
Remote or Hybrid
US
15100 Employees
105K-150K Annually

CDW Logo CDW

Senior Business Analyst

Information Technology
Remote or Hybrid
US
15100 Employees
88K-123K Annually

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account