Staff Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
3 Locations
In-Office
121K-151K Annually
Senior level
Fintech • Payments
The Role
Leads large-scale site reliability engineering strategy, architecting highly available and scalable systems, improving observability, automation, incident response, capacity planning, performance, and cloud costs. Builds self-healing mechanisms and AI agents that automate operational workflows, reduce TOIL, and support incident response and anomaly detection. Establishes AI security and governance controls, advises engineering leadership, leads cross-functional reliability initiatives, and mentors engineers developing production-grade SRE and agentic solutions.
Summary Generated by Built In

About the Role

We are looking for a highly motivated, high-potential Staff Site Reliability Engineer (SRE) to join our team as a technical leader and drive transformative impact across WEX’s platform reliability and operational excellence.

This is a particularly exciting time to be part of the SRE function at WEX. Our diverse product ecosystem supports a wide array of customer businesses and generates rich, complex telemetry across applications, infrastructure, and platforms. Ensuring these systems are scalable, observable, and resilient is critical to unlocking business value and customer success.

As a Staff SRE, you will play a pivotal role in shaping the reliability engineering strategy at WEX. You’ll architect and lead efforts that improve availability, performance, and efficiency at scale, driving initiatives across observability, automation, incident management, problem management, capacity planning, and performance optimization. You’ll be hands-on in building foundational tooling and frameworks while also acting as a multiplier, mentoring engineers, aligning cross-functional teams, and influencing platform decisions with a strong reliability lens.

You’ll also help define how WEX applies AI to reliability engineering, building agents and reusable skills that automate high-TOIL work, integrating safely into our AI ecosystem, and establishing security and operational guardrails so intelligent automation is trustworthy, measurable, and scalable. Our team embraces agile development, a strong product mindset, and modern engineering practices, including AI-assisted operations and intelligent automation.
 

You’ll take on some of the most complex, high-impact challenges at WEX, supported by a team of highly skilled engineers and technical leaders invested in your success and growth.

If you’re a senior technical leader passionate about building reliable systems, leading through influence, and making a meaningful impact with AI-enabled operations, this is a fantastic opportunity for you.
 

What You’ll Do
  • Architect and oversee the implementation of mission-critical systems with a focus on availability, scalability, and operational excellence.

  • Define and enforce SRE best practices and operational standards across engineering and platform teams.

  • Lead cross-functional initiatives to enhance system reliability, performance, and efficiency at scale.

  • Serve as a technical advisor for engineering leadership on reliability, architecture, and operational risk.

  • Develop capacity planning and load testing strategies that proactively identify and mitigate scalability risks.

  • Design self-healing and auto-recovery mechanisms that reduce manual intervention during failures.

  • Drive cloud cost optimization and budgeting initiatives without compromising reliability.

  • Design, build, and govern AI agents and reusable skills that automate operational workflows and reduce TOIL.

  • Evaluate and integrate AI ecosystems, including models, agent frameworks, orchestration, tooling interfaces, and evaluation practices, into SRE and platform workflows.

  • Apply AI security and governance controls, including least-privilege tool access, secure data and prompt handling, auditability, and safe automation boundaries.

  • Lead AI-enabled initiatives for incident response, runbook automation, anomaly detection, and capacity/performance insights, with clear measurement of TOIL reduction and reliability outcomes.

  • Mentor engineers on production-grade agentic solutions and help embed AI into day-to-day reliability practices.
     

What You’ll Bring
  • 8+ years of experience with a focus on large-scale system reliability.

  • Expertise in system architecture, cloud platforms, and automation frameworks.

  • Deep knowledge of Kubernetes, service meshes, and distributed tracing.

  • Experience with monitoring and logging platforms (Grafana, ELK stack, Splunk, etc.).

  • Knowledge of containerization and orchestration (Docker, Kubernetes).

  • Experience designing high-availability, fault-tolerant architectures.

  • Strong understanding of database reliability engineering (MySQL, PostgreSQL, NoSQL), plus networking, databases, and storage architectures.

  • Excellent incident command and crisis management skills.

  • Hands-on experience building AI agents and skills/tools that integrate with operational systems (APIs, observability, ticketing, CI/CD).

  • Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context/memory, evaluation, and human-in-the-loop patterns.

  • Practical understanding of AI security and governance for production use, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions.

  • Demonstrated ability to reduce TOIL with AI by automating repetitive operational work and delivering measurable efficiency and reliability gains.
     

Nice to Have

  • Experience with multi-region and multi-cloud deployments.

  • Deep expertise in scalable microservices and event-driven architectures.

  • Strong experience with advanced observability tools (OpenTelemetry, Jaeger, Prometheus).

  • Leadership in driving large-scale SRE transformations.

  • Experience designing and developing AI agents, skills, and copilots for SRE/platform engineering, including evaluation and safe rollout practices.

  • Familiarity with enterprise agent platforms, skill registries, and observability for AI/agent workflows.

  • Ability to influence engineering culture and process improvements, including adoption of AI-assisted operations under change control, safety, and audit requirements.

The base pay range represents the anticipated low and high end of the pay range for this position. Actual pay rates will vary and will be based on various factors, such as your qualifications, skills, competencies, and proficiency for the role. Base pay is one component of WEX's total compensation package. Most sales positions are eligible for commission under the terms of an applicable plan. Non-sales roles are typically eligible for a quarterly or annual bonus based on their role and applicable plan. WEX's comprehensive and market competitive benefits are designed to support your personal and professional well-being. Benefits include health, dental and vision insurances, retirement savings plan, paid time off, health savings account, flexible spending accounts, life insurance, disability insurance, tuition reimbursement, and more. For more information, check out the "About Us" section.Pay Range: $120,600.00 - $150,900.00

Skills Required

  • 8+ years of experience focused on large-scale system reliability
  • Expertise in system architecture, cloud platforms, and automation frameworks
  • Deep knowledge of Kubernetes, service meshes, and distributed tracing
  • Experience with monitoring and logging platforms such as Grafana, ELK Stack, or Splunk
  • Knowledge of containerization and orchestration using Docker and Kubernetes
  • Experience designing high-availability and fault-tolerant architectures
  • Strong understanding of database reliability engineering, MySQL, PostgreSQL, NoSQL, networking, databases, and storage architectures
  • Excellent incident command and crisis management skills
  • Hands-on experience building AI agents and tools integrating with APIs, observability, ticketing, and CI/CD systems
  • Working knowledge of AI ecosystems and agent architectures, including orchestration, tool calling, context, memory, evaluation, and human-in-the-loop patterns
  • Practical understanding of AI security and governance, secure permissions, data leakage prevention, secrets handling, and guarded autonomous actions
  • Demonstrated ability to reduce TOIL with AI and deliver measurable efficiency and reliability gains
  • Experience with multi-region and multi-cloud deployments
  • Deep expertise in scalable microservices and event-driven architectures
  • Advanced observability experience with OpenTelemetry, Jaeger, and Prometheus
  • Leadership in driving large-scale SRE transformations
  • Experience designing and developing AI agents, skills, and copilots for SRE or platform engineering
  • Familiarity with enterprise agent platforms, skill registries, and observability for AI or agent workflows
  • Ability to influence engineering culture and process improvements involving AI-assisted operations, change control, safety, and audit requirements

WEX Inc. Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about WEX Inc. and has not been reviewed or approved by WEX Inc..

  • Leave & Time Off Breadth Leave offerings are portrayed as a standout, with generous PTO and additional paid time for volunteering. Time-off flexibility is also positioned as a meaningful part of the overall rewards experience.
  • Retirement Support Retirement benefits are presented as strong, including a 401(k) match that is described as competitive. This element appears to materially strengthen the total rewards package even when cash compensation feels less compelling.
  • Strong & Reliable Incentives Variable compensation is sometimes framed positively through bonuses and uncapped earning potential in sales-oriented roles. Stock options are also cited as an additional reward component that can improve perceived total compensation.

WEX Inc. Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Portland, ME
4,900 Employees
Year Founded: 1983

What We Do

We simplify complex payment systems for fleets, corporate payments, and healthcare—unlocking insights, opportunities, and efficiencies to give you greater control of your business. Powered by the belief that complex payment systems can be made simple, WEX (NYSE: WEX) is a leading financial technology service provider across a wide spectrum of sectors, including fleet, travel and healthcare. WEX operates in more than 10 countries and in more than 20 currencies through approximately 4,900 associates around the world. WEX fleet cards offer approximately 14 million vehicles exceptional payment security and control; our travel and corporate solutions business processes over $35 billion of purchase volume annually; and the WEX Health financial technology platform helps 343,000 employers and more than 28 million consumers better manage healthcare expenses.

Similar Jobs

Core Scientific Logo Core Scientific

Site Reliability Engineer

Blockchain • Fintech • Cryptocurrency
In-Office
Austin, TX, USA
290 Employees

Ping Identity Logo Ping Identity

Site Reliability Engineer

Cloud • Security • Software
Remote or Hybrid
USA
2300 Employees
170K-227K Annually

Domino Data Lab Logo Domino Data Lab

Site Reliability Engineer

Artificial Intelligence • Machine Learning
Remote or Hybrid
US
200 Employees
200K-230K Annually

AlphaSense Logo AlphaSense

Site Reliability Engineer

Artificial Intelligence • Fintech • Machine Learning • Natural Language Processing • Business Intelligence
Remote or Hybrid
United States
2000 Employees
150K-225K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account