Lead Systems Operations Engineer

Posted 15 Hours Ago
Be an Early Applicant
Irving, TX, USA
Hybrid
Senior level
Fintech • Financial Services
Wells Fargo: Tech-powered. Innovation-led. We're transforming financial services.
The Role
Leads site reliability and systems operations for critical consumer-facing platforms. Establishes SLOs, SLIs, and error budgets; improves resilience, observability, automation, monitoring, and recovery. Leads major incident response, root-cause analysis, operational readiness, vendor dependency management, and reliability reporting. Provides technical mentorship and influences architecture, platform engineering, CI/CD, and Infrastructure-as-Code practices across cross-functional teams.
Summary Generated by Built In
About this role:
Wells Fargo is seeking a Lead Site Reliability Engineer (SRE) / Lead Systems Operations Engineer within the Consumer Technology (CT) organization. This role will provide technical leadership for operational excellence, platform reliability, resiliency, observability, and support readiness across critical consumer-facing applications and platforms.
The Lead SRE will serve as a senior technical leader responsible for driving reliability engineering practices, reducing operational risk, improving service availability, and enabling scalable platform operations. This role will partner closely with Application Development, Platform Engineering, Infrastructure teams, Shared Services, and External Vendors to ensure highly resilient, supportable, and observable solutions.
The ideal candidate combines deep technical expertise with strong operational leadership and will play a critical role in advancing Site Reliability Engineering practices across the organization.
In this role, you will support:
Reliability Engineering & Platform Stability
  • Lead reliability initiatives across critical business platforms and customer journeys.
  • Establish and drive Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budget practices.
  • Improve platform resilience through automation, self-healing capabilities, capacity planning, and fault-tolerant designs.
  • Identify and eliminate single points of failure across applications, infrastructure, vendor integrations, and customer flows.
  • Champion engineering solutions that improve availability, scalability, recoverability, and operational maturity.

Incident Management & Operational Excellence
  • Serve as a technical lead during major production incidents, providing coordination, technical direction, and recovery leadership.
  • Drive improvements in Mean Time to Detect (MTTD), Mean Time to Diagnose (MTTDiag), and Mean Time to Recover (MTTR).
  • Lead Root Cause Analysis (RCA) efforts and ensure corrective actions are implemented and tracked to completion.
  • Identify recurring operational patterns and develop preventive solutions to reduce production incidents.
  • Develop and maintain incident playbooks, recovery procedures, and operational readiness standards.

Observability & Monitoring
  • Lead enterprise observability initiatives leveraging Splunk, Grafana, GCP Monitoring, AppDynamics, and related platforms.
  • Define monitoring standards, alerting strategies, dashboards, and customer journey observability solutions.
  • Partner with application and infrastructure teams to improve telemetry, logging, tracing, and synthetic monitoring capabilities.
  • Develop actionable operational metrics and executive-level reliability reporting.

Automation & Engineering Excellence
  • Drive automation strategies that reduce manual effort and improve operational consistency.
  • Design and implement self-service operational capabilities and automated recovery solutions.
  • Utilize AI-assisted tools and engineering practices to improve incident detection, diagnosis, and remediation workflows.
  • Promote Infrastructure-as-Code (IaC), CI/CD best practices, and platform engineering principles.

Vendor & Dependency Management
  • Partner with internal and external service providers to improve reliability, support responsiveness, and recovery performance.
  • Evaluate vendor operational performance and contribute to service improvement initiatives.
  • Establish and monitor operational readiness expectations for critical vendor dependencies.
  • Drive resilience planning and support strategies for third-party integrations.

Technical Leadership
  • Provide technical leadership and mentorship to SREs, Systems Operations Engineers, and Platform Support Engineers.
  • Lead technical reviews, operational readiness assessments, and production support governance activities.
  • Influence architecture decisions to ensure supportability, resiliency, observability, and operational sustainability.
  • Collaborate with engineering leaders to establish and mature Site Reliability Engineering practices across Consumer Technology.
Required Qualifications:
  • 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 5+ years of Site Reliability Engineering, Platform Engineering, Production Support, or equivalent experience demonstrated through work experience, military experience, training, or education.
  • 5+ years supporting mission-critical production applications in large enterprise environments.
  • 3+ years leading major incident management, operational support, or reliability engineering initiatives.
  • 3+ years of experience with observability and monitoring platforms such as Splunk, Grafana, AppDynamics, Dynatrace, GCP Monitoring, or similar technologies.
  • 2+ years of experience driving automation, operational improvements, and reliability initiatives.
  • 3+ years of experience supporting distributed systems, cloud-based platforms, infrastructure, networking, and application architectures.
  • 1+ year of experience supporting highly regulated or customer-facing financial services platforms
Desired Qualifications:
  • Consumer Technology, Credit Card, Payments, Lending, or Digital Banking experience.
  • Experience implementing Site Reliability Engineering (SRE) principles, SLOs, SLIs, and Error Budgets.
  • Experience with DevOps, CI/CD, Infrastructure-as-Code, and cloud-native architectures.
  • Experience with AI-assisted engineering, incident management automation, or observability platforms.
  • Strong executive communication and stakeholder management skills.
  • Experience leading cross-functional technical teams without direct authority.
  • Experience supporting vendor governance and third-party operational readiness initiatives.
Job Expectations:
  • Relocation assistance is not provided for this position
  • Visa sponsorship is not available for this position
  • Position requires onsite presence at one of the posted Wells Fargo locations.
Locations:
  • 401 W. Las Collinas Blvd, Irving, Texas
  • 300 S. Brevard St. Charlotte, North Carolina
  • 2600 S. Price Rd. Chandler, Arizona
Posting End Date:
10 Sep 2026
*Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements.
Candidates applying to job openings posted in Canada: Applications for employment are encouraged from all qualified candidates, including women, persons with disabilities, aboriginal peoples and visible minorities. Accommodation for applicants with disabilities is available upon request in connection with the recruitment process.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo .
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.

Skills Required

  • 5+ years of Systems Engineering, Technology Architecture, or equivalent experience
  • 5+ years of Site Reliability Engineering, Platform Engineering, Production Support, or equivalent experience
  • 5+ years supporting mission-critical production applications in large enterprise environments
  • 3+ years leading major incident management, operational support, or reliability engineering initiatives
  • 3+ years of experience with observability and monitoring platforms such as Splunk, Grafana, AppDynamics, Dynatrace, GCP Monitoring, or similar technologies
  • 2+ years driving automation, operational improvements, and reliability initiatives
  • 3+ years supporting distributed systems, cloud-based platforms, infrastructure, networking, and application architectures
  • 1+ year supporting highly regulated or customer-facing financial services platforms
  • Consumer Technology, Credit Card, Payments, Lending, or Digital Banking experience
  • Experience implementing SRE principles, SLOs, SLIs, and error budgets
  • Experience with DevOps, CI/CD, Infrastructure-as-Code, and cloud-native architectures
  • Experience with AI-assisted engineering, incident management automation, or observability platforms
  • Strong executive communication and stakeholder management skills
  • Experience leading cross-functional technical teams without direct authority
  • Experience supporting vendor governance and third-party operational readiness initiatives

Wells Fargo Compensation & Benefits Highlights

  • Healthcare Strength Health coverage begins on day one with comprehensive medical, dental, and vision options, and the company subsidizes a substantial share of premiums for U.S. employees (varying by compensation band).
  • Retirement Support A robust 401(k) program includes an employer match for eligible employees, with specifics laid out in plan materials and filings.
  • Parental & Family Support Paid parental leave extends up to 16 weeks for eligible primary caregivers, alongside fertility coverage, adoption/surrogacy reimbursement, and lactation support.

Wells Fargo Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
205,000 Employees
Year Founded: 1852

What We Do

Wells Fargo & Company (NYSE: WFC) is a leading financial services company that has approximately $2.2 trillion in assets. We provide a diversified set of banking, investment and mortgage products and services, as well as consumer and commercial finance, through our four reportable operating segments: Consumer Banking and Lending, Commercial Banking, Corporate and Investment Banking, and Wealth & Investment Management. Wells Fargo ranked No. 33 on Fortune’s 2025 rankings of America’s largest corporations. Our technology professionals drive innovation, information security, and big data analytics while maintaining a network that handles more than 12 billion customer interactions a year. Join us! Are you looking for more? Find it here. At Wells Fargo, we're more than a financial services leader – we’re a global trailblazer committed to driving innovation, empowering communities, and helping our customers succeed. We believe that a meaningful career is much more than just a job – it’s about finding all of the elements to help you thrive, in one place. Living the Well Life means you’re supported in life, not just work. It means having robust benefits, competitive compensation, and programs designed to help you find work-life balance and well-being. You’ll be rewarded for investing in your community, celebrated for being your authentic self, and empowered to grow. And we’re recognized for it — Wells Fargo continues to rank on the LinkedIn Top Companies lists of best workplaces “to grow your career.” All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic. © 2026 Wells Fargo Bank, N.A. All rights reserved. Member FDIC.

Why Work With Us

We're known for our “Well Life” approach to supporting employees’ career aspirations, work-life balance, and mental and physical health. Wells Fargo continues to rank on the LinkedIn Top Companies lists of best workplaces “to grow your career.”

Gallery

Gallery
Gallery
Gallery
Gallery

Wells Fargo Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: 3 days a week
HQSan Francisco, CA
Bangalore, Bangalore
Belfast, GB
Bengaluru, Karnataka
Chandler, AZ
Charlotte, NC
Technology Center
Hyderabad, Telangana
Irving, TX
New York, NY
New York, NY
Phoenix, AZ
Learn more

Similar Jobs

Wells Fargo Logo Wells Fargo

Lead Systems Operations Engineer

Fintech • Financial Services
Hybrid
Irving, TX, USA
205000 Employees
Hybrid
San Antonio, TX, USA
205000 Employees
Hybrid
Pearland, TX, USA
205000 Employees
Hybrid
San Antonio, TX, USA
205000 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account