Lead Systems Operations Engineer

Posted 6 Hours Ago
Be an Early Applicant
Iselin, NJ, USA
Hybrid
119K-224K Annually
Senior level
Fintech • Financial Services
Wells Fargo: Tech-powered. Innovation-led. We're transforming financial services.
The Role
Leads systems operations and reliability initiatives across platform, application, and engineering teams. Designs scalable, highly available infrastructure, resolves complex production issues, drives automation and self-service, and improves observability, performance, resiliency, and operational processes. The role supports Red Hat and Kubernetes environments, cloud-native architectures, ITSM practices, and hybrid infrastructure. It also requires technical consultation, initiative leadership, cross-functional collaboration, and occasional on-call support.
Summary Generated by Built In
About this role:
Well Fargo is seeking a highly skilled and forward-thinking Lead Systems Operations Engineer to join our Quality & Test Engineering Operations team within CTO Platform Services team. This role is ideal for someone passionate about building scalable, resilient, and intelligent infrastructure solutions. You will play a key role in driving automation, reducing operational toil, and enabling self-service capabilities through cutting-edge technologies including Generative AI and Agent development.
In this role, you will:
  • Lead complex, broad impact initiatives including provision of high level systems consultation for the technology teams
  • Work as key participant in large scale planning of computer systems and network infrastructure for Systems Operations functional area
  • Review and analyze complex technical challenges, as well as escalated support issues related to core business solutions that require in depth evaluation of multiple factors, such as alternatives, enhancements, periodic systems reviews, or improvements to existing systems
  • Make decisions on technical changes and enhancements
  • Consult with engineering team on change design requiring solid understanding of technical process controls or standards that influence and drive new initiatives
  • Collaborate and consult with technical peers, colleagues, and mid to more experienced level managers to resolve systems support issues and achieve goals
  • Collaborate and partner across platform, application, and engineering teams.
  • Ability to manage multiple priorities in a fast-paced, high-impact production environment.
  • Consistent delivery of high-quality reliability outcomes within expected timelines.
  • High attention to detail, data-driven problem-solving, and operational rigor.
  • Prior project or initiative leadership experience is highly desirable.
Required Qualifications:
  • 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 3+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 3+ years of Proficiency in leveraging observability platforms such as BigPanda, ThousandEyes, Grafana, Prometheus, ELK, Splunk Observability, and AppDynamics to enhance service reliability and performance monitoring
  • 3+ years of experience in IT Service Management (ITSM), with a strong background in incident, problem, and change management processes.
  • 3+ years of experience working with Red Hat Enterprise Linux and Kubernetes, with a strong focus on Red Hat OpenShift Container Platform (OCP).
  • 3+ years of experience with Site Reliability Engineering and supporting production grade.
  • 3+ years of experience with cloud-native architectures, high-availability systems, Cloud & Container Technologies like GCP or Azure and familiarity with Kubernetes.
  • 3+ years of experience with Automation & Scripting: including developing and maintaining playbooks.
Desired Qualifications:
  • Strong hands-on experience applying SRE practices, including SLI/SLO definition, error budgets, and reliability metrics.
  • Proven experience troubleshooting and resolving large-scale, distributed production systems.
  • Hands-on experience with observability and monitoring tools such as Grafana, Splunk, Prometheus, Cribl, ThousandEyes, AppDynamics, or equivalent, including dashboards, alerting, logs, and metrics.
  • Strong scripting and automation skills using Python, Bash, and/or PowerShell to reduce operational toil.
  • Experience building automation or reliability tooling using APIs, Git-based workflows, and modern engineering practices.
  • Solid understanding of incident, problem, and change management in enterprise production environments.
  • Strong communication and influencing skills across engineering teams and senior leadership.
  • Experience with capacity management, performance engineering, and resiliency design (HA, fault tolerance, RTO/RPO).
  • Experience operating in hybrid environments (on-prem + cloud) with complex enterprise dependencies.
  • Familiarity with infrastructure automation / IaC tools such as Ansible or Terraform.
  • Ability to drive technical debt remediation for critical legacy platforms using structured backlogs.
  • Experience mentoring or leading senior engineers in reliability, operations, or SRE-focused roles.
  • Experience with Generative AI and Agent development technologies.
Role Expectations:
  • Some on-call nights and weekends may be required.
Pay Range
Reflected is the base pay range offered for this position. Pay may vary depending on factors including but not limited to demonstrated examples of prior performance, skills, experience, or work location. Employees may also be eligible for incentive opportunities.
$119,000.00 - $224,000.00
Benefits
Wells Fargo provides eligible employees with a comprehensive set of benefits, many of which are listed below. Visit Benefits - Wells Fargo Jobs for an overview of the following benefit plans and programs offered to employees.
  • Health benefits
  • 401(k) Plan
  • Paid time off
  • Disability benefits
  • Life insurance, critical illness insurance, and accident insurance
  • Parental leave
  • Critical caregiving leave
  • Discounts and savings
  • Commuter benefits
  • Tuition reimbursement
  • Scholarships for dependent children
  • Adoption reimbursement
Posting End Date:
3 Sep 2026
* Job posting may come down early due to volume of applicants.
We Value Equal Opportunity
Wells Fargo is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic.
Employees support our focus on building strong customer relationships balanced with a strong risk mitigating and compliance-driven culture which firmly establishes those disciplines as critical to the success of our customers and company. They are accountable for execution of all applicable risk programs (Credit, Market, Financial Crimes, Operational, Regulatory Compliance), which includes effectively following and adhering to applicable Wells Fargo policies and procedures, appropriately fulfilling risk and compliance obligations, timely and effective escalation and remediation of issues, and making sound risk decisions. There is emphasis on proactive monitoring, governance, risk identification and escalation, as well as making sound risk decisions commensurate with the business unit's risk appetite and all risk and compliance program requirements.
Applicants with Disabilities
To request a medical accommodation during the application or interview process, visit Disability Inclusion at Wells Fargo .
Drug and Alcohol Policy
Wells Fargo maintains a drug free workplace. Please see our Drug and Alcohol Policy to learn more.
Wells Fargo Recruitment and Hiring Requirements:
a. Third-Party recordings are prohibited unless authorized by Wells Fargo.
b. Wells Fargo requires you to directly represent your own experiences during the recruiting and hiring process.
#DNP-IND

Skills Required

  • 5+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through work experience, training, military experience, or education
  • 3+ years of Systems Engineering, Technology Architecture experience, or equivalent
  • 3+ years of experience leveraging observability platforms such as BigPanda, ThousandEyes, Grafana, Prometheus, ELK, Splunk Observability, and AppDynamics
  • 3+ years of experience in IT Service Management, including incident, problem, and change management
  • 3+ years of experience with Red Hat Enterprise Linux and Kubernetes, particularly Red Hat OpenShift Container Platform
  • 3+ years of experience with Site Reliability Engineering and supporting production-grade systems
  • 3+ years of experience with cloud-native architectures, high-availability systems, GCP or Azure, and Kubernetes
  • 3+ years of experience in automation and scripting, including developing and maintaining playbooks
  • Experience applying SRE practices, including SLI/SLO definition, error budgets, and reliability metrics
  • Experience troubleshooting and resolving large-scale distributed production systems
  • Experience with observability and monitoring tools, dashboards, alerting, logs, and metrics
  • Strong scripting and automation skills using Python, Bash, and/or PowerShell
  • Experience building automation or reliability tooling using APIs, Git-based workflows, and modern engineering practices
  • Understanding of incident, problem, and change management in enterprise production environments
  • Experience with capacity management, performance engineering, and resiliency design, including HA, fault tolerance, and RTO/RPO
  • Experience operating in hybrid on-premises and cloud environments with complex enterprise dependencies
  • Familiarity with infrastructure automation and IaC tools such as Ansible or Terraform
  • Ability to drive technical debt remediation for critical legacy platforms using structured backlogs
  • Experience mentoring or leading senior engineers in reliability, operations, or SRE-focused roles
  • Experience with Generative AI and Agent development technologies
  • Strong communication and influencing skills across engineering teams and senior leadership
  • Prior project or initiative leadership experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
205,000 Employees
Year Founded: 1852

What We Do

Wells Fargo & Company (NYSE: WFC) is a leading financial services company that has approximately $2.2 trillion in assets. We provide a diversified set of banking, investment and mortgage products and services, as well as consumer and commercial finance, through our four reportable operating segments: Consumer Banking and Lending, Commercial Banking, Corporate and Investment Banking, and Wealth & Investment Management. Wells Fargo ranked No. 33 on Fortune’s 2025 rankings of America’s largest corporations. Our technology professionals drive innovation, information security, and big data analytics while maintaining a network that handles more than 12 billion customer interactions a year. Join us! Are you looking for more? Find it here. At Wells Fargo, we're more than a financial services leader – we’re a global trailblazer committed to driving innovation, empowering communities, and helping our customers succeed. We believe that a meaningful career is much more than just a job – it’s about finding all of the elements to help you thrive, in one place. Living the Well Life means you’re supported in life, not just work. It means having robust benefits, competitive compensation, and programs designed to help you find work-life balance and well-being. You’ll be rewarded for investing in your community, celebrated for being your authentic self, and empowered to grow. And we’re recognized for it — Wells Fargo continues to rank on the LinkedIn Top Companies lists of best workplaces “to grow your career.” All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, status as a protected veteran, or any other legally protected characteristic. © 2026 Wells Fargo Bank, N.A. All rights reserved. Member FDIC.

Why Work With Us

We're known for our “Well Life” approach to supporting employees’ career aspirations, work-life balance, and mental and physical health. Wells Fargo continues to rank on the LinkedIn Top Companies lists of best workplaces “to grow your career.”

Gallery

Gallery
Gallery
Gallery
Gallery

Wells Fargo Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: 3 days a week
HQSan Francisco, CA
Bangalore, Bangalore
Belfast, GB
Bengaluru, Karnataka
Chandler, AZ
Charlotte, NC
Technology Center
Hyderabad, Telangana
Irving, TX
New York, NY
New York, NY
Phoenix, AZ
Learn more

Similar Jobs

Hybrid
Iselin, NJ, USA
205000 Employees
Hybrid
Hackensack, NJ, USA
205000 Employees
23-31 Hourly
Hybrid
Rutherford, NJ, USA
205000 Employees
23-31 Hourly
Hybrid
Newark, NJ, USA
205000 Employees
23-31 Hourly

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account