System Reliability Engineering Lead

Posted Yesterday
Be an Early Applicant
Hiring Remotely in USA
Remote
152K-228K Annually
Expert/Leader
Energy • Manufacturing • Solar • Renewable Energy
GE Vernova is accelerating the path to more reliable, affordable, and sustainable energy.
The Role
Leads system reliability engineering for a global grid software SaaS portfolio. Owns cloud infrastructure, platform standardization, SLOs, release governance, incident command, disaster recovery, FinOps, and capacity planning. Serves as final production deployment authority and customer-facing reliability lead. Builds a reliability enablement center, mentors a distributed SRE team, and drives automation, progressive delivery, observability, compliance, and operational excellence across critical utility applications.
Summary Generated by Built In
Job Description SummaryAs the Tech Lead for System Reliability Engineering within the GridOS SaaS Products organization, you will be the hands-on technical authority on production stability for our global grid software SaaS portfolio. You will bridge the gap between architectural design and real-world operations, driving a culture of high reliability and engineering excellence across a distributed team spanning three geographies. You are the "Gatekeeper" for production environments — owning the Change Management process, holding final authority to approve or halt deployments based on system health, and accountable for meeting SLA/SLO targets for critical infrastructure applications serving major North American utility customers.
This is a player-coach role. You will architect and build alongside your team while setting technical direction, mentoring engineers, and serving as the primary customer-facing SRE point of contact. You will own FinOps for the SaaS platform, driving cloud cost optimization and capacity planning as the customer base scales.

Job DescriptionDay 0 — Strategic Provisioning and Design

Standardized Cloud Infra Provisioning

Architect and implement standardized, secure cloud infrastructure provisioning. Drive extreme automation to reduce account provisioning timelines and accelerate customer onboarding to the SaaS platform.

The Golden Path

Define and build the standardized "Middle-Mile" software delivery platform (IDP) using Backstage, ArgoCD, and GitHub Actions. Eliminate bespoke deployment methodologies and establish a single, repeatable path to production.

Follow-the-Sun Architecture

Design and operate the global handover protocols and 24/7 operational coverage model across US, India, and Mexico time zones. Ensure seamless support continuity without graveyard shifts.

Reliability Targets

Establish and own enterprise-wide Service Level Objectives (SLOs) and Service Level Indicators (SLIs) aligned with critical user journeys for global utility customers. Define error budgets and enforce them.

Day 1 — Release Governance and Deployment

Final Approval Authority

Serve as the final technical authority for all production releases. Enforce rigorous change control and validate that all security and performance quality gates are met before any deployment proceeds.

Progressive Delivery

Implement and operate advanced deployment strategies including Canary and Blue/Green rollouts. Build and verify automated rollback capabilities. Hands-on with deployment tooling and pipeline configuration.

SRE Center for Enablement (C4E)

Build and mature the C4E to provide coaching, standardized templates, and repeatable reliability patterns that uplift practices across all product teams. Act as the go-to technical resource for reliability engineering across the organization.

Day 2 — Operational Excellence and Optimization

Incident Command

Serve as the Lead Incident Commander for high-severity (Sev1/Sev2) events. Lead the technical direction, communication, and containment efforts. Available for P1 escalations around the clock.

Blameless Culture

Own the post-incident lifecycle. Facilitate blameless Root Cause Analysis (RCA) to ensure systemic fixes replace recurring operational risks. Build a team culture where incidents drive improvement, not blame.

Business Continuity

Architect and validate end-to-end Backup and Disaster Recovery (DR) strategies, including cross-region failover and automated recovery testing. Hands-on with DR runbook development and execution.

FinOps and Capacity Planning

Own financial operations for the SaaS platform. Drive cloud cost optimization through reserved instances, right-sizing, and waste elimination. Perform long-term capacity planning based on customer growth trajectory and application scaling requirements.

Customer Engagement and Team Leadership

Customer-Facing Accountability

Serve as the primary SRE point of contact for North American utility customers. Own customer satisfaction and NPS for SaaS reliability. Participate in customer-facing reviews, incident communications, and service health reporting. Must meet customer-mandated background check requirements for access to critical infrastructure data and environments.

Player-Coach Team Leadership

Lead a distributed team of 8 SRE engineers across Hyderabad Technical Center and Querétaro, scaling with SaaS application and customer growth. Set technical direction, assign tasks, own team deliverables, and drive day-to-day execution. Mentor engineers on SRE practices, cloud architecture, and operational discipline. Provide performance feedback to the people leader of record. Foster a culture of high performance and continuous learning.



Required Qualifications

Technical Qualifications

  • Cloud Ecosystem: Deep expertise in AWS core services (EC2, EKS, RDS, S3, IAM) and management tools (CloudTrail, CloudWatch)
  • Orchestration: Advanced mastery of Kubernetes internals and EKS cluster operations across multi-region architectures
  • Continuous Delivery: Expert knowledge of ArgoCD, GitHub Actions, and GitOps-first workflows
  • Automation: Proficiency in Infrastructure as Code (IaC) using Terraform and configuration management via Ansible
  • Observability: Hands-on experience with Prometheus, Grafana, observability platforms (Splunk or Datadog), and OpenTelemetry standard to build comprehensive telemetry pipelines
  • FinOps: Demonstrated experience in cloud cost optimization, reserved instance management, right-sizing, and long-term capacity planning for multi-tenant SaaS platforms

Experience and Leadership

  • Overall Experience: 12+ years in software engineering, cloud operations, or infrastructure roles
  • Domain Depth: 8–10 years of hands-on experience in SRE, Platform Engineering, Cloud Operations, or Production Support for large-scale, distributed SaaS applications
  • Technical Leadership: Proven track record of leading distributed engineering teams as a player-coach — setting technical direction while remaining hands-on with architecture, automation, and incident response
  • Operational Discipline: Exceptional troubleshooting skills under pressure and a "Fire Marshal" mindset toward investigation and proactive inspection
  • Customer Engagement: Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
  • Background Check: Must be able to pass customer-mandated background screening for access to critical infrastructure environments


Desired Characteristics

Regulated Environments

  • Practical knowledge of NERC CIP compliance standards in a SaaS context
  • Experience with SOC2, ISO 27001, or IEC 62443 compliance frameworks
  • Familiarity with operating in highly regulated industries such as utilities, financial services, or critical national infrastructure

Certifications

  • AWS Certification: DevOps Engineer — Professional or Solutions Architect — Associate/Professional
  • CKA: Certified Kubernetes Administrator
  • SRE Practitioner Certification
  • AWS FinOps Practitioner or equivalent cloud financial management certification

Additional Information

About the SRE Team: The Grid Software SRE function is a newly established capability supporting the organization's SaaS transformation. The team currently supports Field Damage Assessment (FDA) and Distributed Dynamic Line Rating (DDLR) applications, with the portfolio expanding as GE Vernova Grid Software scales from its initial SaaS customers to a target of 20+ customers by end of 2027. The team operates a follow-the-sun model with engineers based in Hyderabad Technical Center (India) and Querétaro (Mexico).

Why US-Based: North American utility customers operating critical national infrastructure require that production environments and customer data be managed by US-based personnel who have completed customer-mandated background screening. This role exists to meet that requirement while providing hands-on technical leadership to the global SRE team.


Work Schedule

General shift, US business hours. On-call availability required for P1/Sev1 incidents. Follow-the-sun handoff protocols with Hyderabad and Querétaro teams.

Travel Requirements

Up to 10% — customer sites and team locations as needed (estimated 2–4 trips per year)


Additional Information

GE Vernova offers a great work environment, professional development, challenging careers, and competitive compensation. GE Vernova is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, national or ethnic origin, sex, sexual orientation, gender identity or expression, age, disability, protected veteran status or other characteristics protected by law.

GE Vernova will only employ those who are legally authorized to work in the United States for this opening. Any offer of employment is conditioned upon the successful completion of a drug screen (as applicable).

Relocation Assistance Provided: No

#LI-Remote - This is a remote position

For candidates applying to a U.S. based position, the pay range for this position is between $151,800.00 and $227,700.00. The Company pays a geographic differential of 110%, 120% or 130% of salary in certain areas. The specific pay offered may be influenced by a variety of factors, including the candidate’s experience, education, and skill set.

Bonus eligibility: discretionary annual bonus.

This posting is expected to remain open for at least seven days after it was posted on September 25, 2026.

Available benefits include medical, dental, vision, and prescription drug coverage; access to Health Coach from GE Vernova, a 24/7 nurse-based resource; and access to the Employee Assistance Program, providing 24/7 confidential assessment, counseling and referral services. Retirement benefits include the GE Vernova Retirement Savings Plan, a tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions, as well as access to Fidelity resources and financial planning consultants. Other benefits include tuition assistance, adoption assistance, paid parental leave, disability benefits, life insurance, 12 paid holidays, and permissive time off.

GE Vernova Inc. or its affiliates (collectively or individually, “GE Vernova”) sponsor certain employee benefit plans or programs GE Vernova reserves the right to terminate, amend, suspend, replace, or modify its benefit plans and programs at any time and for any reason, in its sole discretion. No individual has a vested right to any benefit under a GE Vernova welfare benefit plan or program. This document does not create a contract of employment with any individual.

Skills Required

  • 12+ years of experience in software engineering, cloud operations, or infrastructure roles
  • 8-10 years of hands-on experience in SRE, platform engineering, cloud operations, or production support for large-scale distributed SaaS applications
  • Deep expertise in AWS core services including EC2, EKS, RDS, S3, IAM, CloudTrail, and CloudWatch
  • Advanced Kubernetes and EKS cluster operations experience across multi-region architectures
  • Expert knowledge of ArgoCD, GitHub Actions, and GitOps workflows
  • Proficiency with Terraform and Ansible
  • Hands-on experience with Prometheus, Grafana, Splunk or Datadog, and OpenTelemetry
  • Experience with cloud cost optimization, reserved instances, right-sizing, and capacity planning for multi-tenant SaaS platforms
  • Experience leading distributed engineering teams as a player-coach
  • Exceptional troubleshooting and incident response skills
  • Experience working directly with enterprise customers on production reliability, incident communication, and service-level reporting
  • Ability to pass customer-mandated background screening for access to critical infrastructure environments
  • Practical knowledge of NERC CIP compliance standards
  • Experience with SOC 2, ISO 27001, or IEC 62443 compliance frameworks
  • Experience in utilities, financial services, or other highly regulated industries
  • AWS DevOps Engineer Professional or Solutions Architect certification
  • Certified Kubernetes Administrator certification
  • SRE Practitioner certification
  • AWS FinOps Practitioner or equivalent certification

GE Vernova Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about GE Vernova and has not been reviewed or approved by GE Vernova.

  • Retirement Support — The 401(k) plan includes company matching contributions and additional company retirement contributions, with access to Fidelity resources and financial planning consultants. Feedback suggests this structure supports long-term savings beyond a basic match.
  • Parental & Family Support — Paid parental leave is available with flexible, continuous or non-continuous usage, and is complemented by adoption resources and Work/Life Connections guidance. Maternity leave is described as extended relative to typical workplace norms.
  • Leave & Time Off Breadth — Time-off programs include 12 paid holidays, permissive time off for many salaried roles, and dedicated personal, illness, and caregiving time for U.S. new hires. Some hourly roles start with a defined PTO bank, while other roles may offer unlimited time off.

GE Vernova Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Cambridge, MA
75,000 Employees
Year Founded: 2024

What We Do

GE Vernova is a planned purpose-built company on a mission to electrify the planet while simultaneously working to decarbonize it. If we want our energy future to be different…we must be different. Our mission is embedded in our name. We retain our treasured legacy, “GE,” in our name as an enduring and hard-earned badge of quality and ingenuity. “Ver” / “verde” signal Earth’s verdant and lush ecosystems. “Nova,” from the Latin “novus,” nods to a new, innovative era of lower carbon energy that GE Vernova will help deliver. GE Vernova brings together GE’s portfolio of energy businesses including Power, Wind, Electrification and Digital businesses. With focus, GE Vernova is accelerating the path to more reliable, affordable, and sustainable energy, while helping our customers power economies and deliver the electricity that is vital to health, safety, security, and improved quality of life. Together, we have The Energy to Change the World.

Why Work With Us

Join our team, to evolve and grow, surrounded by some of the brightest minds in the industry who help you get better every day. You’ll get the chance to rewrite the rules, work on cutting-edge technology, and be part of a global team for positive change.

Gallery

Gallery

Similar Jobs

Comcast Logo Comcast

Sales Representative

Digital Media • Information Technology • News + Entertainment
Remote or Hybrid
Pennsylvania, USA
115000 Employees
26K-28K Hourly

Block Logo Block

Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
New York, NY, USA
12000 Employees
39K-233K Annually

Block Logo Block

Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
Pittsburgh, PA, USA
12000 Employees
129K-233K Annually

Block Logo Block

Mid-market Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
Chicago, IL, USA
12000 Employees
130K-234K Annually

Similar Companies Hiring

Fortune Brands Innovations Thumbnail
Manufacturing
Deerfield, IL
10000 Employees
Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees
Amalgamated Sugar Thumbnail
Food • Greentech • Agriculture • Industrial • Manufacturing
Boise, Idaho
768 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account