Senior Site Reliability Engineer

Posted 13 Days Ago
Be an Early Applicant
Redwood City, CA, USA
Hybrid
180K-230K Annually
Senior level
Artificial Intelligence • Information Technology • Software • Automation
The Role
Own the reliability, scalability, and observability of production systems. Design and operate AWS infrastructure with Terraform and Kubernetes, build monitoring and SLOs, automate deployments and capacity management, ensure database and data pipeline reliability, partner on architecture reviews, and drive infrastructure security, compliance, and incident response.
Summary Generated by Built In

About Us

GridCARE is a leading venture-backed startup solving the most critical constraint in AI’s growth trajectory: immediate access to power. As demand for computing skyrockets, access to energy has become the defining bottleneck in the AI infrastructure race. While leading tech companies invest billions in speculative, long-term solutions that may take decades to arrive, GridCARE’s pioneering physics-based generative AI platform unlocks gigawatts of hidden capacity in today’s electric grid — enabling hyperscalers, data center developers, and utilities to power AI infrastructure years sooner than conventional approaches and without costly upgrades.

Founded at Stanford’s Doerr School of Sustainability and backed by leading climate-tech and deep-tech investors, GridCARE has assembled a world-class team spanning power systems, AI, and infrastructure.

At GridCARE, you will:

⚡ Work at the intersection of AI, energy, and infrastructure — the foundation of the next industrial revolution.

🤝 Partner with hyperscalers, developers, and utilities on high-impact, real-world deployments.

🌎 Help shape a more abundant, efficient, and resilient energy future for the digital era.

🚀 Join a company defining a new category — capacity acceleration for AI.

💰 Receive competitive compensation, equity, and benefits in a fast-growth, mission-driven environment.

Learn more about GridCARE:

  • TechCrunch: GridCARE thinks more than 100 GW of data-center capacity is hiding in the grid

  • Utility Dive: Portland General Electric invests in AI-powered flexibility to speed data-center connection

  • Data Center Dynamics: From Years to Months — Creating an AI Fast Lane for Data Centers

  • GridCARE Raises $64 Million Series A to Create a New Category: Power Acceleration

Job Description

We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the availability our customers (utilities, data center operators) require.

Responsibilities
  • Design and operate infrastructure on AWS using Terraform and Kubernetes

  • Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs

  • Automate away toil — deployment pipelines, capacity management, self-healing systems

  • Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship

  • Manage database and data pipeline reliability for large-scale, real-time grid data processing

  • Drive security and compliance best practices across infrastructure

Qualifications

Required

  • 5+ years in SRE, DevOps, or infrastructure engineering roles

  • Deep experience with Kubernetes, Terraform/IaC, and cloud platforms (AWS Preferred)

  • Strong scripting/programming ability (Python, Bash)

  • Observability Experience (Prometheus, Grafana, Datadog)

  • Track record of running on-call for production systems and leading incident response

  • Experience with CI/CD pipelines (Github Actions) and infrastructure automation

  • Experience with Gitops concepts and tooling (ArgoCD/Flux)

  • Solid understanding of networking, distributed systems, and database reliability

  • Comfortable operating in a fast-moving startup environment with ambiguity

Preferred

  • Experience with data-intensive or real-time processing systems

  • Background in energy, climate tech, or critical infrastructure

  • Experience scaling infrastructure through hypergrowth

  • On-Prem Kubernetes Deployment Experience

  • Windows Server Administration Experience

What We Offer
  • Competitive salary, performance bonus, and equity.

  • Comprehensive health, dental, and vision coverage.

  • Lunch provided three days a week in office.

  • Hybrid schedule for local employees: 3 days in office for collaboration, 2 days remote for focused work.

  • Access to leading academic, industry, and government partners in the AI-energy ecosystem.

  • A mission-driven team focused on shaping the future of the energy transition.

Salary Range

$180,000-$230,000 Total

Join us in tackling one of the most important infrastructure challenges of our time — enabling the energy foundation for the age of AI.

Skills Required

  • 5+ years of experience in SRE, DevOps, or infrastructure engineering roles
  • Deep experience with Kubernetes, Terraform or infrastructure as code, and cloud platforms; AWS preferred
  • Strong scripting or programming ability in Python and Bash
  • Observability experience with Prometheus, Grafana, Datadog, or similar tools
  • Experience running on-call for production systems and leading incident response
  • Experience with CI/CD pipelines, including GitHub Actions, and infrastructure automation
  • Solid understanding of networking, distributed systems, and database reliability
  • Comfort operating in a fast-moving startup environment with ambiguity
  • Experience with data-intensive or real-time processing systems
  • Background in energy, climate tech, or critical infrastructure
  • Experience scaling infrastructure through hypergrowth
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Redwood, California
22 Employees
Year Founded: 2024

What We Do

GridCARE is the essential one-stop power solution for AI data centers, delivering energy security and resilience at unprecedented speed. The company's AI-powered proprietary platform identifies and unlocks previously untapped grid capacity to accelerate AI infrastructure deployment. By creating the critical bridge between power availability and AI expansion, GridCARE is enabling the next generation of AI innovation.

Similar Jobs

Formation Bio Logo Formation Bio

Senior Site Reliability Engineer

Artificial Intelligence • Big Data • Healthtech • Biotech • Pharmaceutical
Easy Apply
Hybrid
3 Locations
150 Employees
186K-232K Annually
Remote or Hybrid
United States
1750 Employees

Zocdoc Logo Zocdoc

Senior Site Reliability Engineer

Healthtech • Information Technology • Software • Telehealth
Easy Apply
Remote or Hybrid
USA
900 Employees
180K-220K Annually

PwC Logo PwC

Site Reliability Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
58 Locations
370000 Employees
151K-187K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account