Senior Site Reliability Engineer NEX

Posted 2 Days Ago
Be an Early Applicant
Houston, TX, USA
In-Office
Senior level
Other • Energy
The Role
Design, build, and operate highly available, scalable systems on Google Cloud. Improve reliability via SLIs/SLOs, capacity planning, DR, and observability. Automate infrastructure with Terraform, improve CI/CD, run on-call rotations, lead incident response and postmortems, and partner with engineering and security teams to reduce toil and optimize cloud cost and performance.
Summary Generated by Built In

Reliability Engineering

● Design, implement, and operate scalable, resilient, and highly available systems on Google Cloud Platform. 

● Improve service availability, latency, performance, scalability, and operational resilience. 

● Define, implement, and track service-level indicators, service-level objectives, and error budgets. 

● Perform capacity planning, performance analysis, and workload forecasting. 

● Design and validate disaster recovery, backup, failover, and service-restoration capabilities. 

● Implement and maintain secure cloud networking, IAM, workload identities, service accounts, and access-control practices. 

● Partner with cybersecurity and identity teams to ensure infrastructure and services follow organizational security standards. 

● Monitor cloud consumption and optimize resource utilization, performance, and cost efficiency. 

● Identify operational risks and recommend improvements to cloud architecture and service design.


Automation and Platform Engineering

● Build and maintain cloud infrastructure using Terraform or comparable infrastructure-as-code tools. 

● Automate repetitive operational activities and systematically identify, measure, and reduce manual toil. 

● Build reusable infrastructure modules, deployment patterns, and operational tooling. 

● Improve CI/CD pipelines to enable secure, repeatable, and reliable software delivery.


Observability and Incident Management 

● Develop actionable alerts that identify meaningful service degradation while reducing alert fatigue and unnecessary operational noise. 

● Create and maintain dashboards, runbooks, operational procedures, and troubleshooting documentation. 

● Participate in a sustainable on-call rotation supporting production systems. 

● Respond to production incidents, coordinate service restoration, and lead incident response when appropriate. 

● Facilitate blameless postmortems and identify corrective and preventive actions. 

● Use incident and operational data to improve system design, automation, monitoring, and response processes.


Collaboration and Service Ownership 

● Partner with software engineering, data engineering, security, and product teams to improve application reliability and production operations. 

● Promote shared responsibility for production reliability between application development and platform teams. 

● Establish and document reliability standards, operational practices, and reusable engineering patterns. 

● Provide technical guidance and coaching on SRE, cloud, Kubernetes, observability, and incident-management practices.


Required Knowledge, Skills, and Abilities 

● Three or more years of experience in Site Reliability Engineering, platform engineering, DevOps, cloud engineering, production software engineering, or a similar role. 

● Experience operating highly available systems in a 24/7 production environment. 

● Hands-on experience operating workloads on Google Cloud Platform or another major public cloud platform. 

● Strong experience managing compute, networking and data GCP services workloads 

● Strong experience with containerization and orchestration technologies, including Docker and Kubernetes. 

● Experience building and managing infrastructure with Terraform or a comparable infrastructure-as-code tool. 

● Proficiency in Python, Go, Java, or another comparable programming language.

● Experience implementing or operating CI/CD pipelines using GitHub Actions,Azure DevOps, Bitbucket Pipelines, or comparable tools. 

● Experience implementing observability using metrics, logs, traces, dashboards, and alerts. 

● Experience participating in on-call rotations, responding to incidents, and contributing to postmortems. 

● Understanding of SLIs, SLOs, error budgets, and other SRE principles. 

● Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud-service layers. 

● Ability to communicate effectively with engineering teams, business stakeholders, and operational personnel.


Minimum Qualifications 

● Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience. 

● 3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, or DevOps. 

● 3+ years of experience operating production workloads in GCP. 

● Ability to understand and communicate in English at a level sufficient to issue, receive, and respond to safety-related and operations-related instructions. 


Preferred Qualifications 

● Google Cloud and/or Kubernetes certifications. 

● Experience supporting data-intensive, streaming, analytics, or event-driven platforms. 

● Experience establishing production-readiness, incident-management, or reliability-review processes. 

● Experience working in the energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries. 

● Experience supporting technology environments that integrate cloud platforms with remote sites, field equipment, industrial systems, or edge computing.



About Us
The Evolving Oil Field Demands Evolving Service Providers

NexTier is a leading provider of integrated completions that employs sustainable practices and equipment to support our customers’ ESG goals while accelerating production in the most demanding US land basins.

Patterson-UTI is committed to a workplace free from discrimination and harassment, offering equal employment opportunities to all individuals regardless of personal characteristics protected by law. Employees are encouraged to report any concerns through multiple channels.


Skills Required

  • 3+ years of experience in Site Reliability Engineering, platform engineering, cloud engineering, DevOps, or similar roles
  • 3+ years operating production workloads in Google Cloud Platform (GCP)
  • Experience operating highly available systems in a 24/7 production environment
  • Strong experience managing compute, networking, and data workloads on GCP
  • Hands-on experience with Docker and Kubernetes (containerization and orchestration)
  • Experience building and managing infrastructure with Terraform or comparable IaC tools
  • Proficiency in Python, Go, Java, or another comparable programming language
  • Experience implementing or operating CI/CD pipelines (GitHub Actions, Azure DevOps, Bitbucket Pipelines, or comparable)
  • Experience implementing observability using metrics, logs, traces, dashboards, and alerts
  • Experience participating in on-call rotations, incident response, and contributing to postmortems
  • Understanding of SLIs, SLOs, error budgets, and other SRE principles
  • Ability to troubleshoot complex issues across application, infrastructure, network, data, and cloud-service layers
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field, or equivalent practical experience
  • Ability to understand and communicate in English sufficient for safety and operations instructions
  • Google Cloud and/or Kubernetes certifications
  • Experience supporting data-intensive, streaming, analytics, or event-driven platforms
  • Experience establishing production-readiness, incident-management, or reliability-review processes
  • Experience working in energy, oil and gas, industrial, IoT, field operations, or other operationally critical industries
  • Experience integrating cloud platforms with remote sites, field equipment, industrial systems, or edge computing
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Houston, TX
1,900 Employees
Year Founded: 1978

What We Do

Patterson-UTI Energy subsidiaries provide onshore contract drilling and pressure pumping services. Patterson-UTI Energy, Inc. pushes the boundaries of innovation so you can embrace new possibilities. With expertise and scale in major operational areas, we provide a diverse network of drilling and pressure pumping services, directional drilling, rental equipment and technology to forge your path to success. Our oilfield solutions deliver results that lead your business into the next generation of oil and gas. With headquarters in Houston, Texas and regional offices throughout our operating areas, let’s team up to advance your business.

Similar Jobs

Samsara Logo Samsara

Senior Analyst, GTM Business Operations – Public Sector

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
92K-139K Annually

Ericsson Logo Ericsson

Operational Sourcing Intern

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Plano, TX, USA
88000 Employees

Ericsson Logo Ericsson

Artificial Intelligence Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Plano, TX, USA
88000 Employees

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
45K-85K Annually

Similar Companies Hiring

Runwise Thumbnail
Greentech • Hardware • Real Estate • Software • Energy • PropTech
New York, NY
199 Employees
Energy CX Thumbnail
Greentech • Professional Services • Business Intelligence • Consulting • Energy • Financial Services • Utilities
Chicago, IL
108 Employees
Rosendin Thumbnail
Other • Manufacturing
San Jose, CA
6219 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account