Site Reliability Engineer – Observability & Automation

Reposted Yesterday
Be an Early Applicant
Santiago, Metropolitana de Santiago, CHL
In-Office
Mid level
Information Technology
The Role
Design, build, and maintain automation and IaC to improve reliability, observability, and incident response for cloud environments. Build CI/CD pipelines, monitor SLOs/SLIs, support cloud infrastructure, perform root cause analysis, and collaborate with development teams to increase system availability and scalability.
Summary Generated by Built In

Job Description:

About DXC

DXC helps global companies run their mission-critical systems and operations while modernizing IT, optimizing data architectures and ensuring security and scalability across public, private and hybrid clouds. The world's largest companies and public sector organizations trust DXC to deploy services to drive new levels of performance, competitiveness, and customer experience across their IT estates.

Our more than 125,000 people in 70-plus countries are entrusted by our customers to deliver transformative technologies to ensure the success, safety and well-being of businesses and people around the world. By combining strengths and expertise globally, we create solutions and deliver greater outcomes for customers across their entire IT estate. Learn more about how we deliver excellence for our customers and colleagues at DXC.com.

About the team

Our Site Reliability Engineering team ensures the reliability, scalability, performance, and availability of critical systems through deep observability, automation, and continuous engineering improvement. We act as the eyes and the hands of production: we instrument and monitor everything that matters, and we automate everything that shouldn't be done by hand. The team partners with development, operations, cloud, infrastructure, networking, and security stakeholders to reduce toil, increase system visibility, accelerate incident response, and keep highly available cloud and on-prem environments healthy.

About the Role

As a Site Reliability Engineer – Observability & Automation, you will be responsible for designing end-to-end observability across our platforms and building the automation that keeps them stable and self-healing. You will own monitoring, logging, tracing, and alerting strategies using tools such as Dynatrace, the Elastic Stack, Zabbix, and Grafana, and you will turn manual operational work into repeatable, auditable automation with Ansible and AWX. You'll partner closely with development and infrastructure teams to make systems more transparent, more reliable, and less dependent on human intervention.

Essential Job Functions

  • Design, implement, and maintain end-to-end observability (metrics, logs, traces, and alerting) across cloud and on-prem systems.
  • Build and tune dashboards and visualizations in Grafana, correlating data from multiple sources for fast diagnosis.
  • Implement APM, full-stack monitoring, and root-cause analysis with Dynatrace, including service-level and infrastructure-level visibility.
  • Design and operate centralized log management and analytics with the Elastic Stack (Elasticsearch, Kibana, Logstash/Beats), including parsing, indexing, and retention strategies.
  • Configure and maintain infrastructure and network monitoring with Zabbix, including templates, triggers, and intelligent alerting.
  • Develop automation with Ansible and AWX/Tower: playbooks, roles, inventories, workflows, and scheduled jobs to eliminate repetitive operational tasks.
  • Build automated incident detection and self-remediation flows that reduce mean-time-to-detect (MTTD) and mean-time-to-resolve (MTTR).
  • Write scripts and tooling (Python, Bash, PowerShell) to integrate systems, enrich monitoring data, and extend automation.
  • Define, instrument, and track SLOs, SLIs, and SLAs, and reduce alert noise by improving signal quality.
  • Collaborate with development teams to improve application reliability, performance, and instrumentation, contributing code and reviewing changes when needed.
  • Support cloud and network infrastructure to ensure high availability, capacity, and performance across environments.
  • Participate in incident management, root cause analysis, and blameless post-mortems, turning findings into automation and monitoring improvements.

Basic Qualifications

  • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related field, or an equivalent combination of education and experience.
  • 2-3 years of experience in Site Reliability Engineering, DevOps, Monitoring/Observability, Cloud, or related roles.
  • Hands-on experience with observability and monitoring platforms, ideally including several of: Dynatrace, Elasticsearch / Elastic Stack, Zabbix, and Grafana.
  • Proven experience building automation with Ansible (playbooks, roles, inventories) and orchestrating it through AWX or Ansible Tower/AAP.
  • Strong scripting and programming skills in Python, Bash, PowerShell, or similar, with the ability to integrate APIs and build operational tooling.
  • Software development knowledge (version control with Git, code review practices, and comfort reading/contributing to application code).
  • Experience working with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Networking fundamentals: TCP/IP, DNS, HTTP/S, load balancing, firewalls, and basic troubleshooting of connectivity and latency issues.
  • Solid understanding of system reliability practices, distributed systems, and SLO/SLI/SLA concepts.
  • Advanced English level required for global collaboration.

Other Qualifications (Preferred / Nice to Have)

  • Experience designing centralized logging and alerting at scale (log pipelines, retention, cost/performance tuning).
  • Experience with containerization and orchestration (Docker, Kubernetes) and how to observe and automate them.
  • Familiarity with Infrastructure as Code (Terraform, CloudFormation) complementing Ansible-based configuration management.
  • Experience with CI/CD tooling (GitLab CI, GitHub Actions, Jenkins, or similar).
  • Experience implementing automated incident response and self-healing systems.
  • Exposure to security and compliance in cloud environments.
  • Certifications in Dynatrace, Elastic, Red Hat (Ansible/RHCE), or a cloud provider (AWS/Azure/GCP) are a plus.
  • Strong analytical thinking, troubleshooting, and a proactive mindset focused on automation, efficiency, and continuous improvement.
  • Ability to work under pressure and handle critical incidents.

Why DXC? / Life at DXC

DXC is an employer of choice with strong values, and fosters a culture of inclusion, belonging and corporate citizenship. We inspire and take care of our people. We work to create a culture of learning, diversity and inclusion and are dedicated to strong ethics and corporate citizenship. DXC is where brilliant people seize opportunities to advance their careers and amplify customer success.

Full-time hires are eligible to participate in the DXC benefit program.  DXC offers a comprehensive, flexible, and competitive benefits program which includes, but is not limited to, health, dental, and vision insurance coverage; employee wellness; life and disability insurance; a retirement savings plan, paid holidays, paid time off.

At DXC Technology, we believe strong connections and community are key to our success. Our work model prioritizes in-person collaboration while offering flexibility to support wellbeing, productivity, individual work styles, and life circumstances. We’re committed to fostering an inclusive environment where everyone can thrive.

Recruitment fraud is a scheme in which fictitious job opportunities are offered to job seekers typically through online services, such as false websites, or through unsolicited emails claiming to be from the company. These emails may request recipients to provide personal information or to make payments as part of their illegitimate recruiting process. DXC does not make offers of employment via social media networks and DXC never asks for any money or payments from applicants at any point in the recruitment process, nor ask a job seeker to purchase IT or other equipment on our behalf. More information on employment scams is available here.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, IT, or equivalent experience
  • 3-7+ years of experience in Site Reliability Engineering, DevOps, or Cloud Engineering
  • Proven experience building and managing automation solutions in production environments
  • Experience with cloud platforms (AWS, Azure, Google Cloud Platform)
  • Programming or scripting skills using Python, Bash, or Go
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation
  • Experience with CI/CD tools such as Jenkins, GitLab CI, or GitHub Actions
  • Experience with monitoring and observability tools (Datadog, Prometheus, Splunk, New Relic)
  • Understanding of system reliability practices and distributed systems
  • Experience with containerization technologies such as Docker and Kubernetes
  • Advanced English level for global collaboration
  • Experience with Kubernetes and cloud-native architectures
  • Knowledge of DevOps and SRE best practices
  • Experience implementing automated incident response and remediation
  • Exposure to security and compliance in cloud environments
  • Cloud platform certification (AWS, Azure, or GCP)
  • Strong analytical thinking, troubleshooting, and problem-solving skills
  • Ability to work under pressure and handle critical incidents

DXC Technology Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about DXC Technology and has not been reviewed or approved by DXC Technology.

  • Healthcare Strength Health coverage includes multiple national carrier options and plan types, with HSA eligibility where applicable. Feedback suggests the medical, dental, and vision lineup is broad and comparable to large-firm offerings.
  • Retirement Support A 401(k) program with employer matching and an annual true-up is available, with standard vesting provisions. This structure can help employees capture matching contributions over the year if contribution rates vary.
  • Leave & Time Off Breadth Flexible or “unlimited” vacation is offered for many U.S. roles instead of accrual-based PTO. Feedback suggests the approach can support work-life balance when team norms allow adequate time away.

DXC Technology Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Ashburn, VA
86,261 Employees
Year Founded: 2017

What We Do

DXC Technology is a Fortune 500 global IT services leader. Our more than 130,000 people in 70-plus countries are entrusted by our customers to deliver what matters most. We use the power of technology to deliver mission critical IT services across the Enterprise Technology Stack to drive business impact. DXC is an employer of choice with strong values, and fosters a culture of inclusion, belonging and corporate citizenship. We are DXC.

Similar Jobs

Mastercard Logo Mastercard

Technical Project Manager

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Santiago, Metropolitana de Santiago, CHL
38800 Employees
Remote or Hybrid
2 Locations
289097 Employees

JPMorganChase Logo JPMorganChase

Controller

Financial Services
Remote or Hybrid
2 Locations
289097 Employees

Mastercard Logo Mastercard

Consultant

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Santiago, Metropolitana de Santiago, CHL
38800 Employees

Similar Companies Hiring

Scrunch  Thumbnail
Artificial Intelligence • Information Technology • Marketing Tech • Software • SEO
Salt Lake City, Utah
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account