CloudOps Engineer IV

Posted 4 Days Ago
Be an Early Applicant
2 Locations
In-Office
Expert/Leader
Software
The Role
Build and operate multi-cloud infrastructure across AWS and Azure, maintain Kubernetes clusters and Helm deployments, and develop GitOps, CI/CD, and infrastructure-as-code automation. Improve reliability through observability, SLI/SLO tracking, incident response, capacity planning, cost optimization, and resilience initiatives. Participate in on-call rotations, postmortems, documentation, and mentoring while supporting high-traffic payments, loyalty, and fuel-pricing platforms.
Summary Generated by Built In
PDI Technologies is looking for a CloudOps Engineer IV to join the SRE organization supporting Paylo, PDI’s payments, loyalty, and fuel-pricing product suite. This is a senior, hands-on individual-contributor role focused on keeping high-traffic, customer- and partner-facing platforms reliable, secure, and running efficiently across multi-cloud infrastructure.
 
 You will bring strong, hands-on experience across AWS, Azure, Kubernetes, Helm, Argo CD, Terraform/OpenTofu, Jenkins, and Datadog, and apply it directly — building and operating infrastructure, improving deployment pipelines, and strengthening observability. You will work under the direction of the Senior Manager, SRE, executing against the team’s reliability and infrastructure roadmap while bringing your own judgment and technical leadership to the problems in front of you.

Job Responsibilities:

    Cloud Infrastructure & Operations: 
     
  • Build, operate, and troubleshoot infrastructure across AWS and Azure in support of production workloads. 
  • Operate and maintain Kubernetes clusters, including deploying and maintaining Helm charts for the services you support. 
  • Participate in on-call rotation, respond to incidents, and drive them to resolution within your area of ownership.
  • Contribute to capacity planning, cost optimization, and resilience improvements for the systems you support.
  •  Automation & Continuous Delivery: 
     
  • Build and maintain GitOps-based deployment pipelines using Argo CD/Argo Workflows, including rollout and promotion configuration across environments. 
  • Write and maintain Infrastructure-as-Code (Terraform, OpenTofu) for the infrastructure you own, following team module standards. 
  • Build and maintain CI/CD pipelines in Jenkins, improving build/deploy automation and reliability. 
  • Support progressive delivery practices (blue-green/canary, automated rollback) for the services you support.
  •  
    Reliability & Observability:
     
  • Build and maintain Datadog dashboards, monitors, and alerts for the services you support, tuning alert thresholds to reduce noise. 
  • Contribute to defining SLIs/SLOs for your services and help track them over time. 
  • Participate in postmortems for incidents you're involved in, and follow through on assigned remediation items.
  • Collaboration & Mentorship:
     
  • Partner with engineers across the SRE team and with product engineering teams to troubleshoot issues and improve system design. 
  • Share knowledge with and mentor less-experienced engineers on the team (CloudOps Engineer I III) on cloud infrastructure, Kubernetes, and CI/CD practices. 
  • Contribute to documentation, runbooks, and onboarding materials for the systems you support. 

Required Qualifications:

  • 10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles.
  • Hands-on experience with AWS — you can build, troubleshoot, and operate cloud infrastructure directly. 
  • Hands-on experience with Kubernetes and Helm — deploying, operating, and troubleshooting workloads in production clusters. 
  • Hands-on experience with Argo CD/Argo Workflows for GitOps-based continuous delivery. 
  • Hands-on experience with Infrastructure as Code (Terraform, OpenTofu). 
  • Hands-on experience with Jenkins and Rancher for CI/CD pipeline development and maintenance. 
  • Hands-on experience with Datadog (or equivalent observability platform), including building dashboards, monitors, and alerts. 
  • Experience participating in an on-call rotation and responding to production incidents. 
  • Strong communication skills and the ability to work effectively across teams. 

Preferred Qualifications:

  • Experience supporting payments, fuel/retail, or loyalty platforms, or other systems with PCI DSS or similar compliance obligations.
  • Relevant certifications such as CKA/CKAD, AWS Certified Solutions Architect – Associate, Microsoft Certified: Azure Administrator, or HashiCorp Terraform Associate. 
  • Experience with messaging systems (Kafka/SQS/SNS) and multi-region/multi-AZ resilience patterns. 
  • Prior experience mentoring junior engineers or leading small technical initiatives. 

Behavioral Competencies:

  • Cultivates Innovation
  • Decision Quality
  • Manages Complexity
  • Drives Results
  • Business Insight 

Skills Required

  • 10+ years of experience in Cloud Operations, Site Reliability Engineering, DevOps, or Infrastructure Engineering roles
  • Hands-on experience building, troubleshooting, and operating AWS cloud infrastructure
  • Hands-on experience deploying, operating, and troubleshooting Kubernetes workloads in production clusters
  • Hands-on experience with Helm
  • Hands-on experience with Argo CD or Argo Workflows for GitOps-based continuous delivery
  • Hands-on experience with Terraform or OpenTofu
  • Hands-on experience developing and maintaining Jenkins and Rancher CI/CD pipelines
  • Hands-on experience with Datadog or an equivalent observability platform, including dashboards, monitors, and alerts
  • Experience participating in on-call rotations and responding to production incidents
  • Strong communication skills and ability to work effectively across teams
  • Experience supporting payments, fuel or retail, loyalty platforms, or systems with PCI DSS or similar compliance obligations
  • CKA/CKAD, AWS Certified Solutions Architect Associate, Microsoft Certified Azure Administrator, or HashiCorp Terraform Associate certification
  • Experience with Kafka, SQS, SNS, and multi-region or multi-AZ resilience patterns
  • Experience mentoring junior engineers or leading small technical initiatives
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Alpharetta, GA
1,905 Employees

What We Do

PDI Technologies resides at the intersection of productivity and sales growth, delivering powerful solutions that serve as the backbone of the convenience retail and petroleum wholesale ecosystem. By “Connecting Convenience” across the globe, we empower businesses to increase productivity, make more informed decisions, and engage faster with their customers. www.pditechnologies.com

Similar Jobs

Hybrid
Chennai, Tamil Nadu, IND
175633 Employees

TransUnion Logo TransUnion

Technical Product Manager

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
4 Locations
13000 Employees

TransUnion Logo TransUnion

Technical Product Manager

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
4 Locations
13000 Employees
Hybrid
Chennai, Tamil Nadu, IND
175633 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account