Lead SRE / Platform Engineer - APP/PROD Support

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office or Remote
Expert/Leader
Internet of Things • Mobile • Retail
The Role
Lead SRE/Platform Engineer responsible for platform reliability, incident escalation, mentoring Tier 2/3 engineers, driving CI/CD and automation, owning observability and streaming stack governance, cloud governance, capacity and DR planning, and platform tooling/upgrade roadmaps.
Summary Generated by Built In
Lead SRE / Platform Engineer - APP/PROD SupportJob Summary / About the RoleAT&T is looking for a Lead SRE / Platform Engineer to own the technical direction, reliability roadmap, and platform governance for enterprise integration, messaging, and event-driven ecosystems. This is the senior-most IC role in the support organization: the candidate is the final technical escalation point for high-severity incidents, mentors Tier 2 and Tier 3 engineers, and drives cross-team architecture and reliability decisions that shape how the platform evolves.Key Responsibilities
  • Own end-to-end platform reliability strategy for availability, resilience, latency, and operational efficiency across the support organization.
  • Serve as the final technical escalation and decision-making authority for high-severity/complex incidents, driving resolution and post-incident reliability improvements.
  • Provide technical leadership and mentoring to Tier 2 and Tier 3 engineers, including skill development, code/config review, and troubleshooting coaching.
  • Drive DevOps and automation strategy including Golden Image improvements and support automation use cases across teams.
  • Define and enforce GitHub Actions pipelines and CI/CD reliability standards platform-wide.
  • Lead JFROG Helm chart automation and JFROG images/ACR migration initiatives.
  • Own microservices deployment enablement strategy and platform/tooling upgrade roadmap.
  • Own and govern monitoring, alerting, observability, and logging stack architecture:
Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.
  • Own health-check framework strategy including Airflow health-check requirements.
  • Partner with architecture and delivery leadership on reliability, scalability, and platform evolution decisions.
  • Own cloud infrastructure governance: creation, maintenance, access controls, and policy enforcement.
  • Own capacity planning, DR planning/exercises, and platform best-practice documentation sign-off.
  • Own cost management governance, role enforcement, and license management decisions.
  • Set and maintain SOP standards for alerts and incident patterns across Tier 1-3.
Required Qualifications / Must-Have Skills
  • 10+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles, including demonstrated technical leadership.
  • Proven experience mentoring or technically leading engineers, with the ability to set standards and review the work of others.
  • Deep hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations at scale.
  • Advanced experience with CI/CD engineering and GitHub Actions, including defining organization-wide standards.
  • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks, including architecture-level ownership.
  • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.
  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).
  • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).
  • Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.
  • Experience owning governance controls: access management, role enforcement, and separation of duties.
  • Proven track record leading high-severity incident response and driving post-incident reliability improvement programs.
  • Strong stakeholder communication skills to represent platform reliability decisions to architecture and leadership audiences.
Good-to-Have / Nice-to-Have
  • Postgres performance and reliability operations.
  • Telecom-scale high-availability systems experience.
  • Prior people-management or formal team-lead experience.
What We Offer
  • Highest IC-level ownership over platform reliability strategy and standards.
  • Direct influence on architecture, governance, and cross-team technical decisions.
  • Mentorship scope across Tier 2 and Tier 3 engineering teams.
  • Enterprise-scale impact across observability, automation, and resilience engineering.

Weekly Hours:

40

Time Type:

Regular

Location:

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City

It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

Skills Required

  • 10+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles with technical leadership
  • Proven experience mentoring or technically leading engineers and setting standards
  • Deep hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS)
  • Advanced experience with CI/CD engineering and GitHub Actions, defining organization-wide standards
  • Deep observability experience with Prometheus, Grafana, AlertManager and logging stacks including architecture-level ownership
  • Strong Python automation scripting skills for reliability engineering and operational toil reduction
  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis and troubleshooting
  • Familiarity with Java, React, and Spring Boot services for production troubleshooting
  • Strong hands-on experience with Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS MSK, and Apache Flink
  • Experience owning governance controls: access management, role enforcement, separation of duties
  • Proven track record leading high-severity incident response and driving post-incident reliability improvements
  • Strong stakeholder communication skills to represent platform reliability decisions to architecture and leadership
  • Postgres performance and reliability operations
  • Telecom-scale high-availability systems experience
  • Prior people-management or formal team-lead experience

AT&T Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about AT&T and has not been reviewed or approved by AT&T.

  • Healthcare Strength Health coverage spans medical, dental, vision, and mental health services, plus a personal healthcare team, wellness apps, and supplemental options such as fertility care, cancer support, doula services, and wigs for chemotherapy. These comprehensive offerings are portrayed as supporting a wide range of employee needs.
  • Leave & Time Off Breadth Paid time off includes vacation, holidays, sick days, caregiver time, parental leave, and adoption assistance, with some roles reaching about 23 days of PTO after several years. Community volunteer days and flexible time off options add further support for work-life balance.
  • Wellbeing & Lifestyle Benefits Employees receive sizable service discounts like 50% off most wireless plans and broadband, along with savings on travel, event tickets, and insurance. Additional workplace perks such as hybrid work models and relocation assistance contribute to overall value.

AT&T Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Dallas, TX
150,000 Employees

What We Do

Bring us your biggest career aspirations. Share your boldest dreams. This is a moment to get energized. Through 5G and Fiber, AT&T provides connectivity that leads to smarter homes, safter communities, higher quality health care and more life-changing innovations. With AT&T, Connecting Changes Everything.

Gallery

Gallery

Similar Jobs

Capco Logo Capco

Product Manager

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

CSC Logo CSC

Security Engineer

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Capco Logo Capco

Visualisation Analyst (Liquidity & Treasury)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

CrowdStrike Logo CrowdStrike

Infrastructure Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account