Senior App/Prod Support (Tier 3 Site Reliability Engineer (SRE) / Platform Engineer)

Posted Yesterday
Be an Early Applicant
3 Locations
In-Office or Remote
Senior level
Internet of Things • Mobile • Retail
The Role
Lead platform reliability and DevOps automation: implement CI/CD with GitHub Actions, automate JFrog/Helm and image migrations, enable microservices deployments, and operate observability and logging stacks. Provide Tier 3 troubleshooting and incident leadership, manage cloud infrastructure governance, capacity and DR planning, cost/license governance, and maintain SOPs and reliability best practices.
Summary Generated by Built In

Key Responsibilities

- Own platform reliability practices for availability, resilience, latency, and operational efficiency.

- Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases.

- Implement and maintain GitHub Actions pipelines and CI/CD reliability standards.

- Lead JFROG Helm chart automation and JFROG images/ACR migration work.

- Support microservices deployment enablement and platform/tooling upgrades.

- Own and optimize monitoring, alerting, observability, and logging stack components:

  Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools.

- Support health-check frameworks including Airflow health-check requirements.

- Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents.

- Collaborate with architecture and delivery teams on reliability and scalability patterns.

- Lead cloud infrastructure creation, maintenance, governance, and access controls.

- Drive capacity planning, DR planning/exercises, and platform best-practice documentation.

- Support cost management, role enforcement, and license management governance.

- Maintain SOP documentation for established alerts and incident patterns.

Required Qualifications / Must-Have Skills

- 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles.

- Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations.

- Advanced experience with CI/CD engineering and GitHub Actions.

- Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks.

- Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction.

- End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required).

- Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role).

- Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink.

- Experience in governance controls: access management, role enforcement, and separation of duties.

- Proven high-severity incident leadership and post-incident reliability improvement execution.

Good-to-Have / Nice-to-Have

- Postgres performance and reliability operations.

- Telecom-scale high-availability systems experience.

Experience Level

Senior to Lead IC (typically 10 to 17 years)

Location / Work Mode

Onsite (Hyderabad / Bangalore or designated AT&T location)

What We Offer

- Opportunity to define and scale platform reliability standards.

- High technical ownership and strong cross-functional influence.

- Enterprise-scale impact across observability, automation, and resilience engineering.

Weekly Hours:

40

Time Type:

Regular

Location:

IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Whitefield Road:Intl Tech Park, Navigator Bldg

It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

Skills Required

  • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles
  • Hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations
  • Advanced experience with CI/CD engineering and GitHub Actions
  • Experience with JFrog, Helm chart automation, and image/ACR migration
  • Deep observability experience with Prometheus, AlertManager, Grafana and logging stacks (OpenSearch, FluentBit, Azure Monitor, Thanos)
  • Strong Python automation scripting skills for reliability engineering and operational toil reduction
  • End-user proficiency with AI-assisted productivity and operations tools for incident analysis and documentation
  • Familiarity with Java, React, and Spring Boot services for production troubleshooting
  • Hands-on experience with streaming stack: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS MSK, Apache Flink
  • Experience in governance controls: access management, role enforcement, separation of duties
  • Proven high-severity incident leadership and post-incident reliability improvement execution
  • Postgres performance and reliability operations
  • Telecom-scale high-availability systems experience

AT&T Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about AT&T and has not been reviewed or approved by AT&T.

  • Healthcare Strength Health coverage spans medical, dental, vision, and mental health services, plus a personal healthcare team, wellness apps, and supplemental options such as fertility care, cancer support, doula services, and wigs for chemotherapy. These comprehensive offerings are portrayed as supporting a wide range of employee needs.
  • Leave & Time Off Breadth Paid time off includes vacation, holidays, sick days, caregiver time, parental leave, and adoption assistance, with some roles reaching about 23 days of PTO after several years. Community volunteer days and flexible time off options add further support for work-life balance.
  • Wellbeing & Lifestyle Benefits Employees receive sizable service discounts like 50% off most wireless plans and broadband, along with savings on travel, event tickets, and insurance. Additional workplace perks such as hybrid work models and relocation assistance contribute to overall value.

AT&T Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Dallas, TX
150,000 Employees

What We Do

Bring us your biggest career aspirations. Share your boldest dreams. This is a moment to get energized. Through 5G and Fiber, AT&T provides connectivity that leads to smarter homes, safter communities, higher quality health care and more life-changing innovations. With AT&T, Connecting Changes Everything.

Gallery

Gallery

Similar Jobs

Deepgram Logo Deepgram

Research Staff, LLMs

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
49 Locations
150 Employees
150K-250K Annually

Cloudflare Logo Cloudflare

Account Executive

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
India
4400 Employees

Coinbase Logo Coinbase

Country Director, India

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
India
4700 Employees
15M-15M Annually

Coinbase Logo Coinbase

Engineering Manager

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
India
4700 Employees
9M-9M Annually

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account