Key Responsibilities
- Provide 24x7 support for incidents, alerts, and operational issues with a strong bias towards independent identification and resolution.
- Triage and diagnose incidents using logs, monitoring dashboards, and platform knowledge; resolve issues directly without defaulting to escalation.
- Engage Tier 2 only when incidents involve architectural complexity, infrastructure-level failures, or changes beyond Tier 1 resolution authority.
- Support client-submitted iTrack incident tickets and maintain end-to-end ticket ownership including resolution and closure.
- Respond to PagerDuty and automated alerts; validate, investigate, and remediate before escalating.
- Monitor production and non-production environment health; proactively identify anomalies and take corrective action.
- Monitor application support mailboxes and manage operational follow-through.
- Provide C2W emergency support as the first responder; independently handle and resolve wherever possible.
- Share client profile data and usage reports on request.
- Maintain clear incident communication to clients even when Tier 2 is engaged.
- Track and report operational metrics including MTTR and ticket resolution trends.
- Develop and maintain Tier 1 SOPs and operational runbooks based on real resolution patterns.
Required Qualifications / Must-Have Skills
- 3+ years of experience in application production support with a strong track record of independently diagnosing and resolving incidents.
- Solid working knowledge of the full technology stack in scope including event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling.
- Hands-on experience with incident lifecycle management in ticketing systems (iTrack or equivalent), including root cause identification and resolution documentation.
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and issue resolution, not just monitoring.
- Hands-on operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics, restarts, and health investigation.
- Practical working knowledge of Confluent Kafka and Azure Event Hub: consumer lag analysis, topic health checks, and message flow troubleshooting.
- Solid SQL/Postgres skills for data-level investigation and validation during incidents.
- Working ability to read and interpret Java, Spring Boot, and React application logs for issue identification.
- Basic Python scripting capability for operational checks and quick-fix automation.
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics.
- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination.
- Willingness to work in rotational 24x7 shifts.
Good-to-Have / Nice-to-Have
- Awareness of hybrid streaming ecosystems including Confluent Cloud, AWS-MSK, and Apache Flink.
- Exposure to IBM Sterling Integrator integration flows for context during incident triage.
- Telecom or high-availability enterprise support experience.
- Experience with CI/CD-driven deployment pipelines in a support context.
Location / Work Mode
Onsite (Hyderabad / Bangalore or designated AT&T location)
Weekly Hours:
40Time Type:
RegularLocation:
IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge CityAT&T and its subsidiaries are committed to equal employment opportunity. All hiring, promotion, and other employment decisions remain merit-based and free from discrimination on the basis of race, color, religion, religious creed, national origin, ancestry, age, sex, sexual orientation, gender, gender identity, gender expression, physical disability, mental disability, pregnancy, medical condition, genetic information, marital status, citizenship status, military status, veteran status, or any other characteristic protected by federal, state, or local laws. In addition, AT&T will provide reasonable accommodations to qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made. Click here to learn more or request an application accommodation here.
Skills Required
- 3+ years of experience in application production support with independent incident diagnosis and resolution
- Working knowledge of event streaming platforms, integration middleware, AKS-hosted microservices, and observability tooling
- Hands-on incident lifecycle management in ticketing systems (iTrack or equivalent) including root cause documentation
- Proficient with Splunk, PagerDuty, Prometheus, and Grafana for active troubleshooting and resolution
- Operational experience with Kubernetes, especially Azure Kubernetes Service (AKS), including pod-level diagnostics
- Practical working knowledge of Confluent Kafka and Azure Event Hub (consumer lag analysis, topic health checks, message flow troubleshooting)
- Solid SQL/Postgres skills for data-level investigation and validation during incidents
- Ability to read and interpret Java, Spring Boot, and React application logs for issue identification
- Basic Python scripting capability for operational checks and quick-fix automation
- Good Linux/Unix command-line skills for real-time log analysis and system diagnostics
- Strong written and verbal communication skills for incident updates, resolution documentation, and client coordination
- Willingness to work in rotational 24x7 shifts (on-site)
- Awareness of Confluent Cloud, AWS-MSK, and Apache Flink
- Exposure to IBM Sterling Integrator integration flows
- Telecom or high-availability enterprise support experience
- Experience with CI/CD-driven deployment pipelines in a support context
AT&T Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about AT&T and has not been reviewed or approved by AT&T.
-
Healthcare Strength — Health coverage spans medical, dental, vision, and mental health services, plus a personal healthcare team, wellness apps, and supplemental options such as fertility care, cancer support, doula services, and wigs for chemotherapy. These comprehensive offerings are portrayed as supporting a wide range of employee needs.
-
Leave & Time Off Breadth — Paid time off includes vacation, holidays, sick days, caregiver time, parental leave, and adoption assistance, with some roles reaching about 23 days of PTO after several years. Community volunteer days and flexible time off options add further support for work-life balance.
-
Wellbeing & Lifestyle Benefits — Employees receive sizable service discounts like 50% off most wireless plans and broadband, along with savings on travel, event tickets, and insurance. Additional workplace perks such as hybrid work models and relocation assistance contribute to overall value.
AT&T Insights
What We Do
Bring us your biggest career aspirations. Share your boldest dreams. This is a moment to get energized. Through 5G and Fiber, AT&T provides connectivity that leads to smarter homes, safter communities, higher quality health care and more life-changing innovations. With AT&T, Connecting Changes Everything.
Gallery









