IT Security Engineering Associate Director

Posted Yesterday
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Expert/Leader
Financial Services
The Role
Lead IAM reliability engineering, observability, automation, resilience, and operational readiness across enterprise identity platforms. Responsibilities include designing monitoring and telemetry strategies, developing automation and testing frameworks, improving high availability and disaster recovery, conducting incident response and root cause analysis, and establishing SRE standards. The role partners with engineering, security, infrastructure, and architecture teams to improve IAM platform performance, scalability, recoverability, and service health.
Summary Generated by Built In
Are you ready to make an impact at DTCC?

Do you want to work on innovative projects, collaborate with a dynamic and supportive team, and receive investment in your professional development? At DTCC, we are at the forefront of innovation in the financial markets. We are committed to helping our employees grow and succeed. We believe that you have the skills and drive to make a real impact. We foster a thriving internal community and are committed to creating a workplace that looks like the world that we serve.

The Information Technology group delivers secure, reliable technology solutions that enable DTCC to be the trusted infrastructure of the global capital markets. The team delivers high-quality information through activities that include development of essential, building infrastructure capabilities to meet client needs and implementing data standards and governance.

Pay and Benefits:
 

  • Competitive compensation, including base pay and annual incentive
  • Comprehensive health and life insurance and well-being benefits, based on location
  • Pension / Retirement benefits
  • Paid Time Off and Personal/Family Care, and other leaves of absence when needed to support your physical, financial, and emotional well-being.
  • DTCC offers a flexible/hybrid model of 3 days onsite and 2 days remote (onsite Tuesdays, Wednesdays and a third day unique to each team or employee).
Position Summary

We are seeking an experienced Associate Director, IAM Site Reliability Engineering (SRE), Observability & Automation to lead reliability engineering, operational architecture, observability, service automation, and platform resiliency initiatives across the enterprise Identity and Access Management (IAM) ecosystem.

This role is responsible for improving the availability, performance, scalability, recoverability, and operational readiness of critical IAM platforms, including Privileged Access Management (PAM), Active Directory, PKI, Secrets Management, Authentication Services, Cloud Identity Platforms, and Identity Security services.

The ideal candidate combines deep expertise in IAM operations, site reliability engineering, observability, scripting, and test automation to drive proactive monitoring, service health management, operational automation, and continuous improvement across the IAM landscape.

The Associate Director will work closely with Engineering, Security, Infrastructure, and Architecture teams to ensure IAM services are designed, instrumented, tested, and operated for maximum reliability and resilience.

Key Responsibilities

IAM Reliability Engineering & Operations

  • Lead reliability engineering initiatives across IAM platforms and services.
  • Drive platform availability, resiliency, service health, and operational readiness improvements.
  • Establish and implement SRE best practices and operational excellence standards.
  • Partner with engineering teams to embed reliability, monitoring, and automation requirements throughout the development lifecycle.
  • Perform production readiness reviews and operational risk assessments.

Observability & Monitoring Engineering

  • Design and implement monitoring and observability strategies for IAM platforms.
  • Develop standards for: 
    • Infrastructure Monitoring
    • Application Monitoring
    • User Experience Monitoring
    • Transaction Monitoring
    • Dependency Monitoring
    • Security Event Monitoring
    • Cloud Service Monitoring
  • Establish logging, metrics, tracing, telemetry, dashboards, and alerting standards.
  • Build and manage centralized observability solutions leveraging platforms such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, and Application Insights.
  • Develop service health dashboards and executive operational reporting.

Reliability Architecture & Platform Engineering

  • Create operational architecture diagrams, service dependency maps, data flow diagrams, and resiliency models.
  • Review platform designs and identify scalability, performance, availability, and resiliency risks.
  • Define and implement standards for: 
    • High Availability (HA)
    • Disaster Recovery (DR)
    • Failover Design
    • Capacity Planning
    • Fault Tolerance
    • Service Recovery
  • Collaborate with architecture and engineering teams to eliminate single points of failure.

Automation Engineering & Testing

  • Develop automation frameworks to improve operational efficiency and platform reliability.
  • Create scripts and tooling using PowerShell, Python, Bash, or similar technologies for monitoring, remediation, health validation, and operational support.
  • Design and implement automated testing frameworks for IAM platforms, APIs, integrations, and provisioning workflows.
  • Develop automation for: 
    • Platform Health Checks
    • Service Validation
    • Monitoring Configuration
    • Incident Triage
    • Alert Enrichment
    • Automated Recovery Tasks
    • Operational Reporting
  • Drive Infrastructure as Code (IaC) and configuration automation practices.
  • Partner with engineering teams to integrate automated testing into CI/CD pipelines.

Availability & Resilience Management

  • Define and monitor SLIs, SLOs, Error Budgets, and Availability Targets.
  • Drive initiatives to improve uptime, performance, reliability, and recoverability.
  • Lead DR testing, failover exercises, resiliency validation, and business continuity readiness activities.
  • Identify operational bottlenecks and implement preventive controls.

Incident & Service Health Management

  • Lead major incident response, technical troubleshooting, and root cause analysis.
  • Drive post-incident reviews and corrective action programs.
  • Analyze incident trends and recurring service issues.
  • Implement proactive monitoring and predictive alerting solutions.
  • Improve Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

Required Qualifications

  • 10+ years of experience in IAM, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure Engineering, or Cybersecurity.
  • Strong experience supporting enterprise IAM technologies, including: 
    • PAM
    • Active Directory
    • PKI
    • Authentication Services
    • Secrets Management
    • Azure AD / Entra ID
    • Cloud Identity Platforms
  • Expertise in observability, monitoring design, telemetry, and service instrumentation.
  • Hands-on experience with enterprise monitoring platforms including Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, or equivalent.
  • Strong scripting and automation experience using PowerShell, Python, Bash, or similar technologies.
  • Experience building automated testing frameworks and validation solutions for enterprise platforms.
  • Experience with APIs, workflow automation, CI/CD pipelines, and Infrastructure as Code.
  • Strong understanding of High Availability, Disaster Recovery, Resilience Engineering, and Business Continuity.
  • Experience conducting root cause analysis and driving operational improvement initiatives.
  • Strong communication, stakeholder management, and leadership skills.
Actual salary is determined based on the role, location, individual experience, skills, and other considerations. We are an equal opportunity employer and value diversity at our company. We do not discriminate on the basis of race, religion, color, national origin, sex, gender, gender expression, sexual orientation, age, marital status, veteran status, or disability status. We will ensure that individuals with disabilities are provided reasonable accommodation to participate in the job application or interview process, to perform essential job functions, and to receive other benefits and privileges of employment. Please contact us to request accommodation.

Skills Required

  • 10+ years of experience in IAM, Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or Cybersecurity
  • Experience supporting enterprise IAM technologies, including PAM, Active Directory, PKI, authentication services, secrets management, Azure AD or Entra ID, and cloud identity platforms
  • Expertise in observability, monitoring design, telemetry, and service instrumentation
  • Hands-on experience with enterprise monitoring platforms such as Splunk, Dynatrace, Datadog, Grafana, Azure Monitor, or equivalent
  • Strong scripting and automation experience using PowerShell, Python, Bash, or similar technologies
  • Experience building automated testing frameworks and validation solutions for enterprise platforms
  • Experience with APIs, workflow automation, CI/CD pipelines, and Infrastructure as Code
  • Strong understanding of high availability, disaster recovery, resilience engineering, and business continuity
  • Experience conducting root cause analysis and driving operational improvement initiatives
  • Strong communication, stakeholder management, and leadership skills

The Depository Trust & Clearing Corporation (DTCC) Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about The Depository Trust & Clearing Corporation (DTCC) and has not been reviewed or approved by The Depository Trust & Clearing Corporation (DTCC).

  • Healthcare Strength Medical, dental, and vision coverage are comprehensive and include mental health support, in-network preventive care, and company-supported HSA/FSA options. Company HSA contributions plus life, AD&D, and disability coverage further reinforce depth of protection.
  • Retirement Support A 401(k) with company matching and references to pension-related benefits provide meaningful long-term savings support. Charitable contribution matching and financial-planning resources complement the core retirement offerings.
  • Parental & Family Support Generous parental leave, adoption assistance, childcare benefits, and dependent care FSAs indicate strong support for families. Back-up care and specialized family health programs are highlighted as part of the package.

The Depository Trust & Clearing Corporation (DTCC) Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Jersey City, NJ
5,075 Employees
Year Founded: 1973

What We Do

With over 45 years of experience, DTCC is the premier post-trade market infrastructure for the global financial services industry. From 21 locations around the world, DTCC, through its subsidiaries, automates, centralizes and standardizes the processing of financial transactions, mitigating risk, increasing transparency and driving efficiency for thousands of broker/dealers, custodian banks and asset managers. Industry owned and governed, the firm simplifies the complexities of clearing, settlement, asset servicing, data management, data reporting and information services across asset classes, bringing increased security and soundness to financial markets. In 2021, DTCC’s subsidiaries processed securities transactions valued at nearly U.S. $2.4 quadrillion. Its depository provides custody and asset servicing for securities issues from 177 countries and territories valued at U.S. $87.1 trillion. DTCC’s Global Trade Repository service, through locally registered, licensed, or approved trade repositories, processes 16 billion messages annually. To learn more, please visit us at www.dtcc.com.

Similar Jobs

Optum Logo Optum

Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Software Engineering Lead

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Software Engineering Lead - Python Fullstack, FastAPI, GenAI, AI ML

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Principal Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account