Executive Director, Platform Foundations & Observability

Reposted 26 Days Ago
Be an Early Applicant
Chicago, IL, USA
Hybrid
217K-342K Annually
Senior level
Big Data • Cloud • Fintech • Information Technology • Financial Services
We clear and settle trades for the options industry.
The Role
The Executive Director will lead Platform Engineering Governance, SRE practices, cloud strategy, metrics reporting, and compliance. The role requires strategic vision, deep technical expertise, and effective management of high-performing teams to ensure operational excellence and risk management.
Summary Generated by Built In

To be considered for this position, applications and resumes are accepted only through our careers site by directly applying to the posted job. We do not accept unsolicited resumes or sales solicitations from staffing agencies. Any OCC employee wishing to submit a referral must do so through their Workday account. Any resume submitted outside of an active job posting will not be considered for employment.

Summary

This role provides executive-level leadership over Platform Foundations, Platform Observability, Site Reliability Engineering, Metrics & Reporting, and PM Coordination within the Platform Engineering organization. The leader in this role is accountable for establishing and scaling strategic platform infrastructure — including Kubernetes, Kafka, and Data Platforms — alongside Platform Observability and a mature SRE practice, ensuring the organization operates with clear, data-driven visibility into platform health and performance. The ideal candidate brings deep technical credibility, executive communication skills, and a proven track record of building and operationalizing platform engineering capabilities at scale in a regulated environment.

Primary Duties and Responsibilities:

To perform this job successfully, an individual must be able to perform each primary duty satisfactorily.

Platform Foundations

  • Serve as Product Manager over the Platform Foundations portfolio, owning the product vision and roadmap encompassing Kubernetes, Kafka, and Data Platforms as the strategic infrastructure backbone enabling all Platform Engineering capabilities.
  • Lead the strategy and delivery of Kubernetes-based container orchestration capabilities, ensuring scalable, reliable, and governed compute infrastructure is available across the organization.
  • Own the Kafka streaming and messaging platform strategy, ensuring the backbone for event-driven architecture and data movement is reliable, scalable, and aligned to organizational needs.
  • Lead the strategy and delivery of Data Platform capabilities that enable Data Engineering and Data Analytics teams, ensuring scalable, reliable, and governed data infrastructure is available across the organization.
  • Define and own the Platform Foundations product roadmap, prioritizing capabilities that reduce engineering toil, accelerate delivery, and strengthen the reliability of core infrastructure across Platform Engineering.
  • Partner with Platform Engineering Product teams and consumers to ensure foundational capabilities align to evolving organizational needs and strategic priorities.
  • Serve as the executive advocate for platform foundations, ensuring Kubernetes, Kafka, and Data Platform capabilities are treated as first-class products with clear ownership, roadmaps, and service levels.

Platform Observability

  • Define and own the golden path for platform observability, establishing standardized logging, metrics, tracing, and alerting patterns that all platform services adopt consistently.
  • Drive adoption of observability standards across platform teams, ensuring every service meets baseline visibility requirements before reaching production.
  • Own the observability tooling strategy and roadmap, ensuring the right platforms and integrations are in place to provide end-to-end visibility across the platform engineering estate.

Site Reliability Engineering & COE

  • Lead the scaling and maturation of the SRE practice, establishing error budgets, SLOs, SLAs, and incident response frameworks across all platform services.
  • Define and enforce reliability standards including on-call models, blameless postmortem processes, and corrective action tracking to drive continuous improvement.
  • Partner with Platform Engineering Product teams to embed reliability principles into build and operate models across all platform domains.
  • Champion toil reduction through automation, ensuring engineering capacity is redirected from manual operations to higher-value platform capabilities.
  • Establish and operate the SRE Center of Excellence (COE), defining shared reliability standards, tooling, and practices that scale consistently across Platform Engineering teams.

Metrics & Reporting

  • Own the platform metrics and reporting function, establishing a consistent framework for measuring platform health, engineering velocity, reliability, and cost efficiency across Platform Engineering.
  • Define and track KPIs aligned to internal SLAs, executive reporting needs, and audit/compliance requirements.
  • Ensure Jira and other platform tooling serve as the single source of truth for work visibility, with dashboards and reporting that enable data-driven prioritization.
  • Build and maintain reporting cadences for leadership, including platform health scorecards, capacity forecasting, and risk transparency.
  • Establish a metrics governance framework including taxonomy, data quality standards, ownership models, and baseline targets to ensure consistency and reliability of platform reporting.
  • Design and maintain executive reporting artifacts that communicate platform health, engineering performance, and compliance posture in a clear, actionable format.
  • Define and operationalize feedback mechanisms including CSAT and NPS survey programs to measure platform team effectiveness and customer satisfaction across Platform Engineering stakeholders.

PM Coordination & Platform Delivery

  • Serve as the primary engineering leadership partner to the Platform Engineering Program Management function, ensuring platform initiatives are properly scoped, sequenced, and resourced.
  • Drive alignment between engineering capacity and roadmap commitments, proactively surfacing dependency risks and trade-off decisions to the Platform Engineering Executive Director.
  • Coordinate across Platform Engineering domains to ensure cross-team delivery dependencies are managed and resolved effectively.
  • Partner with Product and Engineering leaders outside of Platform Engineering to align platform capabilities to broader organizational roadmaps.
  • Represent Platform Engineering at the executive level, delivering clear, compelling reporting and presentations to C-Suite leadership on platform health, roadmap progress, strategic investments, and risk posture.

Leadership & Organizational Excellence

  • Lead, develop, and retain a high-performing team of engineering managers and individual contributors with clear ownership, career paths, and accountability frameworks.
  • Foster a culture consistent with Platform Engineering operating principles: automation-first, full-stack ownership, stability as a prerequisite for velocity, and transparency through tooling.
  • Manage budget for areas of responsibility; ensure adherence to schedules, work plans, and performance requirements.
  • Oversee remediation of audit findings and observations within areas of responsibility, ensuring root cause is addressed, residual risk is reduced, and remediation is completed timely.
Supervisory Responsibilities:
  • Manages a team of engineering managers and senior technical staff across Platform Foundations (Kubernetes, Kafka, and Data Platforms), Platform Observability, Site Reliability Engineering & COE, and Metrics & Reporting functions.

Qualifications:

The requirements listed are representative of the knowledge, skill, and/or ability required.  Reasonable accommodations may be made to enable individuals with disabilities to perform the primary functions.

  • [Required] Proven executive-level leadership of platform engineering, SRE, or infrastructure organizations in a regulated industry environment.
  • [Required] Demonstrated ability to build and scale SRE practices including SLO/SLA frameworks, on-call models, error budgets, and incident response programs.
  • [Required] Experience owning and scaling strategic platform infrastructure including container orchestration (Kubernetes), event streaming (Kafka), and data platform capabilities at enterprise scale.
  • [Required] Demonstrated experience establishing observability standards and golden path tooling across a large platform engineering organization.
  • [Required] Strong track record of cross-functional partnership with Program/Product Management, translating platform capabilities into sequenced, delivery-ready roadmaps.
  • [Required] Ability to design and maintain metrics and reporting frameworks that provide meaningful visibility into platform health, engineering performance, and compliance posture.
  • [Required] Exceptional written and verbal communication skills; ability to translate technical complexity into executive-level insights and business decisions.
  • [Required] Demonstrated ability to lead high-performing, highly technical teams through accountability, coaching, and clear ownership models.
  • [Required] Experience managing work in Agile/Scrum environments with strong prioritization and deadline management discipline.
  • [Required] Experience operating in a production change control process and working directly with audit and compliance functions in a regulated industry environment.
  • [Required] Experience in financial services or similarly regulated industries with exposure to CIS, NIST, and related frameworks.
  • [Preferred] Experience working directly with enterprise risk, legal, or compliance functions in a financial services or similarly regulated environment.

Technical Skills:

  • [Required] Deep knowledge of SRE tooling and observability platforms (e.g., Prometheus, Grafana, PagerDuty, Datadog, or equivalents).
  • [Required] Expert-level knowledge of Kubernetes and container orchestration at enterprise scale, including cluster management, networking, and workload governance.
  • [Required] Strong working knowledge of Kafka and event-driven architecture patterns, including streaming pipelines, topic governance, and consumer group management.
  • [Required] Strong working knowledge of data platform technologies and architectures supporting Data Engineering and Data Analytics use cases (e.g., data lakehouse patterns, pipeline orchestration).
  • [Required] Experience defining and implementing observability golden paths including logging standards, distributed tracing, and alerting frameworks across platform services.
  • [Required] Expert-level knowledge of cloud platforms: AWS, Azure, or GCP; experience with multi-cloud or hybrid environments preferred.
  • [Required] Strong working knowledge of cloud-native architecture patterns and Infrastructure as Code principles.
  • [Required] Experience with metrics and reporting platforms; ability to design KPI frameworks and reporting dashboards for both technical and executive audiences.
  • [Preferred] Experience with GRC tooling or platforms used to manage risk, audit findings, and compliance obligations (e.g., ServiceNow, Archer, or equivalents).
Education and/or Experience:
  • [Required] Bachelor's degree, preferably in a technical discipline (Computer Science, Mathematics, Engineering, or related field), or equivalent combination of education and experience.
  • [Required] 15+ years of progressive experience in cloud engineering, platform reliability, or infrastructure roles with at least 5 years in senior engineering leadership.
  • [Preferred] Depth and breadth of experience in a highly regulated industry such as financial services, with demonstrated understanding of applicable rules and regulatory frameworks.
Certificates or Licenses:

N/A

About Us

The Options Clearing Corporation (OCC) is the world's largest equity derivatives clearing organization. Founded in 1973, OCC is dedicated to promoting stability and market integrity by delivering clearing and settlement services for options, futures and securities lending transactions. As a Systemically Important Financial Market Utility (SIFMU), OCC operates under the jurisdiction of the U.S. Securities and Exchange Commission (SEC), the U.S. Commodity Futures Trading Commission (CFTC), and the Board of Governors of the Federal Reserve System. OCC has more than 100 clearing members and provides central counterparty (CCP) clearing and settlement services to 19 exchanges and trading platforms. More information about OCC is available at www.theocc.com.

Benefits

A highly collaborative and supportive environment developed to encourage work-life balance and employee wellness. Some of these components include:

  • A hybrid work environment, up to 2 days per week of remote work
  • Tuition Reimbursement to support your continued education
  • Student Loan Repayment Assistance
  • Technology Stipend allowing you to use the device of your choice to connect to our network while working remotely
  • Generous PTO and Parental leave
  • 401k Employer Match
  • Competitive health benefits including medical, dental and vision

Visit https://www.theocc.com/careers/thriving-together for more information.

Compensation

  • The salary range listed for any given position is exclusive of fringe benefits and potential bonuses. If hired at OCC, your final base salary compensation will be determined by factors such as skills, experience and/or education.
  • In addition, we believe in the importance of pay equity and consider internal equity of our current team members as part of any final offer.
  • We typically do not hire at the maximum of the range in order to allow for future and continued salary growth. We also offer a substantial benefits package as noted on www.theocc.com/careers
  • All employees may be eligible for a discretionary bonus. Discretionary bonuses are based on various factors, including, but not limited to, company and individual performance and are not guaranteed.

Salary Range

$216,800.00 - $342,400.00

Incentive Range

28% to 35%

This position is eligible for an annual discretionary incentive compensation award, for which the target range is listed above (see Incentive Range). The amount of such award, if any, will be based on various factors, including without limitation, both individual and company performance.

Step 1
When you find a position you're interested in, click the 'Apply' button. Please complete the application and attach your resume.  

Step 2
You will receive an email notification to confirm that we've received your application.

Step 3
If you are called in for an interview, a representative from OCC will contact you to set up a date, time, and location. 

For more information about OCC, please click here.

OCC is an Equal Opportunity Employer

Skills Required

  • Proven executive-level leadership of SRE, cloud engineering, or platform reliability organizations in a regulated industry environment.
  • Demonstrated ability to build and scale SRE practices including SLO/SLA frameworks, on-call models, error budgets, and incident response programs.
  • Deep expertise in cloud architecture strategy and governance, with experience defining and driving enterprise-wide architectural standards.
  • Strong track record of cross-functional partnership with Program/Product Management, translating platform capabilities into sequenced, delivery-ready roadmaps.
  • Demonstrated experience serving in a Product Manager capacity for technical domains such as FinOps, SecOps, or platform tooling, including ownership of roadmap, prioritization, and stakeholder alignment.
  • Experience establishing and managing governance and compliance frameworks within a platform or infrastructure engineering organization, including oversight of incidents, problem management, risk items, and audit obligations.
  • Ability to design and maintain metrics and reporting frameworks that provide meaningful visibility into platform health, engineering performance, and compliance posture.
  • Exceptional written and verbal communication skills; ability to translate technical complexity into executive-level insights and business decisions.
  • Demonstrated ability to lead high-performing, highly technical teams through accountability, coaching, and clear ownership models.
  • Experience managing work in Agile/Scrum environments with strong prioritization and deadline management discipline.
  • Experience operating in a production change control process and working directly with audit and compliance functions in a regulated environment.
  • Experience in financial services or similarly regulated industries with exposure to CIS, NIST, and related frameworks.

What the Team is Saying

Bailey
Daniel
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chicago, IL
1,200 Employees
Year Founded: 1973

What We Do

As the foundation for secure markets, OCC is a customer-driven organization that delivers world-class Risk Management, Clearing, and Settlement Services for a sophisticated mix of financial products that includes standard options, stock loans, and futures contracts.

Why Work With Us

We're bound together by values and behaviors that shape the way we work and live, from team projects to after-hours events and to making a difference in our communities. OCC colleagues thrive in an atmosphere of intellectual curiosity, creative problem-solving and effective interaction.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

OCC Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

A hybrid work environment, up to 2 days per week of remote work

Typical time on-site: 3 days a week
Company Office Image
HQChicago, IL
Company Office Image
Dallas, TX
Company Office Image
Washington, DC
Learn more

Similar Jobs

OCC Logo OCC

Associate Principal, Business Analysis

Big Data • Cloud • Fintech • Information Technology • Financial Services
Hybrid
2 Locations
1200 Employees
100K-153K Annually

OCC Logo OCC

Principal, AI Enablement & Change

Big Data • Cloud • Fintech • Information Technology • Financial Services
Hybrid
Chicago, IL, USA
1200 Employees
168K-226K Annually

OCC Logo OCC

Lead Associate Principal, AI Solutions

Big Data • Cloud • Fintech • Information Technology • Financial Services
Hybrid
Chicago, IL, USA
1200 Employees
123K-168K Annually

OCC Logo OCC

Associate Principal, Quality Assurance

Big Data • Cloud • Fintech • Information Technology • Financial Services
Hybrid
Chicago, IL, USA
1200 Employees
102K-137K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account