Job Title: CaaS Private Site Reliability Engineer
Corporate Title: Vice President
Location: Cary, NC
Who we are:
In short – an essential part of Deutsche Bank’s technology solution, developing applications for key business areas.
Our Technologists drive Cloud, Cyber and business technology strategy while transforming it within a robust, hands-on engineering culture. Learning is a key element of our people strategy, and we have a variety of options for you to develop professionally. Our approach to the future of work champions flexibility and is rooted in the understanding that there have been dramatic shifts in the ways we work.
Having first established a presence in the Americas in the 19th century, Deutsche Bank opened its US technology center in Cary, North Carolina in 2009. Learn more about us here.
Overview
As a CaaS Private Site Reliability Engineer, you will lead reliability, resilience, and operational excellence for the CaaS Private platform in US. You will bring production engineering discipline to Kubernetes, observability, automation, and incident management while helping teams improve service health and platform readiness. You will partner across engineering, operations, and application teams to strengthen SLOs, reduce manual intervention, and ensure platform changes are measurable, supportable, and aligned to enterprise reliability standards.
What We Offer You
- A diverse and inclusive environment that embraces change, innovation, and collaboration
- A hybrid working model, allowing for in-office / work from home flexibility, generous vacation, personal and volunteer days
- Employee Resource Groups support an inclusive workplace for everyone and promote community engagement
- Competitive compensation packages including health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits
- Educational resources, matching gift and volunteer programs
What You’ll Do
- Lead the reliability strategy for the CaaS Private platform in US, including SLO frameworks, operational standards, and incident management maturity
- Drive resilience improvements across observability, capacity planning, upgrade safety, disaster readiness, and operational automation
- Own complex production management issues by leading troubleshooting, identifying root causes, and implementing preventive fixes
- Define and refine service indicators, alert thresholds, dashboard standards, production readiness criteria, escalation paths, and postmortem follow-through
- Develop automation and self-healing workflows that reduce manual intervention, improve recovery times, and strengthen platform supportability
- Partner with cross-functional teams to ensure platform changes are measurable, supportable, and aligned with reliability objectives
How You’ll Lead
- Mentor junior and middle engineers while fostering a culture of blameless learning, measurable reliability, and operational excellence
- Influence platform architecture and roadmap decisions using data from incidents, capacity models, operational trends, and reliability metrics
- Communicate strategic insights and practical recommendations to engineering, operations, and business stakeholders to guide reliability investments
Skills You’ll Need
- Extensive experience with Hands on Bare Metal Kubernetes, Linux, distributed systems reliability, and production platform operations
- Strong hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root cause analysis
- Proven ability to design and implement automation, self-healing workflows, operational checks, maintenance tasks, and runbook improvements
- Experience defining SLOs, service indicators, alert quality standards, production readiness practices, and escalation models
- Ability to lead complex reliability improvements independently while partnering across engineering, operations, and application teams
- Proven ability to leverage AI tools to enhance productivity, optimize workflows to solve business problems, while applying critical judgment to ensure responsible and ethical use of data and AI outputs
Skills That Will Help You Excel
- Strong communication skills with the ability to explain technical findings to technical and non-technical stakeholders
- Sound operational judgment with the ability to balance urgency, risk, and long-term platform stability during incidents
- Growth mindset with a focus on continuous learning, process improvement, and measurable reliability outcomes
- Experience mentoring engineers and improving operational culture across platform or infrastructure teams
- Background with automation tools, infrastructure-as-code practices, capacity planning, disaster readiness, or cloud-native platform operations
Expectations
It is the Bank’s expectation that employees hired into this role will work in the Cary, NC office in accordance with the Bank’s hybrid working model.
Deutsche Bank provides reasonable accommodations to candidates and employees with a substantiated need based on disability and/or religion.
The salary range for this position in Cary is $125,000 to $185,000. Actual salaries may be based on a number of factors including, but not limited to, a candidate’s skill set, experience, education, work location and other qualifications. Posted salary ranges do not include incentive compensation or any other type of remuneration.
Deutsche Bank Benefits
At Deutsche Bank, we recognize that our benefit programs have a profound impact on our colleagues. That’s why we are focused on providing benefits and perks that enable our colleagues to live authentically and be their whole selves, at every stage of life. We provide access to physical, emotional, and financial wellness benefits that allow our colleagues to stay financially secure and strike balance between work and home. Click here to learn more!
Learn more about your life at Deutsche Bank through the eyes of our current employees: https://careers.db.com/life
The California Consumer Privacy Act outlines how companies can use personal information. If you are interested in receiving a copy of Deutsche Bank’s California Privacy Notice please email [email protected].
#LI-HYBRID
We strive for a culture in which we are empowered to excel together every day. This includes acting responsibly, thinking commercially, taking initiative and working collaboratively.
Together we share and celebrate the successes of our people. Together we are Deutsche Bank Group.
We welcome applications from all people and promote a positive, fair and inclusive work environment.
Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status or other characteristics protected by law. Click these links to view Deutsche Bank’s Equal Opportunity Policy Statement and the following notices: EEOC Know Your Rights; Employee Rights and Responsibilities under the Family and Medical Leave Act; and Employee Polygraph Protection Act.
Skills Required
- Extensive experience with Kubernetes, Linux, distributed systems reliability, and production platform operations
- Hands-on experience with observability, monitoring, alerting, dashboarding, incident response, and root cause analysis
- Proven ability to design and implement automation, self-healing workflows, operational checks, maintenance tasks, and runbook improvements
- Experience defining SLOs, service indicators, alert quality standards, production readiness practices, and escalation models
- Ability to lead complex reliability improvements independently while partnering across engineering, operations, and application teams
- Strong communication skills to explain technical findings to technical and non-technical stakeholders
- Experience mentoring engineers and improving operational culture across platform or infrastructure teams
- Background with automation tools, infrastructure-as-code practices, capacity planning, disaster readiness, or cloud-native platform operations
- Work in Cary, NC office in accordance with the Bank's hybrid working model
Deutsche Bank Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Deutsche Bank and has not been reviewed or approved by Deutsche Bank.
-
Healthcare Strength — Health coverage is positioned as comprehensive, spanning multiple medical plan options along with dental, vision, prescription coverage, life insurance, and disability protection.
-
Leave & Time Off Breadth — Time away is described as generous, including annual leave, sick leave, public holidays, wellbeing leave, volunteering leave in some regions, and expanded bereavement leave in certain locations.
-
Retirement Support — Retirement support is presented as a matched savings plan (401(k)), reinforcing longer-term financial security as part of the rewards package.
Deutsche Bank Insights
What We Do
At Deutsche Bank, we give original thinkers the space and support they need to shine. Merging local knowledge with global vision, in-depth insight with industry-leading digital expertise, if you’re an innovator by nature, we can help you to unleash your potential. We see things differently at Deutsche Bank – and we’re proud of our fresh perspective. Today, we’re driving growth through our strong client franchise, investing heavily in digital technologies, prioritising long-term success over short term gains, and serving society with ambition and integrity. Wherever your interests lie – in investment banking, trading, private wealth, asset management, retail banking - or many of the infrastructure functions that support them – you’ll discover resources, training and opportunities designed to keep you ahead of the curve. Intelligence has no boundaries: we welcome high-achieving, talented individuals from any background. If you’re full of imagination, enjoy solving problems and respond positively to complex challenges, discover a career to look forward to and join us!









