Technology Operations Specialist, VP

Reposted 21 Days Ago
Be an Early Applicant
Bay, Laguna, Calabarzon, PHL
In-Office
Expert/Leader
Fintech • Financial Services
The Role
Lead production support for TST domain, resolve functional and technical incidents, drive automation and monitoring improvements, manage major incident recovery and stakeholder communications, maintain runbooks/KEDB, mentor global team, and apply SRE principles (SLO/SLI).
Summary Generated by Built In
Job Description:

Job Title: Technology Operations Specialist

Corporate Title: VP

Location: Pune, India

Role Description

You will be operating within the Production Support Services team of the TST domain, spanning SES, TAS and Trade Finance Lending within Corporate Bank Production Services. The role is focused on providing hands-on production support, troubleshooting and operational ownership for applications hosted on Google Cloud Platform (GCP), with strong emphasis on Google Kubernetes Engine (GKE), containerised workloads, PostgreSQL/Postgres-backed services and related cloud-native technologies. You will be expected to lead incident recovery, perform deep technical analysis across application, infrastructure, database and platform layers, improve observability, and drive automation to enhance service stability, resilience and operational efficiency.

  • Lead technical and functional troubleshooting for production incidents, user requests and platform issues across GCP-hosted applications, including GKE, Cloud Run, Compute Engine, Cloud Storage, IAM, networking, PostgreSQL/Postgres databases and database connectivity.
  • Drive end-to-end incident recovery for P3H and above incidents by analysing logs, metrics, traces, deployment history, configuration changes, infrastructure events and application behaviour.
  • Use GCP observability tools such as Cloud Logging, Cloud Monitoring, Error Reporting, dashboards, uptime checks and alert policies to identify root cause, reduce mean time to detect and improve service reliability.
  • Partner with application, infrastructure, SRE, security and engineering teams to troubleshoot cloud networking, container runtime, IAM, quota, performance, scaling and availability issues.
  • Drive automation, toil reduction, platform hygiene, monitoring improvements and operational controls for cloud-native applications and supporting infrastructure.
  • Prepare and distribute clear incident communications, technical updates, recovery timelines and service restoration summaries for senior stakeholders.
  • Participate in CAB, change validation and post-change support activities to assess operational risk, ensure rollback readiness and support safer change delivery.
  • Collaborate with Global Incident Management, L2/L3 support, SRE and platform teams to orchestrate recovery of major incidents and implement preventive actions.

What we’ll offer you

As part of our flexible scheme, here are just some of the benefits that you’ll enjoy,

  • Best in class leave policy.
  • Gender neutral parental leaves
  • 100% reimbursement under childcare assistance benefit (gender neutral)
  • Sponsorship for Industry relevant certifications and education
  • Employee Assistance Program for you and your family members
  • Comprehensive Hospitalization Insurance for you and your dependents
  • Accident and Term life Insurance
  • Complementary Health screening for 35 yrs. and above

Your key responsibilities

  • Provide hands-on technical and functional production support for applications deployed on GCP and integrated enterprise platforms within the TST domain.
  • Troubleshoot complex production issues across cloud runtime, containers, application services, middleware, databases, network connectivity, IAM permissions, certificates, secrets and deployment pipelines.
  • Analyse GCP logs, metrics and alerts using Cloud Logging, Cloud Monitoring, dashboards and log-based metrics to identify root cause and restore service quickly.
  • Support containerised workloads running on GKE and Cloud Run, including pod/container restarts, scaling behaviour, node pressure, cluster events, health checks, readiness/liveness failures, resource saturation, ingress/service issues and deployment rollbacks.
  • Troubleshoot PostgreSQL/Postgres production issues, including connection failures, query performance, locks, replication or failover symptoms, storage growth, backup/restore readiness and application-to-database connectivity.
  • Build technical and functional subject matter expertise across supported applications, business flows, cloud architecture, service dependencies and infrastructure configuration.
  • Identify proactive opportunities for automation, self-healing, alert rationalisation, toil reduction and improved operational resilience.
  • Drive service requests and incidents to resolution within L2 scope, ensuring timely escalation to L3, SRE, infrastructure or vendor teams where deeper engineering intervention is required.
  • Review monitoring coverage for critical services, SLIs, SLOs, error budgets, availability, latency, throughput, saturation and business-critical transaction flows.
  • Maintain and continuously improve runbooks, KEDB articles, support procedures, troubleshooting guides and operational readiness documentation.
  • Participate in BCP, DR, EDR, SSRTO and component failure tests for cloud and application services, validating recovery procedures and operational readiness.
  • Understand data flow, request flow and service dependencies across cloud infrastructure to provide effective operational support during incidents and changes.
  • Own managed risk, operational controls, production hygiene and continuous improvement initiatives for supported services.
  • Mentor and coach global support team members on GCP troubleshooting, incident handling, observability and production support best practices.
  • Apply an SRE mindset with practical understanding of SLIs, SLOs, error budgets, incident learning and reliability improvement.

Your skills and experience

Must Have: -

  • 12+ years of IT experience in large corporate environments, with strong exposure to controlled production support environments, preferably within Financial Services Technology.
  • Strong hands-on experience supporting and troubleshooting production workloads on Google Cloud Platform, with deeper focus on GKE, Cloud Run, Compute Engine, Cloud Storage, Cloud SQL/PostgreSQL, IAM, VPC networking, load balancers and service accounts.
  • Proven ability to investigate incidents using Cloud Logging, Cloud Monitoring, Error Reporting, dashboards, alerts, log-based metrics and application traces. Strong troubleshooting skills across application, infrastructure and platform layers, including container failures, scaling issues, latency, memory/CPU saturation, network connectivity, DNS, certificates, secrets, IAM permissions and deployment failures.
  • Strong hands-on experience with Kubernetes concepts and GKE operations, including pods, deployments, services, ingress, node pools, namespaces, autoscaling, config maps, secrets, health checks, cluster events, workloads, logs and kubectl-based diagnostics.
  • Working knowledge of cloud-native application architecture, microservices, REST APIs, service-to-service communication, event-driven flows and enterprise integration patterns.
  • Strong understanding of UNIX/Linux operating systems, shell commands, process analysis, file systems, logs, network utilities and infrastructure troubleshooting.
  • Working knowledge of scripting and automation using UNIX shell, Python, PowerShell, Perl or similar tools for diagnostics, reporting, remediation and toil reduction.
  • Understanding of middleware and messaging platforms such as MQ, Kafka or similar, including operational troubleshooting of connectivity, message flow and performance issues.
  • Understanding of web and application server environments such as Apache, Tomcat, WebLogic or equivalent cloud-hosted runtimes.
  • Strong understanding of relational databases, especially PostgreSQL/Postgres and Cloud SQL for PostgreSQL, including connectivity, query performance, locks, indexing concepts, backup/restore, storage growth, availability and operational troubleshooting.
  • Experience with enterprise monitoring tools such as Geneos, AppDynamics, New Relic, Dynatrace, Grafana or equivalent, alongside GCP observability tools. Strong understanding of ITIL Service Management practices, including Incident, Problem, Change, Request and Major Incident Management.
  • Experience in securities services, asset management or trade finance domains will be an advantage.
  • Flexibility to support incident callouts after office hours and during weekends, including participation in rotational weekend cover.

Nice to Have:

  • Good analytical and problem-solving skills.
  • Excellent communication skills, both written and verbal, with attention to detail.
    • Ability to work in virtual teams and in matrix structures.
    • Team management and program management skills
  • ITIL / best practice service context. ITIL foundation is plus.
  • Ticketing Tool experience – Service Desk, Service Now.
  • Understanding of SRE concepts (SLA, SLO’s, SLI’s)
  • Knowledge and development experience in Ansible automation.
  • Working knowledge of one cloud platform (AWS or GCP).

How we’ll support you

  • Training and development to help you excel in your career
  • Coaching and support from experts in your team
  • A culture of continuous learning to aid progression
  • A range of flexible benefits that you can tailor to suit your needs

About us and our teams

Please visit our company website for further information:

https://www.db.com/company/company.html

We strive for a culture in which we are empowered to excel together every day. This includes acting responsibly, thinking commercially, taking initiative and working collaboratively.

Together we share and celebrate the successes of our people. Together we are Deutsche Bank Group.

We welcome applications from all people and promote a positive, fair and inclusive work environment.

Skills Required

  • 14+ years of IT experience in large corporate or controlled production environments
  • Experience in Financial Services Technology or client-facing production support
  • Experience in securities services, asset management, or payments domains
  • Expert level knowledge of GCP and cloud-native architecture and applications
  • Intermediate knowledge of AI, LLM working principles and agentic AI
  • Working knowledge of scripting: UNIX shell, PowerShell, PERL, Python
  • Flexibility for incident callouts (after hours and weekends) and rotation weekend cover
  • Understanding of Java
  • Understanding of operating systems: UNIX and Linux
  • Understanding of middleware such as MQ or Kafka
  • Understanding of WebLogic and webserver environments (Apache, Tomcat)
  • Understanding of RDBMS: Oracle, MS-SQL, Sybase and NoSQL databases
  • Understanding of batch monitoring tools (Control-M, AutoSys)
  • Understanding of monitoring tools (GCO, Geneos, AppDynamics, New Relic, Dynatrace, Grafana)
  • Familiarity with ITIL service management processes (Incident, Problem, Change)
  • SRE mindset with implementation knowledge of SLOs and SLIs
  • Good analytical and problem-solving skills
  • Excellent written and verbal communication skills
  • Ability to work in virtual teams and matrix structures
  • Team management and program management skills
  • ITIL Foundation certification
  • Ticketing tool experience (ServiceDesk, ServiceNow)
  • Knowledge and development experience in Ansible automation
  • Working knowledge of one cloud platform (AWS or GCP) - AWS preferred if GCP already covered

Deutsche Bank Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Deutsche Bank and has not been reviewed or approved by Deutsche Bank.

  • Healthcare Strength Health coverage is positioned as comprehensive, spanning multiple medical plan options along with dental, vision, prescription coverage, life insurance, and disability protection.
  • Leave & Time Off Breadth Time away is described as generous, including annual leave, sick leave, public holidays, wellbeing leave, volunteering leave in some regions, and expanded bereavement leave in certain locations.
  • Retirement Support Retirement support is presented as a matched savings plan (401(k)), reinforcing longer-term financial security as part of the rewards package.

Deutsche Bank Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Frankfurt am Main
68,787 Employees

What We Do

At Deutsche Bank, we give original thinkers the space and support they need to shine. Merging local knowledge with global vision, in-depth insight with industry-leading digital expertise, if you’re an innovator by nature, we can help you to unleash your potential. We see things differently at Deutsche Bank – and we’re proud of our fresh perspective. Today, we’re driving growth through our strong client franchise, investing heavily in digital technologies, prioritising long-term success over short term gains, and serving society with ambition and integrity. Wherever your interests lie – in investment banking, trading, private wealth, asset management, retail banking - or many of the infrastructure functions that support them – you’ll discover resources, training and opportunities designed to keep you ahead of the curve. Intelligence has no boundaries: we welcome high-achieving, talented individuals from any background. If you’re full of imagination, enjoy solving problems and respond positively to complex challenges, discover a career to look forward to and join us!

Similar Jobs

Deutsche Bank Logo Deutsche Bank

Operations Specialist

Fintech • Financial Services
In-Office
Bay, Laguna, Calabarzon, PHL
68787 Employees
Remote or Hybrid
2 Locations
289097 Employees

Smartly Logo Smartly

Technical Support

AdTech • Artificial Intelligence • Digital Media • Marketing Tech • Social Media • Software • Generative AI
Easy Apply
Remote or Hybrid
Philippines
805 Employees

Zscaler Logo Zscaler

Senior Sales Engineer

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
Philippines
8697 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account