Site Reliability Engineer III

Reposted 28 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Fintech • Financial Services
The Role
Design, build, and operate reliable cloud-native platforms using SRE principles. Implement automation, monitoring, incident response, and scalability improvements. Partner with development, security, and platform teams to drive reliability, migrations, and operational best practices while mentoring peers.
Summary Generated by Built In

Candescent is a forward-thinking technology company transforming how financial institutions deliver Intelligent Banking experiences. We unite digital banking, account opening, and branch solutions that power and connect digital banking, account opening, and branch solutions—creating seamless engagement across digital, remote, and in-person channels.

Our Experience-Led, Intelligence-Driven approach combines human-centered design with data, automation, and cloud-based innovation. Built on an API-first architecture, our extensible ecosystem enables institutions to adapt quickly, integrate easily, and unlock new opportunities for growth—turning every customer interaction into a moment of clarity, confidence, and connection.

Position: Site Reliability Engineer III

Experience: 6-9 Years

Location: Bangalore

We are looking for a strong Application Site Reliability Engineer (SRE) to support and improve the reliability of Java-based production systems running on Kubernetes in cloud environments.

This role focuses on application-level reliability, JVM deep troubleshooting, production incident management, and close collaboration with development teams — not infrastructure provisioning or CloudOps.

The ideal candidate understands how Java applications behave in production and can proactively improve performance, scalability, and operational maturity.

Key Responsibilities:

  • Support and operate production Java applications running on Kubernetes (GKE).
  • Troubleshoot complex application issues using logs, metrics, traces, heap dumps, and thread dumps.
  • Participate in incident response, root cause analysis, and blameless postmortems.
  • Collaborate closely with development teams to understand application architecture, dependencies, and failure patterns.
  • Analyze JVM behavior (heap, GC, memory leaks, OOM, thread contention) and recommend performance improvements.
  • Define and improve SLIs, SLOs, alerts, and dashboards.
  • Support application deployments, rollbacks, and runtime configuration changes.
  • Identify reliability, performance, and scalability gaps in application behavior.
  • Automate repetitive operational tasks to reduce toil.
  • Drive improvements in runbooks, operational readiness, and on-call effectiveness.
  • Advocate and influence adoption of shift-left reliability practices.

Must-Have Skills & Experience:

  • Strong hands-on experience supporting Java applications in production.
  • Deep understanding of JVM internals:
    • Heap & memory management
    • Garbage collection
    • tuning OOM analysis
    • Thread dump and performance analysis
  • Proven experience in incident response and production troubleshooting.
  • Experience operating applications on Kubernetes from an application/runtime perspective.
  • Strong experience with application observability:
    • Logs
    • Metrics
    • Monitoring tools
    • Distributed tracing
  • Solid understanding of SLIs, SLOs, and reliability-driven operations.
  • Experience with deployment strategies (rolling, blue/green, canary).
  • Ability to write scripts/automation (Python, Shell, or similar) to reduce operational toil.
  • Strong understanding of application architecture and service dependencies (databases, messaging systems, external APIs).
  • Ability to analyze and troubleshoot issues holistically across the entire application stack, rather than focusing on isolated components.
  • Strong collaboration and communication skills.
  • Demonstrates accountability and sound judgment during high-pressure production incidents.

Cloud & Platform Exposure

  • Experience working with applications deployed on Kubernetes in GCP.
  • Familiarity with GKE environments from an application operations perspective (not CloudOps or infrastructure engineering).
  • Understanding of cloud constructs relevant to application behavior (networking, IAM, storage, compute).

Good-to-Have Skills

  • CI/CD pipeline exposure (GitHub Actions, Jenkins).
  • Familiarity with GitOps practices.
  • Experience supporting cloud migrations or modernization initiatives.
  • Exposure to platform or infrastructure concepts supporting application workloads.

What We Value

  • Ownership mindset and reliability-first thinking.
  • Curiosity to investigate deep production issues.
  • Strong collaboration with development teams.
  • Bias toward automation and continuous improvement.
  • Clear communication during incidents and stakeholder updates.

Statement to Third Party Agencies
To ALL recruitment agencies: Candescent only accepts resumes from agencies on the preferred supplier list. Please do not forward resumes to our applicant tracking system, Candescent employees, or any Candescent facility. Candescent is not responsible for any fees or charges associated with unsolicited resumes.

Skills Required

  • Building and supporting production level Kubernetes clusters; optimizing containerized workloads
  • Experience with cloud networking; configuring VPCs, firewalls, ingress/egress, CDN
  • Experience in AWS services
  • Hands-on experience with EKS, RDS, Lambda, CloudWatch, EFS, S3, EBS
  • Experience in Terraform and Terragrunt
  • Experience in cloud migrations
  • BS in Computer Science or related field, or equivalent experience
  • High initiative and clear communication skills
  • Ability to set up and troubleshoot environments
  • Extensive experience with Prometheus, Dynatrace, or other logging/observability tools
  • Experience with application and infrastructure automation, orchestration, and configuration management
  • Experience operating within cloud environments
  • Drive DevOps development, automation and deployment practices, policies and standards
  • Container build/management and Kubernetes (desired)
  • Cloud migrations including Google Cloud (desired)
  • Scripting with Python (desired)
  • CI/CD using GitHub (desired)
  • Version control with GIT and GitOps (desired)

Candescent Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Candescent and has not been reviewed or approved by Candescent.

  • Leave & Time Off Breadth Policies include unlimited vacation for full-time exempt staff, tenure-based accrual for non-exempt, plus floating holidays and sick leave. This breadth of time off suggests flexibility across employment classifications.
  • Wellbeing & Lifestyle Benefits A discount program is cited that provides access to deals at over 250 retailers. This perk adds everyday savings beyond core benefits.

Candescent Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Atlanta, Georgia
1,030 Employees
Year Founded: 2024

What We Do

Candescent brings together the transformative technologies that power and connect account opening, digital banking and branch solutions for banks and credit unions of all sizes. And we’re here to help you extend, differentiate and illuminate your digital-first banking experiences. Our industry-leading products and services, cloud architecture and on-demand developer tools give you the power to differentiate and deliver seamless customer journeys.

Similar Jobs

CrowdStrike Logo CrowdStrike

Site Reliability Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
11000 Employees

American Express Logo American Express

Site Reliability Engineer

Fintech • Financial Services
Hybrid
2 Locations
100703 Employees

American Express Logo American Express

Site Reliability Engineer

Fintech • Financial Services
Hybrid
2 Locations
100703 Employees
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
528 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account