We run the production infrastructure behind a regulated, digital-first bank payment rails, core banking workloads, and customer-facing services on a multi-account AWS estate managed through code and GitOps. It's a place where reliability and trust matter deeply, and where the systems you look after have real consequences for real customers.
We're looking for an SRE-2 who enjoys that kind of responsibility: someone who keeps systems stable, automates away the repetitive work, and brings a steady hand when things break. You'll make changes through well-defined, safe paths, debug production with curiosity, and treat reliability and compliance as part of the same craft rather than competing goals.
This is a hands-on production role. You'll own real systems, share the on-call rotation with the team, and grow into the judgment that banking infrastructure rewards.
Operate and maintain production cloud infrastructure with a focus on availability, stability, and risk reduction.
Make infrastructure and configuration changes through approved patterns and GitOps workflows, so changes stay reviewable, repeatable, and safe.
Support Kubernetes-based production workloads across our environments: deployments, scaling, rollouts, and day-2 operations like debugging, recovery, and resource tuning.
Monitor system health through metrics, logs, and alerts, and use them to find and fix the root of a problem.
Participate in incident response, troubleshooting, and recovery for production systems.
Contribute to root cause analysis and help implement the preventive actions that come out of post-incident reviews.
Build automation for operational tasks (runbooks, scripts, Terraform) to reduce manual effort and free the team up for higher-value work.
Help maintain the security, compliance, and governance controls that keep a regulated banking environment healthy.
Take part in disaster recovery work DR drills, failover, cross-region posture as a meaningful part of what we deliver.
Keep runbooks, SOPs, and operational documentation useful and up to date.
Collaborate with DevEx and application teams to help their services land smoothly in production.
Must-Have
Core engineering
4–6 years in SRE / Infrastructure / Production Operations roles.
Hands-on AWS experience in production: core services (EC2, S3, IAM, VPC, RDS/Aurora, load balancing) with genuine operational depth, and comfort working within a multi-account setup.
Working knowledge of Kubernetes with day-2 experience you can debug a CrashLoopBackOff, OOMKill, failed rollout, or in-cluster DNS issue, not only deploy manifests.
Strong Linux fundamentals with real troubleshooting depth you're comfortable reasoning through a hung process, disk or inode exhaustion, a failing systemd unit, or a tricky networking issue.
Networking fundamentals (DNS, TCP/IP, HTTPS/TLS, load balancing) solid enough to reason about latency and connectivity during a live incident.
Experience with observability tooling (Prometheus, Grafana, ELK/OpenSearch), using it to form and test hypotheses when you debug.
Hands-on Terraform and Git for infrastructure change, including writing reusable, scalable modules, with a good feel for plan/apply, state, and code review.
Ability to write automation scripts (Bash required; Python strongly preferred).
How you workYou value safe, reviewable change. You're happy working through approved patterns and change windows, because you know that's what keeps a shared production environment trustworthy.
You think in terms of GitOps Git as the source of truth and you're mindful of blast radius and keeping workloads well isolated.
You're ready to share the on-call rotation, and you tend to stay calm and methodical when things get busy.
You lean toward least-privilege and just-in-time access as a natural way of working, and you're comfortable with the audit trails that come with a bank.
You see compliance (RBI, NPCI, PCI) as a set of guardrails to design within, and you're good at finding solutions that work well inside them.
You treat documentation as part of the job updating runbooks and SOPs so the next person (often future-you) has an easier time.
You use AI tooling well pairing it with your own engineering judgment rather than leaning on it blindly and you're mindful that tokens and compute have a cost, not something to take for granted.
Good-to-Have
Exposure to high-availability and disaster recovery setups (multi-AZ, cross-region replication, active-passive/active-active).
Experience with web servers (nginx, Apache).
Working knowledge of databases (PostgreSQL, DynamoDB, ElasticSearch etc.) enough to triage alongside our DBAs.
Prior experience in a regulated or compliance-heavy environment (banking, fintech, payments).
Multicloud exposure, especially with major cloud providers.
Hands-on experience across a broader set of Terraform providers.
Python or Go for building operational tooling.
Life so good, you’d think we’re kidding:
Competitive salaries. Period.
An extensive medical insurance that looks out for our employees & their dependents. We’ll love you and take care of you, our promise.
Flexible working hours. Just don’t call us at 3AM, we like our sleep schedule.
Tailored vacation & leave policies so that you enjoy every important moment in your life.
A reward system that celebrates hard work and milestones throughout the year. Expect a gift coming your way anytime you kill it here.
Learning and upskilling opportunities. Seriously, not kidding.
Good food, games, and a cool office to make you feel like home. An environment so good, you’ll forget the term “colleagues can’t be your friends”.
We believe in equality. Period.
At slice, we are committed to building a diverse and talented workforce. We never discriminate on the basis of race, sex, religion, colour, national origin, gender, gender identity, sexual orientation, age, marital status, veteran status, medical condition, disability, or any other class or characteristic protected by the applicable law.
We consider all qualified job-seekers with criminal histories in a manner consistent with the applicable law. Additionally, we are committed to providing reasonable accommodations to qualified individuals with physical or mental disabilities in order to participate in the job application or interview process, perform essential job functions, and receive other benefits and privileges of employment.
Come join our crew!
About slice:
slice
A new bank for a new India
slice’s purpose is to make the world better at using money and time, with a major focus on building the best consumer experience for your money. We’ve all felt how slow, confusing, and complicated banking can be. So, we’re reimagining it. We’re building every product from scratch to be fast, transparent, and feel good, because we believe that the best products transcend demographics, like how great music touches most of us.
Our cornerstone products and services: slice savings account, slice UPI credit card, slice UPI, and slice business are designed to be simple, rewarding, and completely in your control. At slice, you’ll get to build things you’d use yourself and shape the future of banking in India. We tailor our working experience with the belief that the present moment is the only real thing in life. And we have harmony in the present the most when we feel happy and successful together.
We’re backed by some of the world’s leading investors, including Tiger Global, Insight Partners, Advent International, Blume Ventures, and Gunosy Capital.
Skills Required
- 4-6 years in SRE / Infrastructure / Production Operations roles
- Hands-on AWS production experience (EC2, S3, IAM, VPC, RDS/Aurora, load balancing) in a multi-account setup
- Kubernetes day-2 operations experience (debug CrashLoopBackOff, OOMKill, failed rollouts, in-cluster DNS)
- Strong Linux fundamentals and troubleshooting (hung processes, disk/inode exhaustion, failing systemd units, networking issues)
- Networking fundamentals (DNS, TCP/IP, HTTPS/TLS, load balancing) for incident reasoning
- Experience with observability tooling (Prometheus, Grafana, ELK/OpenSearch) and using metrics/logs to debug
- Hands-on Terraform and Git for infra change, including reusable modules and plan/apply/state management
- Bash scripting for automation
- Python scripting for automation and tooling
- Exposure to high-availability and disaster recovery setups (multi-AZ, cross-region replication, active-passive/active-active)
- Experience with web servers (nginx, Apache)
- Working knowledge of databases (PostgreSQL, DynamoDB, Elasticsearch) to triage with DBAs
- Prior experience in regulated or compliance-heavy environments (banking, fintech, payments)
- Multicloud exposure and experience with additional Terraform providers
- Go for building operational tooling
What We Do
slice is an RBI-licensed small finance bank focused on reimagining consumer banking in India through simple, transparent, technology-driven products. Its offerings include slice savings accounts, UPI credit cards, slice UPI, and slice business. The company aims to improve how people use money and time while delivering a consumer-friendly banking experience, especially for India’s youth, with an emphasis on digital access, transparency, and customer control.

.png)





