Site Reliability Engineer

Reposted 11 Hours Ago
Be an Early Applicant
2 Locations
In-Office
Junior
Artificial Intelligence • Information Technology • Machine Learning • Cybersecurity
The Role
Own production reliability across cloud environments by managing infrastructure, building zero-downtime CI/CD pipelines, codifying infrastructure with Terraform and Ansible, and driving observability. Define SLIs, SLOs, and error budgets; lead incident response, root-cause analysis, and post-incident reviews; optimize cloud costs and capacity; automate AI-powered operational workflows; and participate in on-call rotations.
Summary Generated by Built In
If it's down, it's on you. If it stays up, that's on you too.

At Techdome, you own production for real Healthcare, FinTech, AI, and SaaS products — not a ticket queue. You'll build pipelines, own incidents, ship zero-downtime releases, and be the person the team trusts at 2am.

What you'll do
  • Keep production available, reliable, and performant across every environment.
  • Manage and optimize cloud environments (AWS, Azure, or GCP).
  • Build zero-downtime CI/CD pipelines using Blue-Green, Rolling, and Canary strategies.
  • Codify infrastructure with Terraform and Ansible.
  • Drive observability with Prometheus, Grafana, ELK, Datadog, and OpenTelemetry.
  • Define and own SLIs, SLOs, and error budgets.
  • Lead incident response, RCA, and post-incident reviews.
  • Optimize cloud cost and plan capacity.
  • Automate operational workflows, including AI-powered alert triage and incident summarization.
  • Join the on-call rotation.

What you bring
  • 2+ years as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer.
  • Hands-on production experience with AWS, Azure, or GCP.
  • Docker and Kubernetes experience under real load.
  • Infrastructure-as-Code expertise (Terraform, Ansible, or equivalent).
  • CI/CD pipelines built from scratch (Jenkins, GitHub Actions, GitLab CI, or similar).
  • Strong Linux, networking, and distributed-systems fundamentals.
  • Scripting skills in Python, Go, or Bash.
  • Experience with Blue-Green, Canary, and Rolling deployments in production.
  • Background in FinTech, Payments, Healthcare, or another high-availability domain.
  • Regular use of AI tools (Copilot, Claude, Cursor, ChatGPT, or similar) to move faster.

Extra credit: You've built AI-powered ops workflows (monitoring, alert triage, incident summarization), and you speak fluent SRE — SLOs, error budgets, chaos engineering.

Why Techdome?
Work across AI, Healthcare, Payments, and SaaS products. Own critical infrastructure from day one. Work directly with founders and senior engineering leadership. Real stakes, fast decisions, real growth.

Skills Required

  • 2+ years of experience as an SRE, DevOps Engineer, Platform Engineer, or Cloud Engineer
  • Hands-on production experience with AWS, Azure, or GCP
  • Production experience with Docker and Kubernetes under real load
  • Infrastructure-as-Code expertise with Terraform, Ansible, or equivalent
  • Experience building CI/CD pipelines from scratch using Jenkins, GitHub Actions, GitLab CI, or similar
  • Strong Linux, networking, and distributed-systems fundamentals
  • Scripting skills in Python, Go, or Bash
  • Production experience with Blue-Green, Canary, and Rolling deployments
  • Background in FinTech, Payments, Healthcare, or another high-availability domain
  • Regular use of AI tools such as Copilot, Claude, Cursor, or ChatGPT
  • Experience building AI-powered operations workflows
  • Experience with SLOs, error budgets, and chaos engineering
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
96 Employees
Year Founded: 2020

What We Do

Techdome is a technology consultancy and IT solutions provider founded in 2020. It helps businesses modernize operations and navigate digital transformation through scalable, end-to-end services spanning planning, building, designing, developing, and launching technology solutions, alongside AI, cloud computing, automation, data engineering, machine learning, custom large language models, and cybersecurity. Its mission is to make complex technology accessible and support business growth across diverse industries and markets.

Similar Jobs

Hybrid
2 Locations
289097 Employees
Hybrid
Hyderabad, Telangana, IND
289097 Employees
Hybrid
Hyderabad, Telangana, IND
289097 Employees

MetLife Logo MetLife

Site Reliability Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Hyderabad, Telangana, IND
43000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account