DevOps / Site Reliability Engineer – Engineering

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Healthtech • Software • Analytics • Business Intelligence
The Role
Own AWS infrastructure through Terraform and build GitHub Actions CI/CD pipelines for Django, Python, Node.js, and React services. Manage production deployments, migrations, observability, incident response, PostgreSQL, Redis, messaging systems, security hardening, and reliability improvements. Develop Python and shell automation, maintain runbooks, and provide detailed shift handovers. The role requires production ownership, on-call participation, autonomy, and either normal or night-shift availability.
Summary Generated by Built In
DevOps / Site Reliability Engineer – Engineering 

Experience: 6+ Years 
Location: Kolkata, India / Gurgaon, India (On-site) 
Shift: Normal or Night — depending on candidate preference and business need Reports to: Lead DevOps / SRE 
Employment Type: Full-time 

About Practice by Numbers 
Practice by Numbers (PbN) is a fast-growing SaaS platform helping healthcare organizations leverage data, automation, patient engagement, and operational intelligence to improve business performance and patient outcomes. We build highly scalable cloud-native products serving thousands of customers across North America. 
We are looking for a DevOps/SRE Engineer who writes Terraform and CI/CD pipelines. This is a builder’s role, not a monitoring seat. 

Role Overview 
You will own Terraform modules across our AWS infrastructure — compute, networking, databases, IAM, DNS — and build the GitHub Actions pipelines that ship our Django, Python, Node.js, and React services. 

This role can be staffed on a normal shift or a night shift, depending on your preference and our coverage needs. On a night shift you get an uninterrupted window for the changes that are hardest to make in daylight — database migrations, pipeline rewrites, provider upgrades, and staged cutovers — and you will often be the only engineer on shift. That means real autonomy, real ownership, and the expectation that you can reason through an unfamiliar system on your own and write up clearly what you did. 

This role suits someone who wants deep infrastructure work with production ownership and prefers building over ticket-shuffling. 

Key Responsibilities 

Infrastructure as Code 
  • Own and extend Terraform modules across AWS — ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, Route 53, CloudWatch. 
  • Manage Terraform state safely; write and review plans that reviewers can trust. 
  • Eliminate manually created resources by bringing them under code (via terraform import, refactors, and module extraction). 
  • Keep environments (dev, QA, production) consistent and reproducible.

CI/CD Engineering 
  • Build and maintain GitHub Actions pipelines for build, test, containerization, and deployment.
  • Migrate legacy pipelines (Jenkins, CircleCI, GitLab CI) onto a single, maintainable platform.
  • Design safe deployment paths — staged rollouts, health-gated releases, fast and reliable rollback. Keep pipelines fast: caching, parallelism, test splitting, and honest gating. 

Reliability & Operations 
  • Own the production change window on your shift: deploys, migrations, cutovers, and maintenance.
  • Run and improve observability — dashboards, SLOs, alerting that fires on real user impact rather than noise. 
  • Debug containerized applications in production: logs, metrics, traces, resource limits, networking.
  • Participate in on-call rotation and incident response; drive blameless post-incident reviews. Improve cost efficiency and resource utilization across the AWS footprint. 

Databases & Data Services 
  • Operate PostgreSQL in production — read query plans, identify slow queries, understand connections, locks, and replication. 
  • Plan and execute schema migrations against live systems with minimal disruption.
  • Operate supporting data services: Redis/ElastiCache, message brokers, object storage. 

Automation & Tooling 
  • Write Python and shell automation to remove repetitive operational work. 
  • Build tooling that makes the wider team faster — self-service scripts, runbooks, guardrails.
  • Harden secrets handling, access control, and infrastructure security posture. 

Handover & Communication 
  • Write clear, complete handover notes at the end of every shift. Where shifts do not overlap, your writing is how the rest of the team learns what happened. 
  • Maintain runbooks and infrastructure documentation as systems change. 
  • Coordinate with the wider engineering team on planned work and follow-ups. 

AI-Enabled Engineering 
  • Use modern AI tools to accelerate infrastructure work, scripting, and troubleshooting.
  • Apply AI-assisted practices while maintaining strong engineering, security, and review standards. 

Required Qualifications 
  • 6+ years in DevOps, SRE, Platform, or Infrastructure engineering with genuine production ownership. 
  • Strong AWS experience — container orchestration (ECS/Fargate or EKS), RDS, VPC networking, IAM, load balancing, S3, CloudWatch. 
  • Hands-on Terraform: writing modules, managing state, reviewing plans.
  • Practical CI/CD experience, ideally GitHub Actions. GitLab CI, CircleCI, or Jenkins backgrounds are fine if you can migrate. 
  • Docker, and comfort debugging containerized applications in production. 
  • Solid Linux fundamentals and shell scripting; Python for automation. 
  • Working PostgreSQL knowledge — query plans, slow queries, connections, locks, replication. Experience deploying and operating Django/Python and Node.js/React applications. Hands-on use of an observability platform in anger (New Relic, Datadog, Grafana, Prometheus, or similar). 
  • Clear written English. Your handover notes are how the rest of the team learns what happened. Openness to either a normal or a night shift, and willingness to participate in an on-call rotation. 

Technical Expertise 

Cloud & Infrastructure 
  • AWS (ECS / Fargate / EKS, EC2, Lambda) 
  • VPC, Subnets, Security Groups, NAT, Peering 
  • ALB / NLB, Route 53, CloudFront 
  • IAM, Roles, Policies, Least-Privilege Design 
  • RDS, ElastiCache, S3 
  • Terraform / Infrastructure as Code 

CI/CD & Automation 
  • GitHub Actions 
  • Jenkins / CircleCI / GitLab CI 
  • Docker, Container Registries 
  • Blue-Green & Staged Deployments, Rollback Strategies 
  • Python, Bash / Shell Scripting 
  • Git and Trunk-Based Workflows 

Observability & Reliability 
  • New Relic / Datadog / Grafana / Prometheus 
  • CloudWatch Metrics, Logs, Alarms 
  • OpenTelemetry 
  • SLOs, Error Budgets, Alert Design 
  • Incident Management & Post-Incident Review 

Data & Messaging 
  • PostgreSQL Administration & Tuning 
  • Schema Migrations on Live Systems 
  • Redis / ElastiCache
  • Kafka / Redpanda / Amazon MSK, RabbitMQ, NATS 

Security & Compliance 
  • Secrets Management (AWS Secrets Manager, SOPS, Vault) 
  • Network and Access Hardening 
  • Vulnerability and Patch Management 
  • Audit Logging 

Preferred Qualifications 
  • Celery, RabbitMQ, NATS, or Kafka in production. 
  • Redis / ElastiCache operations. 
  • Secrets management (AWS Secrets Manager, SOPS, Vault). 
  • Experience with sharded or multi-tenant database architectures. 
  • Healthcare or other compliance-sensitive environments (HIPAA, SOC 2). 
  • Telephony / VoIP infrastructure exposure. 
  • Kubernetes. 
  • Prior experience as the sole engineer on shift, or on a night/off-hours rotation. 

What We Look For 
  • A builder’s instinct — you would rather codify a fix than repeat it. 
  • Comfort with autonomy and sound judgment on when to escalate. 
  • Careful, methodical change management on production systems. 
  • Strong written communication; your handover notes are a first-class deliverable.
  • Curiosity about how systems actually behave, not just how they are supposed to. Ownership, follow-through, and continuous learning. 

Why This Role 
  • Real ownership of infrastructure, not a ticket queue. 
  • Flexibility on shift, and a protected change window for high-impact work. 
  • Modern stack: AWS, Terraform, GitHub Actions, Docker, PostgreSQL, Kafka-compatible messaging.
  • Direct impact on the reliability of a platform used by thousands of healthcare practices.


Skills Required

  • 6+ years of experience in DevOps, SRE, Platform, or Infrastructure Engineering with genuine production ownership
  • Strong AWS experience, including ECS/Fargate or EKS, RDS, VPC networking, IAM, load balancing, S3, and CloudWatch
  • Hands-on Terraform experience writing modules, managing state, and reviewing plans
  • Practical CI/CD experience, ideally with GitHub Actions; experience with Jenkins, CircleCI, or GitLab CI is acceptable
  • Docker experience and ability to debug containerized applications in production
  • Solid Linux fundamentals and shell scripting skills
  • Python experience for automation
  • Working PostgreSQL knowledge, including query plans, slow queries, connections, locks, and replication
  • Experience deploying and operating Django/Python and Node.js/React applications
  • Hands-on experience with an observability platform such as New Relic, Datadog, Grafana, or Prometheus
  • Clear written English
  • Availability for normal or night shifts and participation in an on-call rotation
  • Production experience with Celery, RabbitMQ, NATS, or Kafka
  • Redis or ElastiCache operations experience
  • Secrets management experience with AWS Secrets Manager, SOPS, or Vault
  • Experience with sharded or multi-tenant database architectures
  • Experience in healthcare or other compliance-sensitive environments, including HIPAA or SOC 2
  • Telephony or VoIP infrastructure exposure
  • Kubernetes experience
  • Experience as the sole engineer on shift or on a night/off-hours rotation
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
87 Employees
Year Founded: 2015

What We Do

Practice by Numbers is an all-in-one software solution designed to help dental practices, groups, and organizations consolidate, streamline, and improve their day-to-day operations. Its mission is to transform dental practice management through integrated software solutions that enhance patient experiences and optimize business performance, providing a comprehensive suite of analytics, patient communication, and reputation management tools in a single platform.

Similar Jobs

Expedia Group Logo Expedia Group

Development Engineer

AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Hybrid
Gurugram, Haryana, IND
16000 Employees

Expedia Group Logo Expedia Group

Development Engineer

AdTech • eCommerce • Information Technology • Software • Travel • Generative AI
Hybrid
Gurugram, Haryana, IND
16000 Employees

Mastercard Logo Mastercard

Lead Data Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Gurugram, Haryana, IND
38800 Employees

Optum Logo Optum

Consultant

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Gurgaon, Gurugram, Haryana, IND
160000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account