Site Reliability Engineer

Posted 5 Days Ago
Be an Early Applicant
F L Office, Choudwar, Cuttack, Odisha, IND
Hybrid
Junior
Artificial Intelligence • Healthtech • Software • Generative AI
The Role
Design, build, and maintain cloud infrastructure and CI/CD pipelines; manage containerized Kubernetes/Docker environments; automate provisioning with Terraform and Helm; script in Python/Bash; monitor systems, respond to incidents, perform root cause analysis, and document operational procedures to ensure reliability and security.
Summary Generated by Built In
DevOps / Site Reliability Engineer (SRE)
About Triomics:


Triomics is building the agentic AI layer for oncology EHRs. Cancer hospitals spend billions on highly trained staff manually reading unstructured patient records such as pathology reports, clinical notes, genomic panels to power workflows like trial matching, registry curation, visit prep, and quality reporting. We replace that manual work with task-driven AI agents that sit inside the EMR and process records with >95% accuracy, at scale, in real time.

We have grown 10x in the last one year and are processing millions of documents monthly.
About the Role

We are looking for a DevOps / Site Reliability Engineer to help design, build, and maintain scalable, secure, and reliable infrastructure. In this role, you will work closely with engineering teams to streamline deployments, improve system reliability, and ensure our platforms run efficiently in production.

You will be responsible for building automation, managing cloud infrastructure, monitoring systems, and responding to incidents to maintain high availability and performance.

Key Responsibilities
  • Design, implement, and manage cloud-based infrastructure and deployment pipelines.

  • Build and maintain CI/CD pipelines to enable reliable and efficient software delivery.

  • Manage and optimize containerized environments using Kubernetes and Docker.

  • Automate infrastructure provisioning and configuration using Terraform and Helm.

  • Develop and maintain automation scripts using Python and Bash.

  • Monitor system health, performance, and reliability using logging and monitoring tools.

  • Troubleshoot production issues and participate in incident response and root cause analysis.

  • Ensure infrastructure security, network configuration, and system hardening best practices.

  • Collaborate with development teams to improve reliability, scalability, and deployment processes.

  • Maintain clear documentation for infrastructure, processes, and operational procedures.

Requirements
  • 1+ years of experience in DevOps, Site Reliability Engineering, or a related role.

  • Hands-on experience with at least one cloud platform (AWS, Azure, or GCP).

  • Strong experience with Kubernetes, Docker, Jenkins, Terraform, and Helm.

  • Proficiency in Python and Bash scripting.

  • Solid understanding of Linux system administration, networking concepts, and security practices.

  • Experience with monitoring, logging, and incident response systems.

  • Strong communication skills and the ability to document technical processes effectively.

  • Software development experience is a plus.

Nice to Have
  • Experience deploying and scaling AI/ML workloads in production environments.

  • Familiarity with single-tenant deployment models.

  • Experience with MLOps/AIOps platforms such as SageMaker or Kubeflow.

  • Knowledge of chaos engineering and disaster recovery strategies.

  • Experience with cloud cost optimization strategies.

  • Relevant cloud certifications (AWS, Azure, or GCP).

Why Join us ?
  • Impact at scale - the AI you build directly accelerates cancer research and improves patient outcomes worldwide

  • Cutting-edge problems - you’ll work on some of the hardest and most interesting LLM engineering challenges in a highly regulated industry.

  • World-class team - collaborate with experts across AI, engineering, product, and oncology with best-in-industry compensation.

  • Culture that ships - we’re a team that works hard and plays hard (company-sponsored workations in Bali, Sri Lanka, Goa, and more

Perks & Benefits:
  • Lunch Provided at the Office – one less daily decision, one happier employee.

  • Flexible Working Hours – we care about output, not clock-ins.

  • Health Insurance – comprehensive coverage for you and your family.

  • Zomato Meal Benefit – breakfast and dinner can be ordered when you come in early or leave late, because effort deserves fuel.

Skills Required

  • 1+ years of experience in DevOps, Site Reliability Engineering, or a related role
  • Hands-on experience with at least one cloud platform (AWS, Azure, or GCP)
  • Strong experience with Kubernetes
  • Strong experience with Docker
  • Strong experience with Jenkins
  • Strong experience with Terraform
  • Strong experience with Helm
  • Proficiency in Python scripting
  • Proficiency in Bash scripting
  • Solid understanding of Linux system administration
  • Knowledge of networking concepts and security practices
  • Experience with monitoring, logging, and incident response systems
  • Strong communication skills and ability to document technical processes
  • Software development experience
  • Experience deploying and scaling AI/ML workloads in production
  • Familiarity with single-tenant deployment models
  • Experience with MLOps/AIOps platforms such as SageMaker or Kubeflow
  • Knowledge of chaos engineering and disaster recovery strategies
  • Experience with cloud cost optimization strategies
  • Relevant cloud certifications (AWS, Azure, or GCP)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
76 Employees
Year Founded: 2021

What We Do

Triomics is a generative AI platform built for oncology workflows. It helps academic and community cancer centers transform unstructured medical-record data into point-of-care insights, including clinical-trial screening, pre-charting, and clinical-data curation. Its platform uses AI agents to read longitudinal patient records and produce structured, explainable outputs, enabling care teams, research teams, and health systems to act faster on information already in the chart.

Similar Jobs

In-Office or Remote
2 Locations
10000 Employees
In-Office or Remote
2 Locations
10000 Employees

Photon Logo Photon

Software Engineer

Agency • Information Technology
In-Office or Remote
2 Locations
5017 Employees

NVIDIA Logo NVIDIA

Senior Site Reliability Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account