Senior Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
3 Locations
Remote or Hybrid
Senior level
Healthtech
The Role
Own the reliability, availability, security, and performance of cloud software deployed at client sites. Define SLOs, build observability and automation, lead incident response, conduct capacity planning, secure infrastructure and CI/CD pipelines, support deployments, and maintain audit trails. Provide technical leadership, mentor engineers, and collaborate with operational teams and customers. The role is remote-friendly in France or the Netherlands with occasional European travel.
Summary Generated by Built In

Job Summary

The Senior Site Reliability Engineer plays a vital role in ensuring the reliability, availability, and performance of DeepHealth software applications, which integrate AI algorithms to deliver clinically relevant information for enhanced decision support. This role takes ownership of the reliability of cloud components deployed at client sites, ensuring that the DeepHealth solution is scalable, resilient, and secure, providing support to the operational team, and providing technical leadership within the platform engineering practice.

 

Essential Duties and Responsibilities 

  • Own the reliability, availability, and performance of cloud components deployed at client sites and of the DeepHealth solution.

  • Define and monitor service level objectives (SLOs), error budgets, and key reliability metrics.

  • Develop and implement automation tools and processes to eliminate toil and streamline deployment, monitoring, and incident response operations.

  • Design and maintain observability tooling (monitoring, logging, alerting, and tracing), and resolve issues before they impact clients.

  • Contribute to the writing of technical specifications and documentation, ensuring compliance with regulatory requirements and industry best practices.

  • Lead incident response, conduct blameless post-mortems, and drive the implementation of corrective and preventive actions.

  • Perform capacity planning and performance tuning to anticipate growth and ensure optimal resource utilization.

  • Support the deployment of software solutions at customer sites, ensuring smooth implementation and optimal performance.

  • Provide ongoing support and maintenance for deployed solutions and to the operational team, addressing any issues or challenges promptly to maintain high levels of customer satisfaction.

  • Mentor and onboard engineers, providing technical leadership in SRE and platform engineering best practices.

  • Implement and maintain secure infrastructure configurations per approved baselines.

  • Ensure CI/CD pipeline security, including integrity verification and access controls.

  • Perform day-to-day technical security controls including system hardening and log monitoring.

  • Document all infrastructure changes and maintain audit trails.

 

PLEASE NOTE: This is not an exhaustive list of all duties, responsibilities and requirements of the position described above.  Other functions may be assigned and management retains the right to add or change duties at any time.

Minimum Qualifications, Education and Experience

  • Fluency in English, both written and spoken.

  • Significant hands-on experience (7+ years) in Site Reliability Engineering or DevOps.

  • In-depth knowledge of software development practices, including design, implementation, testing, and deployment.

  • Strong knowledge of networking and security best practices.

  • Experience with fleet management and GitOps (e.g., ArgoCD).

  • Proficiency in containerization technologies, such as Docker and Kubernetes and its ecosystem (e.g., Istio, KEDA), and with Linux systems.

  • Experience with virtualization technologies, automation tools (such as Ansible or Terraform) and with SQL database administration.

  • Proficiency with continuous integration and continuous deployment (CI/CD) pipelines.

  • Strong proficiency with cloud-based environments (GCP or AWS).

  • Strong experience with observability and monitoring tools (e.g., Prometheus, Grafana).

  • Excellent communication skills, both written and verbal, and strong problem-solving and analytical skills.

  • Strong technical leadership and mentoring abilities

Preferred:

  • Familiarity with the medical device industry and the specific requirements for software applications within this domain.

  • Understanding of AI and machine learning concepts, with the ability to integrate algorithms into software applications.

  • Experience with information security standards (e.g., ISO 27001, SOC 2).

Travel

Occasional travel may be required (typically less than 10%), primarily within Europe, for audits, customer meetings, partner / vendor visits, or company offsites.

Working Environment

  • France or The Netherlands – Remote-friendly.

  • The role can be based in France or The Netherlands with flexible remote working arrangements.

  • There are offices in Paris, Amsterdam and Rotterdam.

  • Periodic on-site presence may be required for team meetings, audits, or workshops.

Skills Required

  • Fluency in written and spoken English
  • 7+ years of hands-on Site Reliability Engineering or DevOps experience
  • Knowledge of software development practices, including design, implementation, testing, and deployment
  • Strong knowledge of networking and security best practices
  • Experience with fleet management and GitOps, such as ArgoCD
  • Proficiency with Docker, Kubernetes, and related technologies such as Istio and KEDA
  • Proficiency with Linux systems
  • Experience with virtualization technologies and automation tools such as Ansible or Terraform
  • Experience with SQL database administration
  • Proficiency with CI/CD pipelines
  • Strong experience with cloud environments, particularly GCP or AWS
  • Strong experience with observability and monitoring tools such as Prometheus and Grafana
  • Excellent written and verbal communication skills
  • Strong problem-solving and analytical skills
  • Technical leadership and mentoring abilities
  • Familiarity with the medical device industry and related software requirements
  • Understanding of AI and machine learning concepts and algorithm integration
  • Experience with information security standards such as ISO 27001 and SOC 2
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Pak Shek Kok
306 Employees

What We Do

DeepHealth is a wholly-owned subsidiary of RadNet, Inc. (NASDAQ: RDNT) and serves as the umbrella brand for all companies within RadNet’s Digital Health segment. DeepHealth provides AI-powered health informatics with the aim of empowering breakthroughs in care through imaging. Building on the strengths of the companies it has integrated and is rebranding (i.e., eRAD Radiology Information and Image Management Systems and Picture Archiving and Communication System, Aidence lung AI, DeepHealth and Kheiron breast AI and Quantib prostate and brain AI), DeepHealth leverages advanced AI for operational efficiency and improved clinical outcomes in lung, breast, prostate, and brain health. At the heart of DeepHealth’s portfolio is a cloud-native operating system – DeepHealth OS – that unifies data across the clinical and operational workflow and personalizes AI-powered workspaces for everyone in the radiology continuum. Thousands of radiologists at hundreds of imaging centers and radiology departments around the world use DeepHealth solutions to enable earlier, more reliable, and more efficient disease detection, including in large-scale cancer screening programs. DeepHealth’s human-centered, intuitive technology aims to push the boundaries of what’s possible in healthcare.

Similar Jobs

Autodesk Logo Autodesk

Senior Site Reliability Engineer

Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
In-Office or Remote
28 Locations
13285 Employees
38K-55K Annually
In-Office or Remote
33 Locations
200 Employees

Planet Logo Planet

Senior Site Reliability Engineer

Aerospace • Big Data • Greentech • Hardware • Social Impact
In-Office or Remote
3 Locations
747 Employees
77K-96K Annually
Remote
5 Locations
154 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
OneImaging Thumbnail
Healthtech
Miami, FL
62 Employees
Vitalize Thumbnail
Artificial Intelligence • Healthtech • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account