The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.
This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.
Responsibilities:
- Collaborate with our Security Operations teams to define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
- Develop, implement, and coordinate a multi-tenant approach around service offerings for databases, container platforms, authentication, certificates, and product registries.
- Design, develop, and maintain cloud performance testing strategies, frameworks, and environments to validate application scalability, reliability, and resiliency.
- Develop and automate load, stress, spike, and endurance (soak) testing as part of CI/CD pipelines.
- Analyze application and infrastructure performance to identify bottlenecks and recommend performance optimizations across cloud-native services.
- Develop and maintain cost and utilization tracking and attribution processes across Cloud Service Providers.
- Create documentation detailing Cloud Service Provider offerings, implementation patterns, and best practices.
- Develop and maintain technical relationships with our core Cloud Service Providers.
- Implement and maintain secure, scalable infrastructure platforms for delivering cloud services.
- Ensure internal and external SLAs are consistently met or exceeded, while continuously monitoring and improving system performance, reliability, and availability.
- Create tools for automating deployment, monitoring, and platform operations.
- Implement and manage observability solutions (logging, metrics, tracing) using OpenTelemetry, Prometheus, Grafana, Azure Monitor, and related technologies to provide actionable performance insights.
- Plan and execute chaos engineering experiments to evaluate and improve application resiliency and fault tolerance.
Requirements:
- 5+ years of experience with Cloud Service Providers and best practices around implementation and configuration, preferably managing Azure environments supporting SaaS products.
- Experience working across multiple cloud providers (Azure required; AWS and/or Google Cloud Platform considered an asset).
- Strong experience in Cloud Performance Engineering, including performance analysis, capacity planning, scalability testing, and optimization of distributed cloud-native applications.
- Proven experience working with microservices architecture, with a strong focus on Java-based services.
- Experience applying Chaos Engineering practices to evaluate and improve system resiliency.
- Strong experience designing and executing performance testing strategies, including load, stress, spike, and endurance (soak) testing, to validate application scalability and defined latency and error-rate thresholds.
- Hands-on experience with performance testing tools such as JMeter, Gatling, Azure Load Testing, or k6.
- Experience validating application services sustaining 500+ transactions per second (TPS) while meeting defined performance objectives.
- Hands-on experience deploying and managing containerized applications using Docker and Kubernetes, including autoscaling and performance optimization.
- Experience using Terraform to provision and manage cloud infrastructure using Infrastructure as Code (IaC).
- Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing platforms to sustain target transaction rates.
- Hands-on experience implementing and using observability platforms including OpenTelemetry, Prometheus, Grafana, Azure Monitor, Application Insights, and Log Analytics.
- Proven experience with Security and Compliance (SOC 2, HIPAA, ISO 27001) best practices and implementing controls that support high-velocity software delivery teams.
Skills Required
- Expertise in cloud service providers and cloud implementation best practices (multi-cloud)
- Experience managing Azure on behalf of multiple teams (preferred)
- Proven experience with microservices architecture and Java-based services
- Experience applying chaos engineering practices
- Skilled in troubleshooting performance issues and recommending optimizations
- Familiarity with performance testing methodologies and tools
- Experience implementing and operating Prometheus, Grafana, and OpenTelemetry (Otel)
- Proven experience designing and executing performance test plans (load, stress, soak, spike) to validate 500+ TPS
- Hands-on experience with JMeter, Gatling, and Azure Load Testing
- Experience tuning and validating autoscaling (Kubernetes/OpenShift HPA, Azure scale sets)
- Experience tuning Kafka and messaging/queueing components for throughput/latency
- Experience with Azure-native monitoring and diagnostics (Azure Monitor, Application Insights, Log Analytics)
- Proven experience with Security and Compliance (SOC2, HIPAA, ISO27001) controls implementation
- Proficiency with Infrastructure as Code tools (Terraform, Ansible or Chef)
- Experience operating and maintaining production systems in Linux and public cloud environments
- Experience with on-call incident management, troubleshooting, and escalation processes
- Proven ability to prioritize and track multiple projects in parallel
What We Do
Smile Digital Health (doing business as Smile CDR Inc.) specializes in delivering fast, secure, compliant data infrastructures as a service to enable and empower interconnectivity for data-intensive sectors such as healthcare. We are a solutions platform helping organizations like governments, researchers, health systems, healthcare providers, and app developers build connected health solutions and products by leveraging our core expertise in health data and HL7 FHIR. Our flagship product, Smile CDR, is the world’s first FHIR-based clinical data repository (CDR) as a service. Smile CDR is a high performance and secure solution built on the principles of Privacy by Design. Flexibility is a hallmark of the service as it can be hosted in the cloud or on-premises depending on your needs. The design was strongly influenced by real-world experience in both the jurisdictional and organizational setting. Smile CDR provides a rich set of features and capabilities including: -Multiple FHIR versions -FHIR Profiles -Full text indexing of clinical records -Type ahead search functionality -FHIR, HL7v2 and custom ETL for data input -Federated identity and identity provider functionality -International character locale support -Terminology services -Auditing -Smart on FHIR support -Rapid deployment -Extensive administrative tooling Leveraging more than 40 years of experience in building enterprise-class systems, our team includes experts in building integrated healthcare systems including FHIR-based solutions such as HAPI, e-prescriptions and other e-health solutions. We were motivated by the need to improve upon existing options for sharing health data within and across organizations. Smile CDR is the maintainer of HAPI FHIR, the prevailing open source reference implementation of FHIR worldwide








