Senior Infrastructure SRE

Posted 4 Days Ago
Mississauga, ON, CAN
Hybrid
139K-155K Annually
Senior level
Healthtech • Software
The Role
Designs and operates resilient multi-cloud and on-premises infrastructure across Azure, AWS, and GCP. Responsibilities include Infrastructure as Code, automation, Kubernetes, service mesh, identity, storage, observability, SRE metrics, incident response, and reliability improvements. The role participates in on-call rotations, leads complex incidents, drives multi-team infrastructure initiatives, mentors SREs, and applies AI-assisted tooling to reduce operational toil.
Summary Generated by Built In
At PointClickCare our mission is simple: to help providers deliver exceptional care. And that starts with our people. As a leading health tech company that’s founder-led and privately held, we empower our employees to push boundaries, innovate, and shape the future of healthcare.

With the largest long-term and post-acute care dataset and a Marketplace of 400+ integrated partners, our platform serves over 30,000 provider organizations, making a real difference in millions of lives. We also reinvest a significant percentage of our revenue back into research and development, ensuring our employees have the resources to innovate and make a lasting impact. Recognized by Forbes as a top private cloud company and honored as one of Canada’s Most Admired Corporate Cultures, we offer flexibility, growth opportunities, and meaningful work. 

At PointClickCare, we empower our people to be the architects of a smarter healthcare future; one that is human-first and accelerated by AI to create meaningful and lasting change. Employees harness AI as a catalyst for creativity, productivity, and thoughtful decision-making. By integrating AI tools into our daily workflows, collaboration is enhanced, outcomes are improved, and every team member has the proficiency to maximize their impact. It all starts with our hiring practices where we uncover AI expertise that complements our mission, and we continue to invest in training and development to nurture innovation throughout the employee journey.

Join us in redefining healthcare — so it doesn’t just survive, it thrives. To learn more about PointClickCare, check out Life at PointClickCare and connect with us on Glassdoor and LinkedIn.


**Travel to Office expectations**
For Remote Roles: If this role is remote, there will be in-office events that will require travel to and from the Mississauga and/or Salt Lake City office. These will include, but not limited to, onboarding, team events, semi-annual and annual team meetings.

For Hybrid Roles: If this role is Hybrid, there will be an expectation to reside within commutable distance to the office/location specified in the job listing. This will include, but not limited to, weekly/bi-weekly/monthly events in the office with your specific team. This is a requirement for this role.

About the role

PointClickCare builds cloud platforms that power safer, more connected care for millions of patients. You'll design and operate resilient infrastructure services across Azure, AWS, GCP, and on-prem environments, driving reliability, automation, and operational excellence for identity, compute, storage, messaging, and shared services with an SRE-first mindset.

What you'll do
  • Design and implement highly available infrastructure solutions for compute, storage, identity, messaging, and shared services
  • Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
  • Automate operational workflows to eliminate toil: auto-remediation, self-healing systems, capacity planning
  • Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
  • Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
  • Develop observability strategies: metrics, logs, distributed tracing, alerting frameworks
  • Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis, alert triage, runbook and post-mortem drafting, automation scaffolding
  • Own reliability for one or more infrastructure domains end to end, driving multi-team initiatives with product engineering from problem definition through adoption
  • Mentor intermediate SREs; review infrastructure changes; establish operational best practices
What you'll bringMust-haves
  • 5+ years of hands-on experience operating and designing cloud infrastructure
  • Expert-level understanding of the core services of either Azure or AWS
  • Working proficiency in at least one additional platform (Azure, AWS, or GCP)
  • Experience designing and supporting production infrastructure that spans multiple cloud platforms
  • 3+ years of production experience with Infrastructure as Code (Terraform, Pulumi, CloudFormation)
  • Ability to design scalable, reusable IaC modules and enforce GitOps workflows
  • Experience managing IaC across multiple cloud providers, including module design, state layout, and provider-specific resource differences
  • Strong proficiency in at least one programming language (Python, Go, Bash) for production automation
  • Demonstrated ability to write tested, maintainable automation and tooling
  • Practical application of SRE principles in production environments
  • Experience defining and managing SLIs/SLOs, error budgets, toil metrics
  • Track record of improving system reliability (e.g., MTTR reduction, availability improvements)
  • Demonstrated depth across key infrastructure services:
  • Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes (AKS, EKS, or equivalent), plus VM-based compute
  • Practical experience operating a service mesh in production (Istio preferred; Linkerd or equivalent)
  • Strong proficiency with enterprise identity and SSO: SAML, OAuth/OIDC, LDAP, and cloud IAM
  • Working knowledge of an enterprise federation platform such as PingFederate, Entra ID, Okta, or ADFS
  • Strong proficiency in storage solutions (object, block, and file storage; Kubernetes persistent volumes)
  • Working knowledge of designing and operating infrastructure in a regulated environment (HIPAA, SOC 2, PCI, FedRAMP, or equivalent)
  • Familiarity with audit evidence, access controls, encryption in transit and at rest, and data residency constraints
  • Proven track record of measurably reducing operational toil through automation (e.g. ticket volume, manual runbook executions, hours reclaimed)
  • Strong communication and documentation skills; demonstrated ability to influence engineering teams
Nice-to-haves
  • 2+ years in healthcare technology or highly regulated SaaS environments (HIPAA, SOC 2, HITRUST)
  • Cloud certifications: Azure Solutions Architect Expert, AWS Solutions Architect Professional, GCP Professional Cloud Architect, or equivalent
  • Working experience operating Kubernetes at scale across multiple clusters (CKA or CKAD certification a plus)
  • Working knowledge of messaging and event-streaming platforms (Kafka, Azure Service Bus, Event Hubs, SQS, Pub/Sub)
  • Practical experience building CI/CD pipelines and deployment automation (GitLab, GitHub Actions, ArgoCD)
  • Familiarity with AI-assisted engineering and operations tooling: LLM-based coding assistants, AIOps, agentic incident investigation
  • Judgment about where AI belongs in an operational workflow, including safe handling of sensitive data in prompts and reviewing generated changes before they reach production contribution to open-source SRE tools or infrastructure projects (GitHub profile, PRs merged)
Education
  • Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field
  • OR equivalent practical experience with a proven track record in infrastructure and SRE practices evidence of continuous learning and staying current with SRE and cloud-native trends
How we work
Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field
  • Transparent collaboration: We work in the open using OKRs, cross-functional retrospectives, and public roadmaps so everyone knows priorities and progress
  • Blameless culture: We confront problems courageously through structured post-mortems and root cause analyses, focusing on systems improvement not individual blame
  • Data-driven decisions: We use metrics, APM, and observability data to make evidence-based choices about reliability and performance investments
  • Continuous learning: We learn from incidents through retrospectives and RCAs, sharing knowledge across teams to prevent repeat issues
  • Outcome accountability: We're accountable for customer and business results, not just completing tasks, measuring success by impact
  • Thoughtful experimentation: We hold strong opinions loosely, testing concepts and running small experiments before scaling solutions
  • Iterative delivery: We think big but act small, using Scrum, delivery plans, and frequent milestones to ship incrementally and learn fast
  • Inclusive environment: We create space to listen and learn, actively growing our Ally Community to support equity, belonging, and career growth for all
#LI-AV1
#LI-hybrid 

PointClickCare Benefits & Perks:

Benefits starting from Day 1!
Retirement Plan Matching
Flexible Paid Time Off
Wellness Support Programs and Resources
Parental & Caregiver Leaves
Fertility & Adoption Support
Continuous Development Support Program
Employee Assistance Program
Allyship and Inclusion Communities
Employee Recognition … and more!

It is the policy of PointClickCare to ensure equal employment opportunity without discrimination or harassment on the basis of race, religion, national origin, status, age, sex, sexual orientation, gender identity or expression, marital or domestic/civil partnership status, disability, veteran status, genetic information, or any other basis protected by law. PointClickCare welcomes and encourages applications from people with disabilities. Accommodations are available upon request for candidates taking part in all aspects of the selection process. Please contact [email protected] should you require any accommodations. As part of our commitment to a streamlined and equitable hiring experience, PointClickCare uses AI tools to assist with candidate screening and assessment.

When you apply for a position, your information is processed and stored with Lever, in accordance with Lever’s Privacy Policy. We use this information to evaluate your candidacy for the posted position. We also store this information, and may use it in relation to future positions to which you apply, or which we believe may be relevant to you given your background. When we have no ongoing legitimate business need to process your information, we will either delete or anonymize it.  If you have any questions about how PointClickCare uses or processes your information, or if you would like to ask to access, correct, or delete your information, please contact PointClickCare’s human resources team: [email protected] 

PointClickCare is committed to Information Security. By applying to this position, if hired, you commit to following our information security policies and procedures and making every effort to secure confidential and/or sensitive information.

Skills Required

  • 5+ years of hands-on experience operating and designing cloud infrastructure
  • Expert-level knowledge of either Azure or AWS
  • Working proficiency with at least one additional cloud platform: Azure, AWS, or GCP
  • Experience designing and supporting production infrastructure across multiple cloud platforms
  • 3+ years of production experience with Infrastructure as Code using Terraform, Pulumi, or CloudFormation
  • Ability to design reusable Infrastructure as Code modules and enforce GitOps workflows
  • Experience managing Infrastructure as Code across multiple cloud providers
  • Strong proficiency in Python, Go, or Bash for production automation
  • Experience writing tested, maintainable automation and tooling
  • Practical production experience applying SRE principles
  • Experience defining and managing SLIs, SLOs, error budgets, and toil metrics
  • Demonstrated track record of improving system reliability, such as reducing MTTR or improving availability
  • Expert-level experience operating Kubernetes and containerized workloads in production on AKS, EKS, or equivalent
  • Production experience operating VM-based compute
  • Practical experience operating a production service mesh; Istio preferred
  • Strong proficiency with enterprise identity and SSO, including SAML, OAuth/OIDC, LDAP, and cloud IAM
  • Working knowledge of PingFederate, Entra ID, Okta, ADFS, or an equivalent enterprise federation platform
  • Strong proficiency with object, block, and file storage and Kubernetes persistent volumes
  • Working knowledge of infrastructure design and operations in regulated environments
  • Familiarity with audit evidence, access controls, encryption in transit and at rest, and data residency constraints
  • Proven ability to reduce operational toil through automation
  • Strong communication and documentation skills with the ability to influence engineering teams
  • Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or a related technical field, or equivalent practical experience
  • 2+ years of experience in healthcare technology or highly regulated SaaS environments
  • Cloud certification such as Azure Solutions Architect Expert, AWS Solutions Architect Professional, or GCP Professional Cloud Architect
  • Experience operating Kubernetes at scale across multiple clusters
  • CKA or CKAD certification
  • Working knowledge of Kafka, Azure Service Bus, Event Hubs, SQS, or Pub/Sub
  • Practical experience building CI/CD pipelines with GitLab, GitHub Actions, or ArgoCD
  • Familiarity with AI-assisted engineering and operations tooling
  • Contribution to open-source SRE tools or infrastructure projects

PointClickCare Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about PointClickCare and has not been reviewed or approved by PointClickCare.

  • Healthcare Strength Health and dental coverage appear robust, with wellness and assistance programs reinforcing core medical benefits. Coverage quality stands out relative to other benefit elements.
  • Leave & Time Off Breadth PTO and paid holidays are characterized as generous, and flexible work-from-home options are widely available. Occasional extras like summer half‑day Fridays further expand time-off flexibility.
  • Flexible Benefits A customizable mix is evident through remote/hybrid arrangements, day-one eligibility, and a lifestyle or personal spending account. Benefits such as wellness credits and support resources can be tailored to individual needs.

PointClickCare Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Toronto
1,557 Employees
Year Founded: 2000

What We Do

PointClickCare is the market leader driving the transformation of healthcare vulnerable and complex populations through a broad, connected care network powered by deep insights with a commitment to value, outcomes and innovation. We connect post-acute and acute care settings, people and systems like no other company. Our steadfast commitment to our culture and to providing growth opportunities to our employees is evidenced by recent recognition of PointClickCare as one of Canada’s best-managed companies and most admired corporate cultures.

Similar Jobs

Hybrid
Mississauga, ON, CAN
1557 Employees
139K-155K Annually

Tapestry - Coach and Kate Spade Logo Tapestry - Coach and Kate Spade

Sales Associate III

eCommerce • Fashion • Retail • Sales • Wearables • Design
Hybrid
Scarborough, ON, CAN
16000 Employees
18-22 Hourly

Superhuman Logo Superhuman

Software Engineering Intern - Summer 2027

Artificial Intelligence • Information Technology • Machine Learning • Natural Language Processing • Productivity • Software • Generative AI
Hybrid
Toronto, ON, CAN
1500 Employees
50-50 Hourly

Block Logo Block

Director, Revenue Strategy & Planning

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
218K-327K Annually

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account