Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office or Remote
80K-110K Annually
Mid level
Software • PropTech
Tools, data, and services to help Canadian rental housing professionals market smarter, lease faster, and manage better.
The Role
Responds to production incidents across AWS and Kubernetes, diagnosing and remediating infrastructure, networking, database, cache, and deployment issues. Builds monitoring, observability, synthetic and load tests, SLOs, and reliability improvements. Maintains Terraform, Kubernetes, CI/CD, IAM, autoscaling, and security configurations; leads postmortems and incident follow-up. Participates in on-call operations and collaborates with engineering teams to improve application performance and reliability.
Summary Generated by Built In

About Rentsync
Rentsync is an award-winning, high-growth organization that provides high quality websites, marketing services, and software solutions to the rental and property management industry throughout Canada and the United States.


About the role

We’re looking for a hands-on Site Reliability Engineer to lead our response to production incidents. You’ll dig into our AWS and Kubernetes environments, find the root cause, and fix it.


Between incidents, you’ll make sure the same problem doesn’t happen twice by hardening infrastructure, improving monitoring, and working with engineering teams on performance and reliability.


Our environment is bigger and more varied than most companies our size: 10+ products and 100+ services across multiple Kubernetes clusters, mainly on AWS with some Azure and GCP, built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, Python, and Rust.


Technologies you’ll work with: AWS (EKS, EC2, RDS, S3, ALB/NLB, CloudWatch), Azure, GCP, Kubernetes, Terraform, Ansible, GitHub Actions/GitLab CI, PagerDuty, Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry, Cloudflare, Ubuntu & Amazon Linux, MySQL & PostgreSQL, Redis & Memcached, NGINX & Traefik, Bash, Python, and applications built in PHP, Ruby on Rails, JavaScript/TypeScript, .NET, and Rust.


This is a remote position. All qualified candidates are encouraged to apply, however preferential consideration may be given to individuals who are within a reasonable commuting distance of one of our offices.


Duties & Responsibilities

Incident response & first-contact remediation (primary focus)

  • Be the first responder for production alerts and incidents across our services, and take them from triage through to resolution.
  • Diagnose and fix issues directly in AWS (EKS, EC2, RDS, networking, IAM) and Kubernetes, such as failing pods, resource exhaustion, bad deploys, networking/DNS, database and cache problems.
  • Roll back, scale, reconfigure, or patch infrastructure to restore service fast; escalate to development teams only when a code change is truly needed, and give them a clear diagnosis when you do.
  • Own our PagerDuty setup and incident response during business hours, driving down MTTD and MTTR.
  • Run blameless post-mortems and personally drive the technical follow-up work, not just the action-item list.
  • Automate runbooks and repetitive operational work, including using AI tools to speed up triage, investigation, and remediation.


Reliability engineering (preventing the next incident)

  • Build and maintain monitoring for Kubernetes workloads and services (Prometheus/Mimir, Loki, Tempo, Grafana, OpenTelemetry), with low-noise, high-signal alerts.
  • Watch new releases in production and catch regressions in latency, errors, or resource use before they become incidents.
  • Create and maintain production test suites: synthetic checks, smoke tests, health checks, and load/performance tests.
  • Partner with engineering teams to find and fix performance and reliability issues, and define SLOs, SLIs, and error budgets.
  • Harden our platform after incidents: Terraform changes, Kubernetes resource tuning, autoscaling, CI/CD checks, secrets and IAM improvements.
  • Keep service docs and architecture decisions current so any engineer can operate our systems.


Required Knowledge, Skills & Abilities

  • Infrastructure as code with Terraform, and comfort working in CI/CD pipelines.
  • Solid Linux, networking, and container fundamentals.
  • Scripting/automation in Bash, Python, or similar.
  • Calm, clear communication during incidents and across teams


Essential Qualifications

  • 3+ years in a cloud engineering, DevOps, or SRE role supporting production web applications.
  • Hands-on incident response experience where you diagnosed and fixed production issues yourself, not only coordinated them.
  • Strong, hands-on AWS experience in production (EKS, EC2, RDS, VPC networking, IAM, CloudWatch).
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring Kubernetes workloads.
  • A track record of working with engineering teams to identify and resolve performance and reliability issues.
  • Experience with monitoring and observability tools (e.g. Prometheus, Grafana, Loki, Datadog, CloudWatch) and on-call/alerting tools such as PagerDuty.
  • Experience building automated tests or checks for production reliability (synthetics, smoke, health, or load testing).
  • Willingness to take part in an after-hours on-call rotation as we introduce one in the future.


Additional Preferred Qualifications

  • Using AI tools to accelerate SRE work, such as incident triage, log and metric analysis, runbook automation, or infrastructure code.
  • Azure experience (GCP is a plus too).
  • Supporting many tech stacks across multiple teams (PHP, Ruby on Rails, .NET, Python, Rust, JavaScript).
  • AWS certification (e.g. Solutions Architect, DevOps Engineer, or SysOps).
  • Running LGTM (Loki, Grafana, Tempo, Mimir) or OpenTelemetry at scale.
  • Load testing tools such as k6, Locust, or JMeter.
  • MySQL/PostgreSQL operations and Redis/Memcached tuning.
  • Cloudflare (WAF, Workers, Zero Trust).
  • Cloud cost optimization and capacity planning.
  • Preferential consideration may be given to individuals who are within a reasonable commuting distance of one of our offices.


Rentsync is an equal opportunity employer. If you are selected to participate in the interview process and require unique accommodations, please don’t hesitate to let us know.

Successful candidates may be required to complete a criminal background check in the final phase of the interview process.

This is an open-ended job posting and may not represent a specific vacancy within the organization.

Rentsync reserves the right to use Artificial Intelligence to screen and/or assess candidates.

Skills Required

  • 3+ years of experience in cloud engineering, DevOps, or SRE supporting production web applications
  • Hands-on production incident response experience diagnosing and resolving issues
  • Strong hands-on AWS experience, including EKS, EC2, RDS, VPC networking, IAM, and CloudWatch
  • Deep production Kubernetes experience, including troubleshooting, debugging, and monitoring workloads
  • Experience identifying and resolving performance and reliability issues with engineering teams
  • Experience with monitoring, observability, on-call, and alerting tools such as Prometheus, Grafana, Loki, Datadog, CloudWatch, and PagerDuty
  • Experience building automated production reliability tests or checks, including synthetic, smoke, health, or load tests
  • Willingness to participate in a future after-hours on-call rotation
  • Terraform infrastructure as code and CI/CD pipeline experience
  • Linux, networking, and container fundamentals
  • Bash, Python, or similar scripting and automation experience
  • Calm and clear communication during incidents and across teams
  • Experience using AI tools to accelerate SRE work
  • Azure experience; GCP experience is a plus
  • Experience supporting multiple technology stacks across PHP, Ruby on Rails, .NET, Python, Rust, and JavaScript
  • AWS certification such as Solutions Architect, DevOps Engineer, or SysOps
  • Experience operating LGTM components or OpenTelemetry at scale
  • Experience with k6, Locust, or JMeter load testing tools
  • MySQL or PostgreSQL operations and Redis or Memcached tuning
  • Cloudflare experience with WAF, Workers, or Zero Trust
  • Cloud cost optimization and capacity planning experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: St. Catharines
105 Employees
Year Founded: 2010

What We Do

Rentsync equips Canada’s rental housing professionals with the tools, data, and services they need to market smarter, lease faster, and manage more efficiently. Our ecosystem includes Rentals.ca: Canada’s leading online rental marketplace, and Building Stack: the leasing and resident management platform that powers operational success across all portfolio sizes. From listing syndication, lead automation and lease signing to digital advertising, brand development, and custom websites, Rentsync supports the entire list-to-lease (and beyond) lifecycle. Our expert Agency Services team works as an extension of yours, delivering strategic digital campaigns and creative execution tailored to your leasing goals. With built-in performance reporting and a suite of Rental Market Intelligence products and services, we give clients the data insights they need to act quickly and optimize strategy at every stage. List. Lease. Live. All with Rentsync.

Similar Jobs

GitLab Logo GitLab

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
Remote
3 Locations
2500 Employees
223K-380K Annually

Raydar Logo Raydar

Site Reliability Engineer

Professional Services • Consulting
Remote
2 Locations
28 Employees
160K-210K Annually

MLabs Logo MLabs

Site Reliability Engineer

Artificial Intelligence • Blockchain • Information Technology • Consulting
Remote
5 Locations
90K-120K Annually

Inviso Corporation Logo Inviso Corporation

Site Reliability Engineer

Cloud • Information Technology • Business Intelligence • Consulting
Remote
CA
265 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account