Staff Database Reliability Engineer, Open Internet

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Canada
Remote
180K-250K Annually
Senior level
Artificial Intelligence • Fintech • Professional Services • Software
The Role
Design and operate globally distributed database systems across bare-metal and cloud environments. Lead reliability, performance, automation, observability, incident response, capacity planning, security, and database change management. Establish technical direction, simplify the infrastructure footprint, develop data-platform roadmaps, troubleshoot complex distributed-system failures, and mentor engineers across teams.
Summary Generated by Built In
Your Opportunity

Our client builds a high throughput system powering hundreds of billions of daily transactions, each completed within milliseconds, across globally distributed infrastructure designed for reliability and efficiency. They have been quietly bootstrapping and growing in line with revenue for over 2 decades. They maintain independence from external stakeholders, enabling them to chart their course, maintain a long-term perspective and build an enduring, sustainable business that currently employs ~600 team members globally.

Data is central to nearly every part of their platform. High-volume event and transactional data powers customer reporting, billing, marketplace analytics, experimentation, machine-learning workflows, and product experiences. The organization is evolving from centralized, batch-oriented reporting toward a platform-driven architecture combining batch, streaming, and asynchronous processing, with APIs becoming a primary interface to data.

Our client has been stubbornly racking and stacking infrastructure around the world for the duration of their existence, a habit that allows for them to price themselves at the cost of electricity while their competitors are mired in rising cloud infrastructure costs. This long-term, somewhat contrarian, thinking puts them in a position to offer stability and career longevity evidenced by robust benefits that include RRSP matching.

This is an IC role for someone who can move comfortably between deep technical problem solving and broader architectural leadership. You’ll help define how database systems are designed, operated, automated, observed, and evolved across the organization. You’ll work closely with platform engineering, infrastructure, software engineering, operations, and data teams to improve reliability and scalability while championing simplicity in a sophisticated technology landscape.

Key Responsibilities

  • Database architecture & reliability: Design and evolve globally distributed database systems for performance, availability, scalability, and continuity in both bare-metal and cloud environments
  • Technical leadership: Lead complex initiatives, establish technical direction, and influence engineering decisions across teams without relying on direct reporting authority
  • Platform simplification: Identify opportunities to reduce operational complexity and consolidate the database and infrastructure technology footprint
  • Automation: Design and implement automation for provisioning, configuration, deployment, upgrades, testing, maintenance, and database change management use cases
  • Operational excellence: Establish durable practices for operating critical data infrastructure, including monitoring, incident response, capacity planning, maintenance, security, IAM, auditing, and traceability
  • Performance engineering: Diagnose difficult performance and reliability problems across globally distributed, latency-sensitive infrastructure and drive issues through root cause to durable resolution
  • Incident learning: Lead retrospectives and root-cause analysis for significant data-system incidents and convert findings into improvements to tooling, architecture, and operational practices
  • Technical roadmap: Evaluate emerging technologies, develop roadmaps for the data platform, and communicate priorities and trade-offs across engineering and business stakeholders
  • Mentorship: Raise the technical bar for engineers working with data infrastructure through architecture reviews, technical guidance, documentation, and hands-on collaboration

Tech Stack

  • Databases & data systems: MySQL, MariaDB, Galera, PostgreSQL, Aerospike, Redis, Kafka, StarRocks, Vertica, Trino, Iceberg
  • Infrastructure & orchestration: Kubernetes, bare metal, Terraform, Ansible
  • CI/CD & deployment: Helm, GitLab workflows, ArgoCD
  • Observability: Prometheus, Grafana, Loki
  • Programming & automation: Python, Go, Bash/Shell, Kotlin, Java
  • Adjacent data technologies: Kafka, Spark, Flink, Hadoop, Presto/Trino

Your Know-How

  • You have 7+ years of experience spanning database reliability engineering, database administration, site reliability engineering focused on data systems, platform engineering, or a comparable infrastructure role
  • You have operated large-scale production database systems where reliability, latency, availability, and performance genuinely matter
  • You bring deep expertise with relational and distributed data systems and meaningful production experience with several technologies (such as MySQL, MariaDB, Galera, PostgreSQL, Kafka, Aerospike, Redis, StarRocks, or Vertica)
  • You have experience deploying and operating stateful systems across Kubernetes and bare-metal infrastructure
  • You understand relational data architecture and can make informed decisions around schema design, replication, availability, performance, scale, and continuity
  • You have strong automation instincts and experience with infrastructure tooling such as Terraform, Ansible, Helm, GitLab CI/CD, ArgoCD, or comparable technologies
  • You can automate infrastructure and operational workflows using Python, Go, Bash/Shell, Java, Kotlin, or another appropriate programming language
  • You have experience building or operating effective observability systems (using technologies such as Prometheus, Grafana, Loki, or equivalent tooling)
  • You understand database change management and have worked with tooling such as Liquibase, Flyway, Alembic, or similar systems
  • You can troubleshoot complex failures across distributed infrastructure, reason from symptoms through multiple layers of a system, and identify root causes rather than treating symptoms
  • You’re comfortable operating at “staff level” (aka setting direction, navigating ambiguity, mentoring other engineers, balancing immediate operational needs against long-term architecture, and influencing teams outside your immediate domain)
  • You can communicate complex technical decisions clearly to both deeply technical colleagues and stakeholders without the same infrastructure background

It’s a bonus if

  • You have experience operating databases across globally distributed, ultra-low-latency infrastructure
  • You have meaningful experience with streaming and large-scale data technologies such as Kafka, Spark, Flink, Hadoop, Trino/Presto, or Iceberg
  • You’ve operated infrastructure spanning both public cloud and privately operated data centres
  • You have experience simplifying or consolidating a large database technology footprint without compromising reliability or developer productivity
  • You have helped establish security, IAM, auditing, and traceability practices for critical data infrastructure
  • You’ve worked in adtech, financial infrastructure, high-frequency systems, large-scale marketplaces, or another environment where extremely high transaction volume and low latency are fundamental engineering constraints

Interested in learning more?

Please send your resume or LinkedIn profile URL to [email protected] with “Staff Database Reliability Engineer” as the subject line. One of our talent partners will be in contact shortly!
Compensation
The base pay range for this role is CA$180,000 – CA$250,000 per year.

Skills Required

  • 7+ years of experience in database reliability engineering, database administration, site reliability engineering focused on data systems, platform engineering, or comparable infrastructure roles
  • Experience operating large-scale production database systems where reliability, latency, availability, and performance are critical
  • Deep expertise with relational and distributed data systems and production experience with several listed database technologies
  • Experience deploying and operating stateful systems across Kubernetes and bare-metal infrastructure
  • Understanding of relational data architecture, schema design, replication, availability, performance, scale, and continuity
  • Experience with infrastructure automation tools such as Terraform, Ansible, Helm, GitLab CI/CD, or ArgoCD
  • Ability to automate infrastructure and operational workflows using Python, Go, Bash/Shell, Java, Kotlin, or another appropriate language
  • Experience building or operating observability systems using Prometheus, Grafana, Loki, or equivalent tools
  • Understanding of database change management and experience with Liquibase, Flyway, Alembic, or similar tools
  • Ability to troubleshoot complex failures across distributed infrastructure and identify root causes
  • Staff-level experience setting direction, navigating ambiguity, mentoring engineers, balancing operational needs with architecture, and influencing other teams
  • Ability to communicate complex technical decisions clearly to technical colleagues and non-infrastructure stakeholders
  • Experience operating databases across globally distributed, ultra-low-latency infrastructure
  • Experience with streaming and large-scale data technologies such as Kafka, Spark, Flink, Hadoop, Trino/Presto, or Iceberg
  • Experience operating infrastructure across public cloud and privately operated data centers
  • Experience simplifying or consolidating a large database technology footprint
  • Experience establishing security, IAM, auditing, and traceability practices for critical data infrastructure
  • Experience in adtech, financial infrastructure, high-frequency systems, large-scale marketplaces, or similarly high-volume low-latency environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
10 Employees
Year Founded: 2023

What We Do

Lutra provides engineering and software services, including compliance management software and AI-driven automation solutions, to improve operations.

Similar Jobs

Samsara Logo Samsara

Senior AI Security Research Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
CA
4000 Employees
158K-239K Annually

MongoDB Logo MongoDB

Senior Software Engineer

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
2 Locations
5550 Employees
158K-220K Annually

Samsara Logo Samsara

Project Manager

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
CA
4000 Employees
80K-121K Annually

Affirm Logo Affirm

Software Engineer

Big Data • Fintech • Mobile • Payments • Financial Services
Easy Apply
Remote
Canada
2200 Employees
133K-183K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account