Sr Staff Site Reliability Engineer (SRE), Cloud

Sorry, this job was removed at 03:52 p.m. (CST) on Wednesday, Feb 19, 2025
Hiring Remotely in Canada
Remote
Software
Cribl is the AI Platform for Telemetry.
The Role

Cribl does differently. 

What does that mean? It means we are a serious company that doesn’t take itself too seriously; and we’re looking for people who love to get stuff done, and laugh a bit along the way. We’re growing rapidly - looking for collaborative, curious, and motivated team members who are passionate about putting customers first. As a remote-first company we believe in empowering our employees to do their best work, wherever they are. 

As the data engine for IT and Security many of the biggest names in the most demanding industries trust Cribl to solve their most pressing data needs. Ready to do the best work of your career? Join the herd and unlock your opportunity.

About Cribl

Cribl unlocks the value of observability data.
Our products deliver choice and control over the rising volumes of telemetry data, enabling every organization to realize the value of all their observability data without limitation. Backed by the industry’s leading venture capitalists, including CRV, Sequoia Capital, Greylock Partners, Redpoint Ventures, and IVP, our solutions are deployed across organizations of all sizes. Many of the biggest names in the most demanding industries trust Cribl to solve their most pressing observability needs.

At our core, we foster an inclusive, values-aligned culture where all belong. We believe in a remote-first operating model, empowering the flexibility to do your best work, wherever you are. We’re also growing rapidly, welcoming collaborative, curious, and motivated team members who are passionate about putting customers first. 

Join the herd and unlock your opportunity.


About the Team

Cribl Inc is seeking a Senior Staff Site Reliability Engineer to join our mission where you will unlock the value of all observability data, as we expand our team. Cribl provides users a new level of observability, intelligence and control over their real-time data. You will join a team of technical engineers who are committed to shipping only high-quality software and enjoying all the goat gifs the internet has to offer. This role is remote, and you will be part of the engineering organization where you will contribute in our efforts to envision, create, deploy, test, and ship Cribl products.

Not often do you get to be part of something that is fundamentally changing a technology. But here at Cribl we are building the next generation of software that puts our customers in full control of their observability data. If this is something that interests you, and you want to be truly at the center of the wheel helping make this work better every day. Then this opportunity might be something you have been waiting for to be a part of making a real impact.

We are looking for Cloud Site Reliability Engineers and Developers, at all levels at Cribl, who enjoy being in the thick of it. Fixing things at the operational side should always be the last resort, so our SRE engineers are involved from conception to design to development and all the way through production and beyond. You provide your creative input into all things Cloud, Scaling, Reliability, High Availability and much more.

If reliability is your passion, and you have always had strong opinions on how to make things better and have the desire to build consensus around ideas. Then let's talk!


As An Active Member Of Our Team, You Will...

  • Engage with teams and improve service delivery and reliability across their entire lifecycle
  • Measure and monitor all production systems with an eye towards availability, latency and overall system health
  • Design observability systems for different types of applications, using Cribl products and other OpenSource tools
  • Seek out the cause of errors and instability in our production cloud services and drive teams towards better operational excellence
  • Engage with product and platform teams to improve and evolve systems by lobbying for changes that improve reliability, resilience, and observability
  • Lead efforts enabling shift-left monitoring in the organization
  • Help Identify and drive down toil with creative innovation and automation
  • On-call responsibilities


If You Got It - We Want It

  • Extensive experience with enterprise-scale continuous delivery environments
  • Development with JavaScript/Node.js/TypeScript in a Linux/Mac environment
  • Experience with Configuration Management Tools like Terraform (preferred) or Puppet, Chef, Ansible
  • Knowledge of cloud platforms (prefer AWS and Azure, GCP is nice to have) and container + orchestration technologies
  • Extensive experience designing and implementing Observability platforms based on OpenSource tools like Grafana, Prometheus, OpenSearch
  • Experience mentoring engineers, and acting as Subject Matter Expert in areas of Monitoring and Observability.
  • Experience with native monitoring services in AWS, Azure and other popular Cloud Platforms
  • Background in Linux Systems Engineering
  • Experience with Incident response tools, for instance, PagerDuty, FireHydrant etc.
  • Experience with sustainable incident response in a blameless environment
  • Comfortable with a high level of autonomy and working with a distributed team


Preferred Qualifications

  • Knowledge of Cloud and application security
  • Strong knowledge of cloud design patterns for scale, data management, resiliency, etc
  • A love for high quality and a knack for testing
  • Opinions about dashboards, metrics, and SLO’s


Bring Your Whole Self
Diversity drives innovation, enables better decisions to support our customers, and inspires change for the better. We’re building a culture where differences are valued and welcomed. We work together to bring out the best in each other. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which the candidate is applying.

#LI-EL1

Bring Your Whole Self
Diversity drives innovation, enables better decisions to support our customers, and inspires change for the better. We’re building a culture where differences are valued and welcomed, and we work together to bring out the best in each other. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, or any other applicable legally protected characteristics in the location in which the candidate is applying.

Interested in joining the Cribl herd? Learn more about the smartest, funniest, most passionate goats you’ll ever meet at cribl.io/about-us

Cribl Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cribl and has not been reviewed or approved by Cribl.

  • Affordable Benefits Medical and dental premiums are fully covered for individuals in the U.S., with low costs for dependents, and the plans are described as low‑cost overall. This positions healthcare expenses favorably for many employees.
  • Leave & Time Off Breadth Unlimited PTO, paid holidays, and periodic company “refresh” or winter‑break days provide ample time away. Flexible schedules further support taking time when needed.
  • Wellbeing & Lifestyle Benefits A monthly stipend for home office, phone, and internet, plus strong remote‑work setup support, underpin the remote‑first model. Additional perks like recharge days and equipment support bolster day‑to‑day wellbeing.

Cribl Insights

Similar Jobs

Newton.co Logo Newton.co

Site Reliability Engineer

Blockchain • Financial Services • Cryptocurrency • Web3
In-Office or Remote
Toronto, ON, CAN
77 Employees
Remote
Canada
91 Employees

Entrust Logo Entrust

Senior Site Reliability Engineer

Information Technology • Security • Software
In-Office or Remote
12 Locations
2800 Employees
121K-189K Annually

TextNow Logo TextNow

Site Reliability Engineer

Digital Media • Social Media
In-Office or Remote
Open Hall, Subd. F, NL, CAN
239 Employees
113K-162K Annually
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
1,000 Employees
Year Founded: 2018

What We Do

Cribl, the AI Platform for Telemetry, empowers enterprises to manage and analyze telemetry for both humans and agents. Trusted by organizations worldwide, including half of the Fortune 100, Cribl bridges the gap between AI ambition and infrastructure reality. No lock-in. No data loss. No compromises. Cribl’s vendor-agnostic platform ensures data remains portable and interoperable. By cost-effectively handling increasing data volume and variety without delay, Cribl gives enterprises the choice, control, and flexibility to build what’s next.

Why Work With Us

We are building the company that will become the industry leader in IT and Security data. But, doing that doesn’t mean we’re always serious. We approach our work fearlessly, learn quickly, improve constantly, and celebrate our wins at every turn. And more importantly, we laugh a lot.

Gallery

Gallery

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account