Sr. Site Reliability Engineer

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in Canada
Remote
131K-164K Annually
Senior level
Information Technology • Cybersecurity
The Role
Designs and maintains scalable cloud and on-premise infrastructure, Infrastructure as Code, Kubernetes environments, CI/CD pipelines, data streaming, caching, observability, and incident response systems. The role owns AWS reliability and cost efficiency, supports progressive deployments, troubleshoots complex production issues, partners with software teams, and drives automation and continuous improvement across platform operations.
Summary Generated by Built In

Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts who applied their learnings to bring national security-grade technology solutions to commercial customers around the world, Blackpoint Cyber is in hyper-growth mode,  fueled by a recent $190m series C round. 

SUMMARY

We're hiring a Senior Site Reliability Engineer to design, implement, and maintain our cloud and on-premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance. You'll work across cloud platform administration, container orchestration, data streaming, observability, and incident response — partnering with engineering teams to keep our systems reliable, secure, and efficient, and helping foster a culture of continuous improvement.

RESPONSIBILITIES

  • Design, develop, and maintain highly scalable infrastructure using Infrastructure as Code (Terraform and Terragrunt) for automated cloud resource provisioning and orchestration.

  • Own and optimize our AWS cloud environment, ensuring cost efficiency, security best practices, and high-availability standards.

  • Manage and optimize Kubernetes cluster environments (Helm, ArgoCD, Istio, Kustomize) to support continuous delivery and infrastructure-as-code practices.

  • Administer and scale data streaming infrastructure (Confluent Cloud, Apache Kafka) to support enterprise-level data processing.

  • Deploy, configure, and maintain Redis for caching and real-time data processing.

  • Implement and maintain monitoring, alerting, and incident response frameworks (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty) to ensure system reliability and performance.

  • Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog.

  • Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure.

  • Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance and maximize uptime.

  • Drive continuous improvement of automation tooling, operational processes, and engineering methodologies to enhance scalability, reliability, and maintainability.

  • Stay current on emerging SRE trends and tools, and help the team adopt relevant industry advancements and best practices.

REQUIREMENTS

  • 5+ years of experience in a Senior Site Reliability Engineer role or equivalent, with substantial emphasis on cloud infrastructure management and automation.

  • Expertise in Infrastructure as Code (Terraform, Terragrunt) for enterprise-scale deployments.

  • Comprehensive knowledge of AWS, including designing, implementing, and maintaining secure, scalable, resilient cloud architectures.

  • Extensive hands-on experience with distributed data streaming (Confluent Cloud, Apache Kafka).

  • Proven experience with Redis for caching and Amazon RDS for relational database management.

  • Experience with enterprise search and analytics platforms (OpenSearch, Elasticsearch, ChaosSearch).

  • Proficiency designing and implementing monitoring/alerting infrastructure (Prometheus, Grafana, Alert Manager, OpsGenie/PagerDuty).

  • Practical experience with feature flag systems (LaunchDarkly/PostHog) for controlled release management.

  • Extensive experience administering production-grade Kubernetes (Helm, ArgoCD, Istio); working knowledge of Kustomize.

  • Strong problem-solving skills, with the ability to troubleshoot complex systems in production.

  • Strong communication and collaboration skills, with experience working in Agile environments.

NICE TO HAVE

  • Multi-cloud experience (Google Cloud Platform, Microsoft Azure).

  • Understanding of security frameworks and compliance standards for cloud-native/containerized environments.

  • Serverless computing and CI/CD pipeline experience (Jenkins, GitHub Actions).

  • Software development proficiency in Node.js, Python, and/or Go.

Blackpoint Cyber welcomes and encourages applications from qualified individuals of all races, colors, religions, sex, sexual orientation, gender identity or expression, national origin, age, marital status, or any other legally protected status. We are committed to equality of opportunity in all aspects of employment.

For eligible employees in the US, Blackpoint offers competitive Health, Vision, Dental, and Life Insurance plans, a robust 401k plan, Discretionary Time Off, and other minor perks. International employees receive competitive benefits in accordance with local market standards and applicable country requirements.

Blackpoint believes all employees should share in the company’s success – equity participation is available to employees globally, with program details varying by location and employment structure.

Skills Required

  • 5+ years of experience as a Senior Site Reliability Engineer or equivalent, emphasizing cloud infrastructure management and automation
  • Expertise with Terraform and Terragrunt for enterprise-scale Infrastructure as Code deployments
  • Comprehensive experience designing, implementing, and maintaining secure, scalable, resilient AWS architectures
  • Extensive hands-on experience with Confluent Cloud and Apache Kafka
  • Experience with Redis and Amazon RDS
  • Experience with OpenSearch, Elasticsearch, or ChaosSearch
  • Experience designing and implementing monitoring and alerting infrastructure using Prometheus, Grafana, Alert Manager, OpsGenie, or PagerDuty
  • Experience with LaunchDarkly or PostHog feature flag systems
  • Extensive production Kubernetes administration experience with Helm, ArgoCD, and Istio; working knowledge of Kustomize
  • Strong production troubleshooting and complex problem-solving skills
  • Strong communication and collaboration skills, including experience in Agile environments
  • Multi-cloud experience with Google Cloud Platform or Microsoft Azure
  • Understanding of security frameworks and compliance standards for cloud-native and containerized environments
  • Serverless computing and CI/CD pipeline experience with Jenkins or GitHub Actions
  • Software development proficiency in Node.js, Python, or Go
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Denver, Colorado
187 Employees
Year Founded: 2014

What We Do

Blackpoint Cyber is a technology-driven cybersecurity company headquartered in Denver, Colorado. Founded in 2014 by former U.S. Department of Defense and intelligence security experts, Blackpoint leverages decades of real-world experience and deep knowledge of malicious tradecraft to provide proactive, nation-state-grade cybersecurity to organizations worldwide. Our mission is clear: to deliver 24/7, human-powered Managed Detection, Response, and Remediation (MDR) services, empowering IT professionals with the industry’s fastest threat elimination and risk mitigation capabilities. Blackpoint’s proprietary technology and active Security Operations Center (SOC) work together to stop cyber threats in real-time, ensuring organizations of all sizes remain protected in a constantly evolving threat landscape. At Blackpoint, we are a passionate team of cybersecurity professionals dedicated to helping Managed Service Providers (MSPs) become the heroes modern businesses rely on. By arming MSPs with cutting-edge technology, relentless 24/7 support, and a trusted partnership, we help them safeguard their clients and combat cyber threats with confidence and precision. At Blackpoint Cyber, we believe sophisticated cybersecurity should be accessible to all. That’s why we remain deeply committed to the growth and success of the Managed IT and Security community, offering cutting-edge solutions that empower IT professionals to combat cyber threats with confidence. Strike first and secure fast with Blackpoint Cyber.

Similar Jobs

Remote
British Columbia, BC, CAN
729 Employees
Remote
Canada
2969 Employees
In-Office or Remote
3 Locations
456 Employees
120K-170K Annually
In-Office or Remote
Montréal, QC, CAN
1000 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account