As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.
About the Role:
At CrowdStrike, Site Reliability Engineering (SRE) is at the forefront of ensuring the reliability and scalability of our cloud-native security platform. In this role, you'll manage a team of talented engineers, providing technical leadership on key projects and empowering them to excel in their roles.
As an SRE Manager, you will lead a team of SRE engineers ensuring the reliability, scalability, and performance of CrowdStrike's cloud-native security platform. You'll provide technical leadership and mentorship, owning both reliability engineering and software delivery pipelines - driving engineering velocity while maintaining zero tolerance for downtime in security-critical infrastructure.
What you will Do
Define and enforce SLOs, SLIs, and error budgets across distributed systems processing millions of events per second
Drive system reliability by blending software engineering principles with AI-driven automation, moving from reactive firefighting to proactive, automated operations
Lead major incident response and facilitate blameless postmortems, driving systemic reliability improvements
Own capacity planning, traffic management, and load shedding strategies for high-throughput distributed systems
Own the end-to-end software delivery pipeline strategy — designing, building, and maintaining scalable, reliable pipelines using Jenkins, GitLab CI, and Bitbucket Pipelines
Build and maintain observability frameworks including metrics, distributed tracing, and log aggregation across the full stack
Champion chaos engineering and resilience validation practices for security-critical systems
Lead and grow a high-performing SRE team, mentoring engineers and fostering a culture of continuous learning and operational excellence
Partner with cross-functional engineering teams to embed reliability practices early in the software development lifecycle
What You'll Need
Experience & Leadership
Proven track record of building, growing, and retaining high-performing SRE/DevOps engineering teams in a fast-paced, high-growth environment
10+ years of software engineering experience with significant focus on reliability engineering, platform infrastructure, and production operations at scale
3+ years of hands-on management experience overseeing SRE/DevOps engineering teams, including incident command and reliability ownership
Bachelor's degree in Computer Science or related field, or equivalent work experience
Reliability Engineering
Deep understanding of SRE principles including SLOs, SLAs, SLIs, and error budgeting strategies applied to large-scale distributed systems
Proven experience owning reliability for high-throughput distributed systems processing millions of events per second, including capacity planning, traffic management, and load shedding strategies
Strong incident management facilitating blameless postmortems, and driving system reliability improvements
Demonstrated ability to build, operationalize, and maintain highly scalable, security-critical microservices-based distributed systems with zero tolerance for data loss or downtime.
Advanced observability experience including Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), and large-scale log aggregation (ELK/Splunk) with a focus on building custom SLO dashboards and reliability scorecards.
Experience owning disaster recovery strategies including backup automation, failover testing, and business continuity planning for stateful distributed systems
Platform and Delivery Engineering
Proficiency in Python and/or Golang for automation, tooling, and platform services
Hands-on experience designing and managing scalable software delivery pipelines using Jenkins, GitLab CI, Bitbucket Pipelines, or equivalent
Strong proficiency in Infrastructure as Code (IaC) - Terraform, Ansible, Pulumi, or equivalent
Familiarity with GitOps workflows using ArgoCD or Flux for managing infrastructure deployments at scale
Cloud and Big Data Exposure
Proficiency in at least one cloud environment (AWS, Azure, GCP) with emphasis on multi-region architecture, cloud-native reliability patterns, and security-first cloud design
Strong experience with Kubernetes at scale - managing large cluster fleets, workload orchestration, and container lifecycle management
Familiarity with distributed data systems including relational databases (PostgreSQL), NoSQL (Cassandra), OLAP (Pinot), Indexing(OpenSearch) and real-time streaming platforms (Kafka, Flink)
Exposure to Big Data and analytics technologies like Spark,Storm.
#LI-AP1
Benefits of Working at CrowdStrike:
Market leader in compensation and equity awards
Comprehensive physical and mental wellness programs
Competitive vacation and holidays for recharge
Paid parental and adoption leaves
Professional development opportunities for all employees regardless of level or role
Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
Vibrant office culture with world class amenities
Great Place to Work Certified™ across the globe
CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.
CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.
If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at [email protected] for further assistance.
Find out more about your rights as an applicant.
CrowdStrike participates in the E-Verify program.
Notice of E-Verify Participation
Right to Work
CrowdStrike, Inc. is committed to fair and equitable compensation practices. Placement within the pay range is dependent on a variety of factors including, but not limited to, relevant work experience, skills, certifications, job level, supervisory status, and location. The base salary range for this position for all U.S. candidates is $140,000 - $215,000 per year, with eligibility for bonuses, equity grants and a comprehensive benefits package that includes health insurance, 401k and paid time off.For detailed information about the U.S. benefits package, please click here.
Skills Required
- Proven track record of building, growing, and retaining high-performing SRE/Platform engineering teams
- 10+ years of software engineering experience with focus on reliability engineering, platform infrastructure, and production operations at scale
- 3+ years of hands-on management experience overseeing SRE/Platform engineering teams, including incident command
- Deep understanding of SRE principles including SLOs, SLAs, SLIs, and error budgeting strategies
- Experience driving system reliability using software engineering and automation (AI-driven automation preferred)
- Proficiency in at least one cloud environment (AWS, Azure, or GCP) with multi-region architecture experience
- Proven experience owning reliability for high-throughput distributed systems, capacity planning, traffic management, load shedding
- Strong incident management background including leading major incident response and blameless postmortems
- Demonstrated ability to build, operationalize, and maintain highly scalable, security-critical systems with zero tolerance for data loss or downtime
- Bachelor's degree in Computer Science or related field, or equivalent work experience
- Ability to work 2+ days per week in Sunnyvale offices
- Experience operating security platforms, telemetry pipelines, or sensor fleet infrastructure at massive scale
- Proficiency in Golang
- Experience with Kubernetes at scale, including service mesh technologies (Istio/Linkerd) and container security practices
- Familiarity with high-throughput data streaming platforms such as Apache Kafka and Apache Flink
- Experience with hybrid cloud and multi-cloud failover strategies
- Advanced observability experience: Prometheus, Grafana, distributed tracing (Jaeger/OpenTelemetry), ELK/Splunk
- Experience building and running chaos engineering programs
- Contributions to open-source projects or active involvement in tech community
CrowdStrike Compensation & Benefits Highlights
-
Healthcare Strength — Company materials outline medical, dental, and vision coverage alongside HSAs/FSAs and an Employee Assistance Program, indicating a comprehensive health foundation.
-
Parental & Family Support — Official documents include paid family leave, family care resources, and adoption and infertility assistance, signaling meaningful support for caregivers.
-
Equity Value & Accessibility — Equity grants are paired with an Employee Stock Purchase Plan with a lookback discount, enhancing participation in company ownership.
CrowdStrike Insights
What We Do
CrowdStrike has redefined security with the world’s most advanced cloud-native platform that protects and enables the people, processes and technologies that drive modern enterprise. Tested and proven, the world's largest organizations trust CrowdStrike to stop breaches with unparalleled protection against the most sophisticated cyberattacks. The CrowdStrike culture has been built upon our Core Values since the day we began. We are Fanatical About the Customer, Relentlessly Focused on Innovation and believe that our Limitless Passion drives Unlimited Potential for every CrowdStriker. As a purpose-built remote-first company, we believe cultivating a connected culture for every employee, no matter where they are in the world, is a key ingredient in building a high-performing, diverse team. We don’t have a mission statement. We’re on a mission—to stop breaches. Ready to join a mission that matters?
Why Work With Us
We have a culture that celebrates achievement, encourages flexibility and innovation and thrives on teamwork. We all work towards a single mission: to stop breaches. This common goal drives a sense of community and connection among our people across the globe.
Gallery
CrowdStrike Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.



























