As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. We work on large scale distributed systems, processing almost 3 trillion events per day and this traffic is growing daily. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We're proud to work for a mission-driven company leveraging AI to transform the way we work. CrowdStrikers drive their careers through flexibility and autonomy while also being expected to contribute to a culture of responsible AI adoption, experimentation, and innovation. We use an AI-first mindset as a force multiplier to proactively and continuously accelerate execution, build expertise, uncover insights, and solve complex problems. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you.
About the Role:
Our mission is to make all of our customers' security-relevant data continuously available for automated detection and response, threat hunting, and other Falcon platform use cases. To enable this, the systems behind NG-SIEM (next-generation security information and event management) are growing to accommodate >100 PB of event and action data ingested every day, up to 10 years of retention, and dozens of millions of queries per hour across large sections of the data stored, for tens of thousands of customers. As a Senior Engineer on the NG-SIEM Serverless Cell Platform team, you will be responsible for ensuring the health, performance, and reliability of cell infrastructure — operating the security industry's largest SIEM platform while building the automation and self-healing systems that enable it to scale seamlessly.
The NG-SIEM platform is built on a cell-based architecture where each cell is a self-contained LogScale cluster serving customer data. As we scale toward hundreds of petabytes per day, you will own the operational management of these cells through monitoring, incident response, capacity planning, and continuous improvement — while also developing automation systems that reduce toil and enable the fleet to self-heal. You will participate in follow-the-sun on-call rotations (8x7), perform version upgrades and scaling operations, optimize resource utilization, and write Go microservices that transform manual operational procedures into robust, automated workflows. You will join a distributed team of high-ownership technical leaders who share a strong passion for our mission: to stop breaches.
This role is based in Australia and is fully remote. Many team members are located in Sydney.
What You'll Do:
Monitor and maintain the health, performance, and reliability of LogScale cells. Respond to incidents and participate in follow-the-sun on-call rotations with team support and clear escalation paths.
Upgrade LogScale clusters and microservices using ring-based rollout strategies. Deploy patches, update configurations, and recover cells when failures occur.
Build and maintain monitoring with defined Service Level Indicators. Implement alerting, synthetic test suites, and operational dashboards for rapid troubleshooting and root cause analysis.
Write Go microservices and automation scripts that transform manual operational procedures into automated workflows, reducing toil and enabling self-healing capabilities.
Work closely with product engineering and customer support teams to troubleshoot production issues and coordinate changes across the platform.
Learn and grow through mentorship opportunities, knowledge-sharing sessions, and exposure to large-scale distributed systems engineering challenges.
What You'll Need:
A passion for reliability engineering and operational excellence, with strong understanding of how large-scale distributed systems behave under pressure;
7+ years of experience in site reliability engineering, platform engineering, or infrastructure operations, with significant time operating and improving distributed systems at scale;
Experience with capacity planning, resource optimization, and scaling operations for large-scale infrastructure;
Proficiency in at least one programming language (Go, Python, Java, or similar) for writing automation scripts and microservices;
Experience with cloud infrastructure (AWS, OCI, or GCP) including compute, storage, networking, and infrastructure-as-code (Terraform, Pulumi, or similar);
A can-do attitude — you thrive collaborating in a team and are not afraid of taking on responsibilities;
Strong written and verbal communication skills for writing runbooks, incident reports, and collaborating with teams across time zones;
Ability to make pragmatic tradeoffs between operational stability and feature delivery needs.
Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.
If you meet most but not all of these qualifications, we still encourage you to apply.
Bonus Points:
Hands-on experience building monitoring systems, defining SLIs/SLOs, creating dashboards, and implementing alerting;
Strong troubleshooting and root cause analysis skills across distributed system components;
Experience operating database infrastructure, log management platforms, search systems, or SIEM platforms at scale;
Track record of building automation that enables self-healing infrastructure and reduces operational toil;
Experience with streaming platforms (Kafka or similar) and understanding of distributed data pipelines, backpressure, and partition management;
Familiarity with containerization and orchestration (Kubernetes, Docker, or similar);
Background in cybersecurity, security operations, or familiarity with security data workflows.
#LI-MZ1
Benefits of Working at CrowdStrike:
- Market leader in compensation and equity awards
- Comprehensive physical and mental wellness programs
- Competitive vacation and holidays for recharge
- Paid parental and adoption leaves
- Professional development opportunities for all employees regardless of level or role
- Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
- Vibrant office culture with world class amenities
- Great Place to Work Certified™ across the globe
CrowdStrike is proud to be an equal opportunity employer. We are committed to fostering a culture of belonging where everyone is valued for who they are and empowered to succeed. We support veterans and individuals with disabilities through our affirmative action program.
CrowdStrike is committed to providing equal employment opportunity for all employees and applicants for employment. The Company does not discriminate in employment opportunities or practices on the basis of race, color, creed, ethnicity, religion, sex (including pregnancy or pregnancy-related medical conditions), sexual orientation, gender identity, marital or family status, veteran status, age, national origin, ancestry, physical disability (including HIV and AIDS), mental disability, medical condition, genetic information, membership or activity in a local human rights commission, status with regard to public assistance, or any other characteristic protected by law. We base all employment decisions--including recruitment, selection, training, compensation, benefits, discipline, promotions, transfers, lay-offs, return from lay-off, terminations and social/recreational programs--on valid job requirements.
If you need assistance accessing or reviewing the information on this website or need help submitting an application for employment or requesting an accommodation, please contact us at [email protected] for further assistance.
Skills Required
- 7+ years in site reliability, platform engineering, or infrastructure operations working with distributed systems at scale
- Experience with capacity planning, resource optimization, and scaling operations for large-scale infrastructure
- Proficiency in at least one programming language for automation and microservices (Go, Python, Java, or similar)
- Experience with cloud infrastructure (AWS, OCI, or GCP) including compute, storage, networking
- Experience with infrastructure-as-code tools (Terraform, Pulumi, or similar)
- Strong written and verbal communication skills for runbooks, incident reports, and cross-timezone collaboration
- Pragmatic decision-making balancing operational stability and feature delivery
- Proven experience utilizing AI technologies to enhance decision-making and automate workflows
- Hands-on experience building monitoring systems, defining SLIs/SLOs, dashboards, and alerting
- Strong troubleshooting and root cause analysis across distributed system components
- Experience operating database, log management, search systems, or SIEM platforms at scale
- Track record building automation for self-healing infrastructure and reducing operational toil
- Experience with streaming platforms (Kafka or similar) and distributed data pipelines
- Familiarity with containerization and orchestration (Kubernetes, Docker, or similar)
- Background in cybersecurity or familiarity with security data workflows
Similar Jobs
What We Do
CrowdStrike has redefined security with the world’s most advanced cloud-native platform that protects and enables the people, processes and technologies that drive modern enterprise. Tested and proven, the world's largest organizations trust CrowdStrike to stop breaches with unparalleled protection against the most sophisticated cyberattacks. The CrowdStrike culture has been built upon our Core Values since the day we began. We are Fanatical About the Customer, Relentlessly Focused on Innovation and believe that our Limitless Passion drives Unlimited Potential for every CrowdStriker. As a purpose-built remote-first company, we believe cultivating a connected culture for every employee, no matter where they are in the world, is a key ingredient in building a high-performing, diverse team. We don’t have a mission statement. We’re on a mission—to stop breaches. Ready to join a mission that matters?
Why Work With Us
We have a culture that celebrates achievement, encourages flexibility and innovation and thrives on teamwork. We all work towards a single mission: to stop breaches. This common goal drives a sense of community and connection among our people across the globe.
Gallery
CrowdStrike Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.



























