Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!
Site Reliability Engineer (SRE)
The Opportunity
The Site Reliability Engineer (SRE) plays a critical role in ensuring the reliability, scalability, performance, and operational excellence of Qualys platforms and services. This role operates at the intersection of software engineering and operations, applying automation, observability, troubleshooting, and reliability engineering practices to maintain highly available production systems and improve customer experience.
The SRE will work closely with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to monitor production systems, troubleshoot issues, automate operational processes, and continuously improve system reliability.
Where This Role Sits
The SRE role partners closely with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to ensure production systems remain resilient, scalable, observable, and operationally efficient.
The role provides hands-on support for cloud-native and distributed systems, including Kubernetes, Apache Spark, Kafka, and cloud-based data platforms, while contributing to incident response, automation, monitoring, and reliability improvements.
Key Responsibilities
Reliability and Operations
- Monitor production applications, infrastructure, and distributed systems to identify issues and maintain service health.
- Participate in incident response, troubleshooting, escalation, and root-cause analysis.
- Support highly available and scalable production services and workloads.
- Monitor system performance, resource utilization, availability, latency, and overall service health.
- Participate in on-call rotations and follow established incident management processes.
- Support capacity planning and identify potential reliability and performance risks.
- Assist engineering teams in resolving production issues and implementing corrective actions.
- Contribute to post-incident reviews and continuous reliability improvements.
Data Platform and Distributed Systems
- Support the operation and monitoring of Apache Spark workloads and Kafka pipelines.
- Monitor distributed data-processing workloads and investigate performance or availability issues.
- Assist with troubleshooting Spark jobs, Kafka pipelines, application failures, resource constraints, and related production issues.
- Develop an understanding of distributed systems concepts and their operational requirements.
- Support data platforms and technologies such as Delta Lake and S3 where applicable.
- Apply basic SQL skills for troubleshooting, validation, and operational analysis.
Kubernetes and Cloud Operations
- Assist with Kubernetes deployments, resource monitoring, application troubleshooting, and operational support.
- Monitor containerized applications and investigate resource, availability, and performance issues.
- Support cloud environments such as AWS, OCI, or Google Cloud Platform.
- Work with infrastructure and engineering teams on configuration, deployments, and operational workflows.
- Develop familiarity with Infrastructure as Code (IaC) and cloud-native operational practices.
Monitoring and Observability
- Monitor applications, infrastructure, and distributed workloads using tools such as Prometheus and Grafana.
- Create and maintain dashboards, alerts, and monitoring configurations.
- Analyze logs, metrics, and system behavior to identify and troubleshoot production issues.
- Track reliability indicators such as availability, latency, performance, resource utilization, and service health.
- Contribute to improvements in observability and proactive issue detection.
Automation and Platform Engineering
- Develop scripts and automation to reduce repetitive manual operational tasks.
- Create tooling for monitoring, log analysis, troubleshooting, deployments, and operational workflows.
- Use Python, Bash, Go, or similar programming/scripting languages to automate operational processes.
- Support CI/CD and automated deployment workflows.
- Contribute to Infrastructure as Code and self-service capabilities where applicable.
- Continuously improve operational processes through automation and standardization.
Collaboration and Continuous Improvement
- Collaborate with Engineering, DevOps, Infrastructure, Security, and Data Platform teams to improve system reliability.
- Assist development teams in designing and operating production-ready services.
- Document system architectures, operational procedures, troubleshooting steps, and best practices.
- Create and maintain runbooks, SOPs, and troubleshooting guides.
- Proactively identify reliability risks and opportunities for operational improvement.
- Contribute to SRE best practices, knowledge sharing, and continuous improvement initiatives.
What Good Looks Like
- Proactively identifies reliability and performance risks before they significantly impact customers.
- Responds effectively to production incidents and contributes to meaningful root-cause analysis.
- Uses automation to reduce repetitive manual operational work.
- Demonstrates strong troubleshooting skills across Linux, Kubernetes, cloud, and distributed systems.
- Effectively monitors and analyzes system metrics, logs, and application behavior.
- Collaborates effectively with engineering and infrastructure teams to resolve production issues.
- Continuously improves system reliability, observability, scalability, and operational efficiency.
- Builds a strong understanding of production systems and applies SRE principles to day-to-day operations.
Distinguishing Expectations
The SRE is measured by improvements in reliability, automation, operational efficiency, observability, and service availability.
For an early-career SRE, success is demonstrated through:
- Effective monitoring and troubleshooting of production systems.
- Rapid learning of complex production environments.
- Reduction of manual operational effort through automation.
- Consistent participation in incident response and root-cause analysis.
- Improvements to monitoring, documentation, and operational processes.
- Growing ownership of production services and reliability-focused initiatives.
Required Qualifications
- 0–3 years of experience in SRE, DevOps, Cloud Engineering, Systems Engineering, or a related technical role.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Basic understanding of Linux/Unix systems, processes, memory, networking, and system troubleshooting.
- Fundamental understanding of distributed systems and cloud-native technologies.
- Experience or familiarity with Kubernetes and containerized applications.
- Familiarity with Apache Spark, Kafka, or similar distributed data-processing platforms.
- Strong programming and scripting fundamentals in Python, Bash, Go, Java, Scala, or similar languages.
- Ability to write scripts for automation, monitoring, log analysis, and operational tasks.
- Basic SQL knowledge and strong problem-solving skills.
- Familiarity with monitoring and observability tools such as Prometheus and Grafana.
- Basic understanding of cloud platforms such as AWS, OCI, or Google Cloud Platform.
- Understanding of basic networking, security, and distributed systems concepts.
- Willingness to learn production systems, debugging techniques, and SRE practices.
Preferred Qualifications
- Familiarity with Delta Lake, S3, or distributed data platforms.
- Understanding of CI/CD and automated deployment processes.
- Familiarity with Terraform, CloudFormation, or other Infrastructure as Code tools.
- Experience with scripting-based automation or internal tooling projects.
- Familiarity with tools such as Jenkins, GitHub, Bitbucket, ELK, Splunk, or AppDynamics.
- Experience with large-scale or distributed systems.
- Understanding of reliability engineering principles and automation best practices.
- Exposure to incident management, on-call processes, and postmortem practices.
- Cloud certifications are a plus.
- Strong analytical, communication, collaboration, and troubleshooting skills.
Core Competencies
SRE | Linux | Python / Bash / Go | Coding & Scripting | Spark | Kafka | Kubernetes | Cloud Platforms | SQL | Monitoring & Observability | Automation | Distributed Systems | Troubleshooting | Incident Response | CI/CD | Reliability Engineering
Work Environment
- Full-time role.
- On-call participation may be required.
- Hybrid or remote flexibility depending on company policy.
- Cross-functional collaboration with Engineering, DevOps, Infrastructure, Security, and Data Platform teams.
- Opportunity to work with cloud-native, distributed, and large-scale production systems.
ABOUT QUALYS
Qualys, Inc. is a pioneer and leading provider of cloud-based IT, security, and compliance solutions, serving more than 10,000 customers across over 130 countries. Qualys delivers innovative security and compliance solutions that help organizations simplify security operations and reduce risk.
Skills Required
- 0–3 years of experience in SRE, DevOps, Cloud Engineering, Systems Engineering, or a related technical role
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience
- Basic understanding of Linux/Unix systems, processes, memory, networking, and system troubleshooting
- Fundamental understanding of distributed systems and cloud-native technologies
- Experience or familiarity with Kubernetes and containerized applications
- Familiarity with Apache Spark, Kafka, or similar distributed data-processing platforms
- Programming and scripting fundamentals in Python, Bash, Go, Java, Scala, or similar languages
- Ability to write scripts for automation, monitoring, log analysis, and operational tasks
- Basic SQL knowledge and strong problem-solving skills
- Familiarity with Prometheus, Grafana, or similar monitoring and observability tools
- Basic understanding of AWS, OCI, or Google Cloud Platform
- Understanding of networking, security, and distributed systems concepts
- Willingness to learn production systems, debugging techniques, and SRE practices
- Familiarity with Delta Lake, S3, or distributed data platforms
- Understanding of CI/CD and automated deployment processes
- Familiarity with Terraform, CloudFormation, or other Infrastructure as Code tools
- Experience with scripting-based automation or internal tooling projects
- Familiarity with Jenkins, GitHub, Bitbucket, ELK, Splunk, or AppDynamics
- Experience with large-scale or distributed systems
- Understanding of reliability engineering, incident management, on-call, and postmortem practices
- Cloud certifications
- Strong analytical, communication, collaboration, and troubleshooting skills
Qualys Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Qualys and has not been reviewed or approved by Qualys.
-
Affordable Benefits — Benefits costs are widely viewed as low for employees and dependents, with healthcare often described as almost fully paid for. Feedback suggests this affordability helps offset perceptions of lower base pay in some roles.
-
Healthcare Strength — Healthcare offerings are broad, including multiple medical plan options, dental and vision coverage, mental health support, and disability insurance. Benefits are described as “pretty amazing” or “great,” reinforcing perceived quality and coverage depth.
-
Equity Value & Accessibility — Equity participation is accessible through company stock plans and an employee stock purchase plan. Compensation packages commonly include equity alongside salary and bonus, which some consider a meaningful part of total rewards.
Qualys Insights
What We Do
Qualys, Inc. (NASDAQ: QLYS) is a pioneer and leading provider of disruptive cloud-based security, compliance and IT solutions with more than 10,000 subscription customers worldwide, including a majority of the Forbes Global 100 and Fortune 100. Qualys helps organizations streamline and automate their security and compliance solutions onto a single platform for greater agility, better business outcomes, and substantial cost savings. The Qualys Cloud Platform leverages a single agent to continuously deliver critical security intelligence while enabling enterprises to automate the full spectrum of vulnerability detection, compliance, and protection for IT systems, workloads and web applications across on premises, endpoints, servers, public and private clouds, containers, and mobile devices. Founded in 1999 as one of the first SaaS security companies, Qualys has strategic partnerships and seamlessly integrates its vulnerability management capabilities into security offerings from cloud service providers, including Amazon Web Services, the Google Cloud Platform and Microsoft Azure, along with a number of leading managed service providers and global consulting organizations. For more information, please visit http://www.qualys.com








