ON.energy is setting the standard for large load interconnection. ON’s patented AI UPS™ is a medium-voltage UPS that serves as a firewall, protecting the data center from the grid, and the grid from the data center. With multiple gigawatts currently under construction, we are enabling grid-safe data centers.
ROLE SUMMARY
We're looking for A Senior Infrastructure QA Engineer to validate the reliability, performance, and deployment integrity of our server and application infrastructure, including systems running on Red Hat Enterprise Linux (RHEL) and containerized workloads. This role sits at the intersection of QA and systems/infrastructure engineering — testing not just application behavior, but the platforms and deployment pipelines that behavior runs on, including performance under load, correct behavior across upgrades, and validation of containerized and orchestrated deployments.
KEY RESPONSIBILITIES
- Design and execute test plans for infrastructure changes, including OS-level updates, container deployments, and platform upgrades.
- Validate server sizing and performance under load, including load, stress, and soak testing of production-bound systems.
- Test and verify containerized application deployments for correctness, resource behavior, and rollback safety.
- Build and maintain test automation and monitoring scripts to validate infrastructure health and deployment outcomes.
- Collaborate with systems, DevOps, and application engineering teams to define acceptance criteria for infrastructure changes.
- Document test results, defects, and qualification outcomes clearly for technical and non-technical stakeholders.
KEY REQUIREMENTS
- 5+ Years of hands-on Linux administration experience, specifically with Red Hat Enterprise Linux (RHEL) — system administration, package management, service management (systemd), and shell scripting.
- Experience testing or validating containerized deployments — building, running, and troubleshooting Docker containers, and working knowledge of container orchestration fundamentals (Kubernetes or OpenShift).
- Solid QA/testing background — designing test plans and test cases, executing structured test cycles, and tracking/triaging defects to resolution.
- Experience with CI/CD pipelines and validating automated build/deployment processes (e.g. Jenkins, GitLab CI, Azure DevOps, or similar).
- Working knowledge of networking fundamentals (TCP/IP, DNS, firewalls, VPN concepts) sufficient to test connectivity, access, and security requirements.
- Scripting ability in Bash and/or Python, for building test harnesses and automating repetitive validation tasks.
- Experience with performance/load testing tools and methodology (e.g. JMeter, Locust, k6, or comparable).
- Familiarity with version control (Git) and issue tracking systems (e.g. Jira, Azure).
- Strong troubleshooting and root-cause analysis skills spanning OS, network, and application layers.
- Clear written communication — able to produce test plans, test cases, and defect/qualification reports that non-specialists can follow.
PREFERRED EXPERIENCE
- Experience testing industrial/OT systems or SCADA platforms (e.g. Ignition, PLCs, MQTT/OPC UA) — directly relevant to our environment.
- Deeper Kubernetes/OpenShift operational experience — cluster administration, Helm charts, or similar.
- Exposure to infrastructure-as-code tooling (Ansible, Terraform).
- RHEL certification
- Experience with monitoring/observability tooling (Prometheus, Grafana, Nagios, Zabbix, or similar).
- Security testing awareness — vulnerability scanning, hardening validation, or familiarity with CIS benchmarks for RHEL.
- Experience with virtualization platforms (VMware, KVM).
- Prior experience in energy, utilities, or other critical-infrastructure/industrial sectors.
- Experience with cloud platforms (Azure, AWS), especially hybrid on-prem/cloud environments.
- Familiarity with database testing (PostgreSQL, MySQL, or similar).
- Experience with load balancers or reverse proxies (HAProxy, NGINX).
#LI-AD1
For US-based roles - What you’ll get:
- Competitive salary + annual performance-based bonus eligibility
- Medical, dental, and vision insurance
- 401(k) with company match
- Paid time off and company holidays
For Mexico-based roles - What you’ll get:
- Competitive salary + annual performance bonus eligibility
- Christmas Bonus (Aguinaldo): 30 days
- Major medical expenses and life insurance
- Paid time off and holidays (per local policy)
For all roles:
- Professional development and growth opportunities
- Opportunity to grow with a mission-driven team shaping the future of clean energy
- Equal Opportunity: ON.energy is committed to equal employment opportunity and to maintaining a work environment free of harassment, discrimination, or retaliation.
- Benefits vary by role and location and are subject to change.
Agency Notice: ON.energy does not accept unsolicited resumes from staffing agencies, search firms, or third-party recruiters. Resumes submitted without a fully executed Master Services Agreement (MSA) and a written request from an authorized member of our Talent Acquisition team will be considered the property of ON.energy. No placement fees or compensation will be paid for unsolicited candidate submissions.
Skills Required
- 5+ years of hands-on Linux administration experience, specifically with Red Hat Enterprise Linux, including system administration, package management, systemd, and shell scripting.
- Experience testing or validating Docker container deployments and knowledge of Kubernetes or OpenShift orchestration fundamentals.
- Solid QA/testing experience designing test plans and cases, executing structured test cycles, and tracking defects to resolution.
- Experience with CI/CD pipelines and validating automated build and deployment processes.
- Working knowledge of TCP/IP, DNS, firewalls, and VPN concepts.
- Bash and/or Python scripting ability for test harnesses and automation.
- Experience with performance and load testing tools and methodology.
- Familiarity with Git and issue tracking systems such as Jira or Azure.
- Strong troubleshooting and root-cause analysis skills across OS, network, and application layers.
- Clear written communication and ability to produce test and defect documentation for non-specialists.
- Experience testing industrial or OT systems or SCADA platforms such as Ignition, PLCs, MQTT, or OPC UA.
- Deeper Kubernetes or OpenShift operational experience, including cluster administration or Helm charts.
- Exposure to infrastructure-as-code tools such as Ansible or Terraform.
- RHEL certification.
- Experience with monitoring and observability tools such as Prometheus, Grafana, Nagios, or Zabbix.
- Security testing awareness, including vulnerability scanning, hardening validation, or CIS benchmarks for RHEL.
- Experience with virtualization platforms such as VMware or KVM.
- Prior experience in energy, utilities, critical infrastructure, or industrial sectors.
- Experience with Azure, AWS, or hybrid on-premises/cloud environments.
- Familiarity with database testing, including PostgreSQL or MySQL.
- Experience with load balancers or reverse proxies such as HAProxy or NGINX.
What We Do
ON.energy is building the backbone of energy and AI infrastructure, powering grid-safe data centers and mission-critical facilities. The company supplies and operates hyperscale power systems that solve the toughest resilience challenges, delivering custom solutions for AI data centers, mission-critical facilities, and front-of-the-meter assets. Its track record spans industrial, manufacturing, infrastructure, transportation, and grid-scale storage. With patented technology and proprietary software, ON.energy develops projects worldwide that set new benchmarks for resilience.






