Site Reliability Engineer II (NGPOS Operations Support)

Reposted 2 Days Ago
Cincinnati, OH, USA
In-Office
50-52 Hourly
Mid level
Agency • Information Technology • Professional Services • Software
The Role
Lead incident response and serve as Incident Commander for P1/P2 outages, run RCA and corrective actions, improve observability and SLIs/SLOs, reduce alert fatigue, support retail store deployments, create runbooks/playbooks, collaborate with engineering teams, and participate in on-call rotations and maintenance windows.
Summary Generated by Built In

Position Overview

We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support.

The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.

This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements.

Location: Blue Ash, OH (Cincinnati) – Onsite (5 Days/Week)
Employment Type: W2 – Contract-to-Hire
Duration: Full-Time
Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship.
Experience Required: 3+ Years

Nice to Have

  • Enterprise Point of Sale (POS) systems

  • Retail technology experience

  • Automation scripting

  • Monitoring optimization

  • Runbook creation

  • Store technology deployments

Key Responsibilities

  • Lead major incident response during production outages

  • Serve as Incident Commander during P1/P2 incidents

  • Coordinate technical bridge calls

  • Communicate outage status to engineering teams and business leadership

  • Lead Root Cause Analysis (RCA) activities

  • Track corrective actions through completion

  • Improve production reliability and system stability

  • Enhance monitoring and observability

  • Reduce alert fatigue

  • Partner with Software Engineering and Platform Engineering teams

  • Support retail store deployments

  • Develop operational documentation, runbooks, and playbooks

  • Participate in after-hours support rotations and maintenance windows

  • Improve service health using SLIs and SLOs

Technical Environment

Monitoring & Observability

  • Dynatrace

  • Azure Monitor

  • Log Analytics

  • Metrics

  • Dashboards

Cloud

  • Microsoft Azure

  • Google Cloud Platform (GCP)

Containers

  • Kubernetes

  • Docker

Operating Systems

  • Linux

Scripting Languages

  • Bash

  • Python

Agile Tools

  • Jira

Enterprise Environment

  • Retail systems

  • Point of Sale (POS)

  • Production Support

  • Hybrid Infrastructure

Ideal Candidate Profile

The ideal candidate will demonstrate:

  • Strong leadership during production incidents

  • Excellent troubleshooting and analytical skills

  • Effective communication under pressure

  • Ownership and accountability

  • Experience coordinating multiple engineering teams

  • Strong operational discipline

  • Continuous improvement mindset

  • Passion for reliability engineering

  • Excellent documentation skills

Required Skills

Must Have

  • Major Incident Management experience

  • Experience serving as Incident Commander

  • Leading production bridge calls

  • Coordinating cross-functional technical teams

  • Executive communication during P1/P2 outages

  • Root Cause Analysis (RCA)

    • Five Whys

    • Fishbone Analysis

    • Timeline reconstruction

    • Corrective action tracking

  • Production troubleshooting across:

    • Cloud environments

    • On-premise infrastructure

    • Retail/POS systems

  • Observability and Monitoring

    • Dynatrace

    • Azure Monitor

    • Log analysis

    • Dashboards

    • Metrics

  • Linux administration

  • Bash and/or Python scripting

  • Kubernetes

  • Docker

  • Microsoft Azure and/or Google Cloud Platform (GCP)

  • Agile methodology

  • Jira

  • Strong communication and collaboration skills

  • Ability to work onsite five days per week

  • Willingness to travel and participate in on-call rotations

Travel Requirements

  • Local travel initially throughout Cincinnati and Louisville

  • Travel will expand as additional retail locations are deployed

  • Approximately every four weeks as new sites go live

  • Shared on-call and travel rotation with a growing six-person team

Skills Required

  • 3+ years experience
  • Major Incident Management experience
  • Experience serving as Incident Commander
  • Leading production bridge calls and coordinating cross-functional teams
  • Executive communication during P1/P2 outages
  • Root Cause Analysis (Five Whys, Fishbone Analysis, timeline reconstruction, corrective action tracking)
  • Production troubleshooting across cloud environments, on-premise infrastructure, and retail/POS systems
  • Observability and monitoring (Dynatrace, Azure Monitor, log analysis, dashboards, metrics)
  • Linux administration
  • Bash and/or Python scripting
  • Kubernetes
  • Docker
  • Microsoft Azure and/or Google Cloud Platform (GCP)
  • Agile methodology experience
  • Jira
  • Strong communication and collaboration skills
  • Ability to work onsite five days per week in Blue Ash, OH
  • Willingness to travel and participate in on-call rotations
  • Permanent resident work authorization (no sponsorship) and ability to convert to full-time
  • Enterprise POS systems experience
  • Retail technology experience
  • Automation scripting (beyond basic Bash/Python)
  • Monitoring optimization experience
  • Runbook creation and operational documentation
  • Store technology deployments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
28 Employees

What We Do

Hudson Information Technology and Manpower Services, part of The Hudson Group, is a global workforce solutions and software services partner founded in 2019. The company combines HudsonIT Consultancy Ltd, which provides enterprise software and technology consulting, with Hudson Manpower Inc, which specializes in comprehensive technical recruitment across various sectors, including Oil & Gas, IT, and Hospitality.

Similar Jobs

Hybrid
2 Locations
289097 Employees
Hybrid
Columbus, OH, USA
289097 Employees
Hybrid
2 Locations
289097 Employees
Hybrid
Westerville, OH, USA
289097 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account