IT Infrastructure Support Site Reliability Engineer II

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Ireland, IRL
Remote
51K-63K Annually
Senior level
Information Technology
The Role
Build and maintain infrastructure automation, Infrastructure-as-Code, monitoring, CMDB, and network-management systems for servers, cloud environments, Cisco networks, and physical security devices. Develop remediation, patching, compliance, provisioning, logging, and ticketing workflows; create monitoring exporters and dashboards; perform advanced troubleshooting and root-cause analysis; and lead postmortems. Participate in a 24x5 on-call rotation supporting mission-critical infrastructure.
Summary Generated by Built In

About the Job 

We are seeking an experienced Infra Automation Engineer to join our IT Infrastructure Support team, responsible for ensuring the reliability, scalability, and performance of critical physical security infrastructure, including IP camera fleets, access control systems, and a large-scale Cisco switch fleet,  and the servers, network, and cloud environment that support them. In this role, you will combine software engineering expertise with operations knowledge to build and maintain automation tools, a centralized CMDB, monitoring systems, and processes that support enterprise-grade server, network, and security device management within a large-scale, cloud-hosted enterprise environment. You will work closely with cross-functional teams to define and enforce service level objectives, reduce operational toil through automation, and drive continuous improvement in system resilience. This position requires 24x5 availability with on-call rotation to ensure uninterrupted support for mission-critical infrastructure. 


Key Responsibilities :


Partner with leadership to establish, monitor, and enforce Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for infrastructure tooling, including configuration compliance rates, patch success rates, and deployment latency metrics. 

Provide Level 3 expertise for tooling-specific incidents, focusing on automating incident remediation workflows and reducing Mean Time To Repair (MTTR) through intelligent automation and runbook development. 

Identify and automate repetitive manual tasks across managed infrastructure, targeting measurable reductions in operational overhead (e.g., 50% reduction in manual server build time) through scripting and workflow automation. 

Conduct thorough root cause analysis and lead blameless postmortems for all major service-impacting incidents, driving systemic improvements in tooling reliability and infrastructure resilience. 

Engineer and maintain automated processes and scripts to populate, update, and synchronize asset management platforms (e.g., NetBox), configuration management databases, and monitoring systems for internal and external stakeholders. 

Design, develop, and deploy full-stack applications, custom plugins, and automation scripts to extend functionality of management and monitoring systems, enabling direct device interaction for configuration management. 

Develop and maintain fully automated Infrastructure-as-Code configurations for Windows and Linux server roles using tools such as Ansible, Terraform, or Puppet, including drift detection and auto-remediation capabilities. 

Build end-to-end automation pipelines for vulnerability patching, security baseline enforcement (CIS benchmarks), and continuous compliance auditing against internal and regulatory standards for physical security devices. 

Develop API-driven tools for network configuration management, automated firmware updates, zero-touch provisioning, pre/post-change validation, and real-time network health monitoring across the device fleet. 

Deploy and standardize monitoring agents, centralized log collection systems, and custom dashboards with alerts based on critical SLIs (latency, error rate, saturation, traffic) for servers and edge devices. 

Build and maintain custom monitoring exporters for the physical security device fleet, including camera systems, ensuring accurate metrics and structured, multi-severity logging output. 

Build diagnostic tooling to correlate timestamps across distributed log streams and detect clock drift or NTP desync, a recurring root cause of false-positive outages across the device fleet. 

Build automation scripts for intelligent ticket handling, problem validation, and escalation workflows within enterprise ticketing systems, ensuring 2-hour initial response SLAs are consistently met. 

Support foundational security improvements across the device fleet, including managed credential/access controls and automated configuration backup. 

Participate in 24x5 on-call rotation to provide timely support for infrastructure systems, security devices, and related tooling, ensuring service continuity and rapid incident response. 


Required Skills :


6+ years of experience in Infra Automation Engineering, or Infrastructure Engineering. 

Strong proficiency in Python, Bash, and PowerShell for automation scripting, with experience in Go for building high-performance backend services and APIs. 

Hands-on experience with Infrastructure-as-Code tools (Terraform, Ansible, Chef, or Puppet) and configuration management practices, including drift detection, version control, and automated remediation. 

Advanced knowledge of Linux and Windows server environments, including Tier 3 troubleshooting capabilities, system hardening, and enterprise-scale server management. 

Solid understanding of enterprise networking concepts, Cisco device administration, network automation protocols (NETCONF/RESTCONF), and experience with network monitoring and flow analysis tools. 

Experience implementing and managing monitoring solutions (Prometheus, Grafana, Datadog) or comparable proprietary internal monitoring and metrics-streaming systems (e.g., Monarch, Streamz), and centralized logging platforms (ELK Stack), with ability to create custom dashboards and alerting rules. 

Experience deploying and customizing a CMDB/IPAM platform (e.g., NetBox) as a source of truth for device inventory and downstream automation. 

Comfort operating within a large-scale, cloud-hosted enterprise environment (Kubernetes, Terraform, Helm), including familiarity with internal development and code-review tooling (e.g., Cider, Critic) or comparable large-scale internal toolchains. 

Experience writing and maintaining custom monitoring exporters/agents for edge/IoT and physical security devices, including structured, glog-style logging output. 

Proficiency in advanced text-processing and scripting (e.g., awk/gawk) for log parsing and timestamp correlation, with working knowledge of NTP/clock synchronization practices.

Salary Range

50,640.00 - 63,300.00 EUR (Annual)
  • Please note that the salary information provided herein is base pay only (gross); it does not include other forms of compensation which may or may not apply to this specific position, namely, performance-based bonuses, benefits-related payments, or other general incentives - none of which are guaranteed, may be subject to specific eligibility requirements, and are wholly within the discretion of Astreya to remit.
  • Further, the salary information noted above is a range that consists of a minimum and maximum rate of pay for this specific position. Where an applicant or employee is placed on this range will depend and be contingent on objective, documented work-related considerations like education, experience, certifications, licenses, preferred qualifications, among other factors.

Skills Required

  • 6+ years of experience in infrastructure automation engineering or infrastructure engineering
  • Strong proficiency in Python, Bash, and PowerShell
  • Experience with Go for backend services and APIs
  • Hands-on experience with Terraform, Ansible, Chef, or Puppet
  • Experience with configuration management, drift detection, version control, and automated remediation
  • Advanced knowledge of Linux and Windows server environments
  • Tier 3 troubleshooting, system hardening, and enterprise-scale server management experience
  • Understanding of enterprise networking, Cisco device administration, NETCONF/RESTCONF, network monitoring, and flow analysis
  • Experience implementing and managing Prometheus, Grafana, Datadog, or comparable monitoring systems
  • Experience with centralized logging platforms such as the ELK Stack
  • Ability to create custom monitoring dashboards and alerting rules
  • Experience deploying and customizing a CMDB/IPAM platform such as NetBox
  • Experience with Kubernetes, Terraform, and Helm in a large-scale cloud-hosted enterprise environment
  • Familiarity with internal development and code-review tooling or comparable large-scale toolchains
  • Experience maintaining monitoring exporters or agents for edge, IoT, and physical security devices
  • Experience with structured glog-style logging output
  • Proficiency with awk/gawk and advanced text processing for log parsing and timestamp correlation
  • Working knowledge of NTP and clock synchronization practices

Astreya Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Astreya and has not been reviewed or approved by Astreya.

  • Healthcare Strength Health coverage includes multiple medical plan options plus dental and vision, complemented by FSAs, an EAP, disability and life insurance, and wellness programs. Feedback suggests these offerings provide solid core protection across many roles.
  • Wellbeing & Lifestyle Benefits Client-site amenities at some large tech campuses can add non-cash value such as meals or on-site perks that enhance the day-to-day experience. Wellness Days and access to learning resources and tuition reimbursement further support overall wellbeing.
  • Flexible Benefits Choice among medical plan types and tax-advantaged accounts enables some customization to individual needs. Some roles also offer remote or flexible work, adding practical flexibility to the total package.

Astreya Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, California
1,958 Employees
Year Founded: 2001

What We Do

Astreya is the leading IT solutions provider for some of the world's most recognizable and innovative organizations. Our journey started in 2001 in the heart of Silicon Valley and reaches thirty-three countries with over 2200+ IT professionals. We enable businesses to make better decisions, achieve operational efficiency and gain a competitive edge. The Astreya advantage is centered around focus and clear- vision, world-class talent, and innovative technology: Creativity is in our DNA. Our dedicated Software and Service Innovation teams bring best-in-class technology and tools to bear for our clients.

Similar Jobs

Astreya Logo Astreya

Site Reliability Engineer

Information Technology
In-Office or Remote
3 Locations
1958 Employees
51K-63K Annually

Mastercard Logo Mastercard

Manager, Product Experience & Design

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Dublin, IRL
38800 Employees

Mastercard Logo Mastercard

Senior Data Scientist

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Dublin, IRL
38800 Employees

Mastercard Logo Mastercard

Senior Network Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Dublin, IRL
38800 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account