Systems Engineer - Level III

Posted 7 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Financial Services
The Role
Monitor and maintain the operational health of infrastructure, applications, networks, and security services in a hybrid environment. The engineer triages incidents, manages tickets, communicates with stakeholders, supports major incident bridges, and escalates issues within SLA requirements. The role also improves monitoring, alert quality, runbooks, escalation procedures, and incident response through automation, observability, AIOps, and AI-enabled practices. Rotating shifts, off-hours support, and on-call participation are required.
Summary Generated by Built In

With a company culture rooted in collaboration, expertise and innovation, we aim to promote progress and inspire our clients, employees, investors and communities to achieve their greatest potential. Our work is the catalyst that helps others achieve their goals. In short, We Enable Possibility℠.

Role Purpose

The Senior Systems Engineer (NOC Engineer) helps ensure the availability, reliability, and operational health of business-critical infrastructure, applications, and services across a modern hybrid technology environment. This role requires strong technical judgment, proactive operational awareness, clear communication, and the ability to support fast, coordinated response during service-impacting events.

How This Role Has Evolved

Modern eNOC operations now extend beyond incident management and monitoring. Engineers are expected to understand how observability, automation, AIOps, and AI-enabled capabilities can improve detection, reduce alert noise, accelerate troubleshooting, and strengthen service resilience. A strong candidate should be curious about emerging technology trends and willing to adopt new tools and practices that improve operational outcomes.

Core Responsibilities

  • Serve as a first and fast responder for infrastructure, application, network, and security-related operational events.

  • Monitor, triage, and manage global incidents, ensuring events are assessed, prioritized, communicated, and escalated within accepted SLAs.

  • Use modern monitoring, observability, alerting, and incident management platforms to identify trends, reduce alert noise, and detect potential service degradation before customers are impacted.

  • Open, update, and manage tickets with clear documentation, accurate impact details, and timely stakeholder communications throughout the incident lifecycle.

  • Lead or support major incident bridges during production-impacting events, coordinating with technical teams, service owners, and business stakeholders.

  • Partner with engineering, application, infrastructure, and service management teams to improve operational runbooks, escalation paths, alert quality, and response procedures.

  • Stay current with emerging IT operations trends, including AI, AIOps, automation, predictive monitoring, and intelligent incident response capabilities.

Modern Tools & Capabilities

Engineers should be comfortable working with enterprise monitoring, observability, ticketing, alerting, and incident response platforms. Familiarity with tools such as SolarWinds, ServiceNow, PagerDuty, Splunk, Dynatrace, Azure Monitor, LogicMonitor, and similar platforms is highly valuable, along with an understanding of escalation workflows, alert correlation, event enrichment, runbooks, and operational automation.

AI / AIOps Readiness

A strong Senior NOC Engineer should stay informed about current AI and AIOps trends in the IT operations market, including intelligent alerting, anomaly detection, event correlation, automation-assisted triage, automated remediation, and generative AI use cases for documentation and troubleshooting. The role requires a willingness to learn, evaluate, and responsibly adopt these capabilities where they improve operational quality, speed, and consistency.

Required Experience

  • Ability to manage multiple operational priorities, incidents, and project-related tasks in a fast-paced environment.

  • Excellent verbal and written communication skills with the ability to provide clear, concise, and timely updates to technical and non-technical stakeholders.

  • Strong customer service mindset with the ability to build positive and collaborative relationships across technology and business teams.

  • Strong troubleshooting, analytical thinking, and problem-solving skills with the ability to assess impact, urgency, and appropriate escalation paths.

  • Experience documenting incidents, actions taken, timelines, communications, and resolution details in enterprise ticketing systems.

  • Working knowledge of enterprise infrastructure, including Windows Server, Linux, networking fundamentals, DNS, DHCP, firewalls, VPN, load balancing, and cloud or hybrid environments.

  • Familiarity with monitoring, observability, and incident response platforms such as SolarWinds, ServiceNow, PagerDuty, Splunk, Dynatrace, Azure Monitor, or similar tools.

  • Understanding of alert correlation, event enrichment, escalation workflows, runbooks, and operational automation concepts.

  • Willingness to learn and adopt AI-enabled operational practices, including AIOps, intelligent alerting, anomaly detection, automation-assisted troubleshooting, and AI-supported documentation.

  • Organized, detail-oriented, self-motivated, and able to remain engaged through incident closure and post-incident follow-up activities.

  • Ability to participate in rotating shifts, off-hours support, and on-call responsibilities as required.

Preferred Experience

  • Experience with Linux, Windows Server, virtualization, cloud platforms, and enterprise infrastructure support.

  • Knowledge of networking concepts and technologies, including routing, switching, DNS, DHCP, VPN, firewalls, load balancers, and secure connectivity.

  • Hands-on familiarity with operational tools such as SolarWinds, ServiceNow, PagerDuty, , Azure Monitor, Splunk, Dynatrace, LogicMonitor, or similar platforms.

  • Exposure to automation or scripting concepts using PowerShell, Python, REST APIs, workflow automation, or low-code automation platforms.

  • Awareness of AI and AIOps market trends, including predictive monitoring, anomaly detection, event correlation, automated remediation, and generative AI use cases in IT operations.

  • Ability to contribute to operational improvement initiatives, including alert tuning, runbook development, knowledge base documentation, and post-incident review improvements.

Education

  • Bachelor's degree or equivalent working experience

Success Measures

  • Incidents are assessed, communicated, escalated, and documented accurately and within expected response timelines.

  • Monitoring and alerting practices continue to improve through reduced noise, better correlation, and clearer operational visibility.

  • Runbooks, knowledge articles, and escalation procedures remain current, practical, and easy for the team to use during live events.

  • The engineer actively contributes to continuous improvement by adopting relevant tools, automation, and AI-enabled practices that strengthen eNOC operations.

#LI-RS1


Do you like solving complex business problems, working with talented colleagues and have an innovative mindset? Arch may be a great fit for you. If this job isn’t the right fit but you’re interested in working for Arch, create a job alert! Simply create an account and opt in to receive emails when we have job openings that meet your criteria. Join our talent community to share your preferences directly with Arch’s Talent Acquisition team.


10400 Arch Global Services (Philippines) Inc.

Skills Required

  • Ability to manage multiple operational priorities, incidents, and project-related tasks in a fast-paced environment
  • Excellent verbal and written communication skills for timely updates to technical and non-technical stakeholders
  • Strong customer service mindset and ability to build collaborative relationships
  • Strong troubleshooting, analytical thinking, and problem-solving skills
  • Experience documenting incidents, actions, timelines, communications, and resolutions in enterprise ticketing systems
  • Working knowledge of enterprise infrastructure, Windows Server, Linux, networking, DNS, DHCP, firewalls, VPN, load balancing, and cloud or hybrid environments
  • Familiarity with monitoring, observability, and incident response platforms such as SolarWinds, ServiceNow, PagerDuty, Splunk, Dynatrace, or Azure Monitor
  • Understanding of alert correlation, event enrichment, escalation workflows, runbooks, and operational automation
  • Willingness to learn and adopt AI-enabled operational practices, including AIOps and intelligent alerting
  • Organized, detail-oriented, self-motivated, and able to remain engaged through incident closure and post-incident follow-up
  • Ability to participate in rotating shifts, off-hours support, and on-call responsibilities
  • Experience with Linux, Windows Server, virtualization, cloud platforms, and enterprise infrastructure support
  • Knowledge of routing, switching, DNS, DHCP, VPN, firewalls, load balancers, and secure connectivity
  • Hands-on familiarity with SolarWinds, ServiceNow, PagerDuty, Azure Monitor, Splunk, Dynatrace, LogicMonitor, or similar platforms
  • Exposure to PowerShell, Python, REST APIs, workflow automation, or low-code automation platforms
  • Awareness of AI and AIOps trends, predictive monitoring, anomaly detection, event correlation, automated remediation, and generative AI in IT operations
  • Ability to contribute to alert tuning, runbook development, knowledge documentation, and post-incident review improvements
  • Bachelor's degree or equivalent working experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Hamilton
285 Employees
Year Founded: 2001

What We Do

Arch Capital Group Ltd. (Arch Capital or ACGL), a Bermuda public limited liability company, writes insurance and reinsurance on a worldwide basis through operations in Bermuda, the United States, Canada, Europe and Australia, with a focus on specialty lines. Arch Capital Services LLC is owned by ACGL and provides corporate, legal and other support services to Arch Capital. ACGL provides insurance, reinsurance and mortgage insurance on a worldwide basis through operations in Bermuda, the United States, Canada, Europe, Australia and Hong Kong.

Similar Jobs

Remote or Hybrid
2 Locations
289097 Employees
Remote or Hybrid
2 Locations
289097 Employees

Smartly Logo Smartly

Paid Social Specialist

AdTech • Artificial Intelligence • Digital Media • Marketing Tech • Social Media • Software • Generative AI
Easy Apply
Remote or Hybrid
Philippines
805 Employees

Duda, Inc. Logo Duda, Inc.

Technical Support

Agency • Digital Media • eCommerce • Marketing Tech • Software • Design • App development
Remote or Hybrid
Philippines
200 Employees

Similar Companies Hiring

Granted Thumbnail
Artificial Intelligence • Healthtech • Insurance • Mobile • Financial Services
New York, New York
23 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account