Senior Automation & Observability Engineer

Posted 4 Hours Ago
Easy Apply
Be an Early Applicant
Hiring Remotely in United States
Remote or Hybrid
113K-147K Annually
Senior level
Cloud • Information Technology
Support mission critical workloads, modernize IT infrastructure and reduce total cost of ownership.
The Role
Designs and operates enterprise observability, monitoring, telemetry, and automation solutions across infrastructure, applications, IoT platforms, and cloud environments. Responsibilities include managing Grafana and Instana ecosystems, dashboards, alerts, telemetry pipelines, incidents, RCA, databases, and platform health. Automates provisioning, monitoring deployment, remediation, and operational workflows using Python, PowerShell, Bash, Ansible, and Puppet. Supports 24x7 operations, major incidents, ITIL processes, service onboarding, documentation, and cross-functional reliability initiatives.
Summary Generated by Built In

At Ensono, our Purpose is to be a relentless ally, disrupting the status quo and unleashing our clients to Do Great Things!  We enable our clients to achieve key business outcomes that reshape how our world runs. As an expert technology adviser and managed service provider with cross-platform certifications, Ensono empowers our clients to keep up with continuous change and embrace innovation.

We can Do Great Things because we have great Associates. The Ensono Core Values unify our diverse talents and are woven into how we do business. These five traits are the key to achieving our purpose:

Honesty, Reliability, Curiosity, Collaboration, and Passion.

About the role and what you'll be doing: 

We are seeking an experienced IoT / Observability Engineer responsible for monitoring, managing, automating, and optimizing enterprise infrastructure, applications, IoT platforms, and enterprise telemetry ecosystems. The ideal candidate will possess strong expertise in observability platforms, monitoring technologies, automation frameworks, and operational support processes to ensure high availability, reliability, and performance of business-critical systems.

The engineer will support enterprise monitoring operations, FOAK services, ELT platforms, incident management, and automation initiatives while collaborating with Infrastructure, Cloud, Network, Application, and Service Delivery teams.

We want all new Associates to succeed in their roles at Ensono. That's why we've outlined the job requirements below. To be considered for this role, it's important that you meet all Required Qualifications. If you do not meet all of the Preferred Qualifications, we still encourage you to apply. 

Key Responsibilities

Monitoring & Observability

  • Design, implement, and maintain enterprise monitoring and observability solutions.
  • Develop and maintain dashboards, alerts, and visualizations using Grafana.
  • Monitor infrastructure, applications, middleware, and IoT services using IBM Instana, Grafana, SolarWinds, and related observability tools.
  • Configure and manage data collection using Telegraf, Prometheus, and monitoring agents.
  • Analyze metrics, logs, traces, events, and telemetry data to identify performance bottlenecks and service degradation.
  • Support SLO, SLA, and operational health monitoring initiatives.
  • Perform Root Cause Analysis (RCA) and troubleshooting for infrastructure and application issues.

Foak & Enterprise Logging/Telemetry

  • Support onboarding, monitoring, and operational management of FOAK (First Office Application Kit) services and enterprise applications.
  • Configure, validate, and troubleshoot Enterprise Logging & Telemetry (ELT) integrations across infrastructure, middleware, applications, and cloud platforms.
  • Monitor telemetry pipelines, log ingestion, event correlation, and data quality to ensure complete observability coverage.
  • Collaborate with engineering teams to improve telemetry standards, monitoring effectiveness, and proactive incident detection through ELT and observability frameworks.
  • Support FOAK application integrations with Grafana, Instana, Prometheus, and enterprise monitoring platforms.

Infrastructure & Platform Monitoring

Monitor and support:

  • Linux Servers
  • Windows Servers
  • VMware Infrastructure
  • Citrix VDI Platforms
  • DNS Services
  • Proxy Services
  • Middleware Platforms
  • Integration Services
  • Enterprise Applications
  • IoT Platforms

Additional Responsibilities:

  • Investigate performance issues, recurring alerts, and infrastructure anomalies.
  • Validate monitoring platform health and monitoring coverage.
  • Monitor capacity, availability, CPU, memory, storage, and service health metrics.
  • Support platform upgrades, maintenance, and operational readiness reviews.

Database & Data Management

  • Configure and maintain InfluxDB time-series databases.
  • Manage data retention policies, performance tuning, and capacity planning.
  • Develop operational dashboards and reports for infrastructure and application performance insights.
  • Support telemetry data ingestion, storage optimization, and historical trend analysis.

Event & Incident Management

  • Monitor operational alerts, events, notifications, and incidents from enterprise monitoring platforms.
  • Acknowledge, investigate, troubleshoot, and resolve assigned incidents.
  • Coordinate with Infrastructure, Network, Cloud, Security, Application, and Service Delivery teams during incident resolution.
  • Participate in major incident bridges, DR exercises, and 24x7 operations support activities.
  • Follow escalation procedures, SOPs, operational runbooks, and ITIL processes.
  • Support Problem Management activities and contribute to RCA documentation.

Instana & APM Operations

  • Administer and support IBM Instana monitoring environments.
  • Monitor application, API, middleware, and microservices performance using Instana.
  • Validate Instana agent health following server patching and maintenance activities.
  • Support application onboarding and APM configuration standards.
  • Configure alerts, baselines, and performance thresholds.
  • Raise and track incidents related to Instana platform availability and performance.

Automation & Scripting

Develop automation solutions using:

  • Python
  • PowerShell
  • Shell Scripting (Bash)
  • VBScript

 Responsibilities:

  • Automate operational tasks, monitoring deployments, and remediation workflows.
  • Build reusable automation tools to improve operational efficiency.
  • Integrate monitoring platforms with enterprise automation frameworks.
  • Support webhook-based automation and event-driven operational workflows.

Configuration Management & Infrastructure Automation

Implement Infrastructure as Code (IaC) and automation using:

  • Ansible
  • Puppet

Responsibilities:

  • Automate server provisioning and configuration management.
  • Automate monitoring agent deployment and onboarding.
  • Maintain automation playbooks and deployment pipelines.
  • Improve operational consistency and reduce manual efforts across environments.

Knowledge Management

  • Maintain SOPs, runbooks, monitoring procedures, and escalation matrices.
  • Participate in KT sessions, service onboarding, and operational readiness reviews.
  • Support service transition, migration, and continuous improvement initiatives.
  • Maintain observability standards and monitoring documentation.

Required Skills

Monitoring & Observability

  • Grafana
  • IBM Instana
  • SolarWinds
  • Telegraf
  • Prometheus
  • InfluxDB
  • OpenTelemetry
  • Grafana Alloy
  • APM Monitoring
  • Event Management
  • Alert Management
  • Observability Concepts
  • SLO/SLA Monitoring

FOAK & Enterprise Telemetry

  • FOAK (First Office Application Kit) Support
  • Enterprise Logging & Telemetry (ELT)
  • Log Aggregation & Correlation
  • Telemetry Data Analysis
  • Event Correlation
  • Application Onboarding
  • Monitoring Standards & Observability Frameworks

Infrastructure

  • VMware
  • Linux Administration
  • Windows Server
  • Citrix VDI
  • DNS Services
  • Proxy Services
  • Middleware Technologies
  • Infrastructure Performance Monitoring

Scripting & Programming

  • Python
  • PowerShell
  • Shell Scripting (Linux)
  • VBScript

Automation Tools

  • Ansible
  • Puppet
  • Webhooks
  • Infrastructure as Code (IaC)

ITSM & Operations

  • ServiceNow
  • Incident Management
  • Problem Management
  • Change Management
  • ITIL Framework
  • Major Incident Management

Nice to Have

  • Docker
  • Kubernetes
  • AWS
  • Microsoft Azure
  • Google Cloud Platform (GCP)
  • Jenkins
  • GitHub Actions
  • GitLab CI/CD
  • REST APIs
  • Microservices Monitoring
  • DevOps & SRE Practices

Preferred Experience

  • 7 to 10+ years of experience in Monitoring, Observability, Infrastructure Operations, SRE, or Platform Engineering.
  • Experience supporting large-scale enterprise environments and 24x7 operations.
  • Hands-on experience with Grafana, Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB.
  • Experience supporting VMware, Citrix, middleware, enterprise applications, and cloud monitoring platforms.
  • Experience working with FOAK applications and Enterprise Logging & Telemetry (ELT) platforms.
  • Strong troubleshooting, RCA, incident management, and operational support skills.
  • Experience with automation frameworks and Infrastructure as Code (Ansible preferred).
  • Experience integrating observability platforms with enterprise automation solutions.

Key Competencies

  • Observability & Monitoring
  • Infrastructure Automation
  • Enterprise Telemetry Management
  • FOAK Application Support
  • Root Cause Analysis
  • Problem Solving
  • Performance Optimization
  • Service Reliability Engineering (SRE)
  • Cross-Functional Collaboration
  • Operational Excellence
  • Continuous Improvement

Why Ensono?

Ensono is a place to make better happen – for our clients and for your career. You can do great things through innovation or collaboration, by learning or volunteering, or to promote diversity and inclusion. You can do great things for your own health or for a healthier planet. Whatever it means to you to do great things we want Ensono to be the place you can do it. 

We are a client-facing business, but we do encourage clients to allow us to work remotely most of the time so if you are not required to be on a client site, you can choose to work from home or in our Ensono offices.

Some of our benefits include:

  • Unlimited Paid Days Off
  • Three health plan options
  • 401k with company match
  • Eligibility for dental, vision, short and long-term disability, life and AD&D coverage, and flexible spending accounts
  • Family Forming Benefit including fertility coverage and adoption/surrogacy reimbursement
  • Paid childbearing and paternal leave
  • Education Reimbursement, Student Loan Assistance or 529 College Funding
  • Sabbatical leave
  • Wellness program
  • Flexible work schedule

As of the date of this posting, a good faith estimate of the current pay scale for this role is $113,000 to $147,000 annually based on a full-time schedule. Please note that placement in the range may vary based on numerous factors including but not limited to skills, experience, internal equity, and business needs. In addition to base salary, other compensation programs, depending on eligibility, include an annual bonus plan based on company and individual performance and an equity grant under our Associate Equity Appreciation Program.

Ensono is an Equal Opportunity/Affirmative Action employer. We are committed to providing equal employment to our Associates and building a diverse and inclusive workforce. All qualified applicants will be considered without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, protected veteran status, disability, or other legally protected basis, in accordance with applicable law.

Pay transparency nondiscrimination statement/posting OFCCP’s pay transparency policy can be found on OFCCP’s website.

If you need accommodation at any point during the application or interview process, please let your recruiter know or email [email protected].

Skills Required

  • Grafana, IBM Instana, SolarWinds, Telegraf, Prometheus, InfluxDB, OpenTelemetry, and Grafana Alloy
  • APM monitoring, event management, alert management, observability concepts, and SLO/SLA monitoring
  • FOAK support, Enterprise Logging and Telemetry, log aggregation and correlation, telemetry analysis, event correlation, application onboarding, and observability frameworks
  • VMware, Linux administration, Windows Server, Citrix VDI, DNS, proxy services, middleware, and infrastructure performance monitoring
  • Python, PowerShell, Linux shell scripting, and VBScript
  • Ansible, Puppet, webhooks, and Infrastructure as Code
  • ServiceNow, incident management, problem management, change management, ITIL, and major incident management
  • Seven to ten or more years of experience in monitoring, observability, infrastructure operations, SRE, or platform engineering
  • Experience supporting large-scale enterprise environments and 24x7 operations
  • Hands-on experience with Grafana, Instana, SolarWinds, Telegraf, Prometheus, and InfluxDB
  • Experience with VMware, Citrix, middleware, enterprise applications, and cloud monitoring platforms
  • Experience with FOAK applications and Enterprise Logging and Telemetry platforms
  • Strong troubleshooting, root cause analysis, incident management, and operational support skills
  • Experience with automation frameworks and Infrastructure as Code; Ansible preferred
  • Experience integrating observability platforms with enterprise automation solutions
  • Docker, Kubernetes, AWS, Microsoft Azure, Google Cloud Platform, Jenkins, GitHub Actions, GitLab CI/CD, REST APIs, microservices monitoring, and DevOps or SRE practices

Ensono Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Ensono and has not been reviewed or approved by Ensono.

  • Leave & Time Off Breadth Time off provisions include unlimited PTO, paid volunteer time, and a formal sabbatical program. These offerings provide flexibility for rest, community service, and extended renewal.
  • Retirement Support Retirement support includes a 401(k) with company match as part of the core package. This adds long‑term financial value alongside day‑one eligibility for other coverage.
  • Parental & Family Support Family‑forming coverage and paid parental leave are explicitly included, with adoption and surrogacy reimbursement. These benefits support diverse paths to growing a family.

Ensono Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Downers Grove, IL
3,000 Employees
Year Founded: 2015

What We Do

Ensono helps IT leaders be the catalyst for change by harnessing the power of hybrid IT to transform their businesses. Our broad services portfolio from mainframe to cloud, powered by an intelligent governance platform, is designed to help our clients operate for today and optimize for tomorrow. We are award-winning certified experts in AWS & Azure

Why Work With Us

Our culture is collaborative & results-driven. Curiosity, passion, honesty & reliability are values we live by. Career & professional development is encouraged through promotions, learning opportunities, Ensono University - eTalks, training academies, paid tuition and study leave, quarterly Innovator Awards. Thinking Thursdays (no meetings 8 to 12)

Gallery

Gallery

Similar Jobs

Xero Logo Xero

Lead Engineer (Accounting Core Team)

Cloud • Fintech • Information Technology • Machine Learning • Software
Remote or Hybrid
Washington, DC, USA
4500 Employees
206K-258K Annually

ServiceNow Logo ServiceNow

Product Manager

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Addison, TX, USA
29000 Employees

ServiceNow Logo ServiceNow

Senior Customer Success Manager

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
Orlando, FL, USA
29000 Employees
87K-152K Annually

ServiceNow Logo ServiceNow

Sr. Manager, Product Marketing – CRM

Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Remote or Hybrid
San Diego, CA, USA
29000 Employees
149K-261K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account