The Role
Design, build, and troubleshoot large-scale distributed, event-driven cloud services. Manage Linux/Windows VM fleets, automate infrastructure (IaC/CaC), own reliability, security, performance and SLOs, participate in on-call rotations, and collaborate on monitoring, automation, and security response.
Summary Generated by Built In
Responsibilities
KEY RESPONSIBILITIES
- Hands-on design, analysis, development and troubleshooting of highly distributed large-scale production systems and event-driven, cloud-based services
- Primarily Linux Administration, managing a fleet of Linux and Windows VMs as part of the application solutions
- Involved in Pull Requests for site reliability goals
- Advocate IaC (Infrastructure as Code) and CaC (Configuration as Code) practices within Honeywell HCE
- Ownership of reliability, up time, system security, cost, operations, capacity and performance-analysis
- Monitor and report on service level objectives for a given applications services. Work with the business, Technology teams and product owners to establish key service level indicators.
- Ensuring the repeatability, traceability, and transparency of our infrastructure automation
- Support on-call rotations for operational duties that have not been addressed with automation
- Support healthy software development practices, including complying with the chosen software development methodology (Agile, or alternatives), building standards for code reviews, work packaging, etc.
- Create and maintain monitoring technologies and processes that improve the visibility to our applications' performance and business metrics and keep operational workload in-check.
- Partnering with security engineers and developing plans and automation to aggressively and safely respond to new risks and vulnerabilities.
- Develop, communicate, collaborate, and monitor standard processes to promote the long-term health and sustainability of operational development tasks.
- Participate in technical training events, game day scenarios, and professional conferences
YOU MUST HAVE
- 3 Years of experience in system administration, application development, infrastructure development or related areas
- 3 years of experience with programming in languages like Javascript, Python, PHP, Go, Java or Ruby
- 3 years Mastery of infrastructure automation technologies (like Terraform, CodeDeploy, Puppet, Ansible, Chef)
- 5+ years Cloud and container native Linux administration/build/management skills
- 3+ years expertise in container/container-fleet-orchestration technologies (like Kubernetes, Openshift, AKS, EKS, Docker, Vagrant, etcd, zookeeper)
Skills Required
- 3 years experience in system administration, application development, infrastructure development or related areas
- 3 years programming experience in Javascript, Python, PHP, Go, Java or Ruby
- 3 years mastery of infrastructure automation technologies (Terraform, CodeDeploy, Puppet, Ansible, Chef)
- 5+ years cloud and container-native Linux administration, build and management skills
- 3+ years expertise in container and container-fleet orchestration (Kubernetes, Openshift, AKS, EKS, Docker, Vagrant, etcd, zookeeper)
- Hands-on experience designing, developing and troubleshooting highly distributed large-scale production systems and event-driven cloud services
- Experience with Linux administration and managing fleets of Linux and Windows virtual machines
- Proven experience advocating and implementing Infrastructure as Code (IaC) and Configuration as Code (CaC) practices
- Experience creating and maintaining monitoring, SLO/SLA reporting and performance visibility for application services
- Willingness to participate in on-call rotations and support operational duties
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company






