Some of the world’s most innovative global software and technology companies struggle to find engineering partners capable of stepping into complex environments and immediately driving meaningful outcomes. These teams need more than additional hands—they need senior engineers who can quickly understand an environment, identify the path forward, and execute without constant direction.
Enter EverOps – the premier Embedded Service Provider. We partner directly with customer engineering teams to assess and address mission-critical infrastructure, cloud, and delivery challenges.
The ChallengeEverOps is looking for a Lead DevOps Engineer with deep AWS infrastructure experience and unusually strong networking expertise to support a complex cloud migration and modernization initiative within a high-transaction payments environment.
You’ll be joining an active migration already in motion, where timelines are compressed, dependencies are not always fully documented, and the environment spans legacy infrastructure, AWS, application platforms, databases, networking, and production operations.
This role requires someone who can get productive quickly, work through ambiguity, and independently turn broad objectives into executable technical work.
The MissionAs a Lead DevOps Engineer, you will join our U.S.-Based Virtual Operating Center and embed directly with a customer engineering team.
Your immediate priority will be providing hands-on engineering leadership and execution for an active migration from legacy hosted infrastructure into AWS. You’ll troubleshoot migration blockers, assess network and infrastructure dependencies, build cloud infrastructure, and help move production workloads safely.
As the immediate migration effort stabilizes, your focus will expand into modernizing legacy CentOS workloads, building the supporting deployment and observability platform, and implementing AWS-based disaster recovery for a business-critical payments platform.
This is a highly autonomous role. You will be expected to identify what needs to happen, define technical deliverables, communicate risks and dependencies clearly, and drive work through completion without requiring step-by-step direction.
What You’ll DoCloud Migration: Provide hands-on engineering support for the migration of production workloads from legacy hosted infrastructure into AWS.
Network Engineering: Diagnose and resolve complex connectivity, routing, DNS, firewall, VPN, load-balancing, security-group, and hybrid-networking issues impacting migrations and production systems.
Migration Planning: Assess undocumented or partially documented environments, identify dependencies and blockers, and translate findings into practical migration plans and technical workstreams.
AWS Infrastructure: Design, build, troubleshoot, and improve production AWS environments using modern infrastructure-as-code practices.
Legacy Modernization: Help retire legacy CentOS 7 systems and migrate applications onto a modern, supportable platform.
Platform Engineering: Implement and improve CI/CD, monitoring, observability, secrets management, configuration management, and deployment automation.
Disaster Recovery: Design and implement AWS-based disaster recovery capabilities, including infrastructure, replication dependencies, recovery procedures, and operational runbooks.
Production Cutovers: Support migration rehearsals, rollback planning, production cutovers, validation, and post-migration stabilization.
Technical Ownership: Independently define and execute technical deliverables, surface risks early, and drive issues to resolution.
Documentation: Produce useful architecture documentation, migration plans, operational runbooks, dependency maps, and knowledge-transfer materials.
Technical Leadership: Serve as a senior technical partner to customer engineers and EverOps team members, providing direction when ambiguity or complex infrastructure decisions arise.
Experience: 7+ years of professional experience in DevOps, Cloud Engineering, SRE, Infrastructure Engineering, or a related discipline, with significant production AWS experience.
AWS Expertise: Deep hands-on experience designing, operating, troubleshooting, and migrating production workloads in AWS.
Networking Depth: Strong knowledge of TCP/IP, routing, subnetting, DNS, NAT, firewalls, VPNs, proxies, load balancers, security groups, network ACLs, and AWS networking services.
AWS Networking: Production experience with VPC architecture, Transit Gateway, Route 53, ALB/NLB, PrivateLink/VPC endpoints, VPN connectivity, and multi-account or hybrid-network environments.
Migration Experience: Proven experience executing data-center, hosted-infrastructure, lift-and-shift, re-platforming, or cloud-modernization migrations involving live production workloads.
Infrastructure as Code: Advanced proficiency with Terraform and experience managing production infrastructure through version-controlled IaC.
Linux: Strong Linux systems administration and troubleshooting skills, including experience with legacy environments and operating-system modernization.
Containers: Production experience with Docker and container orchestration platforms such as ECS, EKS, or Kubernetes.
CI/CD: Experience designing and operating modern CI/CD pipelines using tools such as GitHub Actions, Jenkins, Argo CD, or similar platforms.
Observability: Experience implementing and troubleshooting monitoring, logging, metrics, and alerting platforms such as Datadog, Prometheus, Grafana, or comparable tooling.
Automation: Strong scripting ability using Python, Bash, or similar languages to automate infrastructure and operational workflows.
Production Operations: Experience supporting business-critical production environments where uptime, change control, and careful migration planning matter.
Autonomy: Demonstrated ability to enter an unfamiliar environment, independently identify priorities, define a path forward, and execute with minimal supervision.
Communication: Ability to clearly explain technical risks, dependencies, tradeoffs, and recommendations to both engineers and technical leadership.
Fintech / Payments: Experience working in payments, financial services, banking, or another highly regulated, high-transaction environment.
Rackspace / Hosted Infrastructure: Experience migrating workloads out of Rackspace or similar managed hosting / colocation environments.
Migration Leadership: Experience owning end-to-end infrastructure migrations, including discovery, dependency mapping, rehearsal, cutover, rollback, and stabilization.
Disaster Recovery: Hands-on experience designing and implementing AWS DR environments, replication strategies, recovery runbooks, and RTO/RPO objectives.
Database Infrastructure: Familiarity with infrastructure supporting Aurora/MySQL, PostgreSQL, Cassandra, or other distributed data platforms.
Security: Experience operating within environments involving PCI, SOC 2, or similar security and compliance requirements.
GitOps: Experience with GitOps and pull-request-driven infrastructure workflows using tools such as Atlantis, Terraform Cloud/Enterprise, Scalr, or Argo CD.
Platform Engineering: Experience building internal platforms that standardize deployment, observability, secrets management, and infrastructure consumption.
Certifications: AWS Certified Solutions Architect – Professional, AWS Certified Advanced Networking – Specialty, CKA, or similar advanced certifications.
100% Remote Workplace: We’ve been remote since Day 1!
Unlimited Paid Time Off.
Equity: Become a true owner of the company.
401K with company contribution and sponsored healthcare.
Professional Growth: Access to training and certification programs to accelerate your career.
Skills Required
- 7+ years professional experience in DevOps, Cloud Engineering, SRE, Infrastructure Engineering, or related discipline
- Significant production AWS experience; hands-on designing, operating, troubleshooting, and migrating production workloads in AWS
- Strong networking knowledge: TCP/IP, routing, subnetting, DNS, NAT, firewalls, VPNs, proxies, load balancers, security groups, network ACLs
- AWS networking production experience: VPC architecture, Transit Gateway, Route 53, ALB/NLB, PrivateLink/VPC endpoints, VPN connectivity, multi-account/hybrid networking
- Proven migration experience (data-center, hosted-infrastructure, lift-and-shift, re-platforming, cloud-modernization) with live production workloads
- Advanced proficiency with Terraform and managing production infrastructure via version-controlled IaC
- Strong Linux systems administration and troubleshooting, including legacy OS modernization (e.g., CentOS 7)
- Production experience with Docker and container orchestration (ECS, EKS, or Kubernetes)
- Experience designing and operating CI/CD pipelines (GitHub Actions, Jenkins, Argo CD, or similar)
- Experience implementing and troubleshooting observability: monitoring, logging, metrics, alerting (Datadog, Prometheus, Grafana, or comparable)
- Strong scripting/automation skills using Python, Bash, or similar languages
- Experience supporting business-critical production environments with emphasis on uptime, change control, and careful migration planning
- Ability to operate autonomously, define technical deliverables, communicate risks, and drive work to completion with minimal supervision
- Experience in payments/fintech or regulated high-transaction environments
- Experience migrating workloads out of Rackspace or similar hosted/colocation providers
- Hands-on AWS disaster recovery design and implementation (replication strategies, runbooks, RTO/RPO)
- Familiarity with database infrastructure: Aurora/MySQL, PostgreSQL, Cassandra, or distributed data platforms
- Experience operating within PCI, SOC 2, or similar security and compliance environments
- Experience with GitOps and pull-request-driven infrastructure workflows (Atlantis, Terraform Cloud/Enterprise, Scalr, Argo CD)
- Advanced certifications (AWS Solutions Architect Pro, AWS Advanced Networking, CKA) or similar
What We Do
Introducing a New Kind of Partner: THE EMBEDDED SERVICE PROVIDER A PARTNER THAT CAN PERFORM COMPLEX DELIVERY AS PART OF YOUR TEAM Companies have a lot of trouble finding partners that can perform complex deliveries and services. A partner that can co-own problems from within their organization. Enter the Embedded Service Provider: An ESP performs a service from within the client team structure. THE EVEROPS TECHPOD For It Operations, Production DevOps and Identity Our TechPod model is what allows us to take on complex parts of your technology from within your team structure. As part of every contract, you get all TechPod elements: - Pod Leader - Architect - Engineering - Project work as part of the monthly cost - Operations ENGINEERED OPERATIONS The foundation of our TechPods is our Engineered Operations group: The relentless pursuit of applying engineering & automations to operations functions. All clients benefit from: - EverOps Labs - Speeds architecting and validates deployments - EverOps GitOps models - EverOps Alternative Compute models - EverOps ZeroTrust models for corp & engineering - EverOps Cloud Governance models - EverOps Deployment Automation - EverOps Site Reliability Engineering - EverOps NOC Automation-monitoring -> Alerting -> Slack / Pagerduty - EverOps Site build & PM templates









