Principal Platform Engineer -Infrastructure Automation & Agentic Engineering

Posted Yesterday
Be an Early Applicant
Salisbury, NC, USA
Hybrid
163K-245K Annually
Expert/Leader
AdTech • eCommerce • Food • Marketing Tech • Retail
We provide cutting-edge, seamless omnichannel experiences for customers—no matter when, where or how they choose to shop
The Role
Owns enterprise hosting-platform strategy, governance, roadmaps, and cross-domain engineering outcomes. Designs reusable infrastructure automation, IaC pipelines, self-service capabilities, observability, lifecycle controls, and governed generative or agentic AI workflows. Leads evaluation, reliability, security, FinOps, disaster recovery, and operational readiness while influencing senior stakeholders and coaching technical teams. The role spans Windows, Linux, AIX, virtualization, storage, middleware, networking, ServiceNow, and production infrastructure across a large enterprise.
Summary Generated by Built In
Category/Area of Expertise: IT & Technology
Job Requisition: 546158
Address: USA-NC-Salisbury-2110 Executive Drive
Store Code: Infrastructure - Hosting (5118678)
Ahold Delhaize USA, a division of global food retailer Ahold Delhaize, is part of the U.S. family of brands, which includes five leading omnichannel grocery brands - Food Lion, Giant Food, The GIANT Company, Hannaford and Stop & Shop. Our associates support the brands with a wide range of services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more.
Primary Purpose
The Principal Platform Engineer is an enterprise-level individual contributor within the Technology Infrastructure Hosting team. The role shapes hosting-platform vision, strategy, roadmaps, governance, priorities, investment recommendations, and delivery outcomes for complex, business-critical initiatives across multiple teams and systems. It develops reusable modules, golden paths, self-service capabilities, engineering standards, evaluation methods, and operational controls across Windows Server, Active Directory and identity integrations; AIX/POWER and Linux; VMware, Nutanix, Hyper-V and related compute; storage, backup, SAN and NAS; middleware; networking and firewalls; Datadog, Nagios and related observability; ServiceNow CMDB, ITSM and workflow integrations; and adjacent infrastructure services. The engineer remains hands-on with code and production systems and converts approved designs, domain standards, runbooks, lifecycle obligations, and recurring work into secure, scalable, reliable, cost-effective, and version-controlled capabilities using IaC, configuration management, CI/CD, deterministic orchestration, and governed AI where it adds value.
This role operates with broad decision-impacting opportunities over platform engineering priorities and outcomes without direct people-management responsibility. It aligns senior leaders and cross-functional stakeholders, establishes governance for consistent delivery, coaches senior technical leaders and engineers, resolves systemic cross-domain problems, and challenges unsupported completion claims with evidence. Engineering-whether manual, scripted, IaC-based, or AI-assisted-must improve scalability, developer and operator experience, reliability, security, lifecycle currency, recoverability, observability, auditability, service and financial outcomes.
Our flexible/hybrid work schedule includes 3 in-person days in our Salisbury, NC office and 2 remote days.
Applicants must be currently authorized to work in the United States on a full-time basis.
Core Responsibilities
  • Platform strategy and roadmap. Define and execute the enterprise hosting-platform vision, multi-year roadmap, capability priorities, target outcomes, and investment sequencing. Use business analysis, demand, risk, lifecycle, service performance, experience, and cost data to recommend priorities and align senior stakeholders.
  • Governance and portfolio leadership. Establish decision rights, engineering standards, intake and prioritization methods, delivery controls, scorecards, and review forums for consistent platform outcomes. Lead complex initiatives across internal teams and external partners using Lean-Agile and SAFe-aligned planning, user stories, estimation, dependency management, and incremental delivery.
  • Cross-domain platform engineering. Translate approved designs into reusable implementation patterns, automation, and acceptance criteria spanning Windows/AD, AIX/Linux, VMware/Nutanix/Hyper-V, storage/backup/SAN/NAS, middleware, network/firewall, observability, and ServiceNow. Identify upstream and downstream dependencies, raise material design gaps to the accountable design authority, and verify implementation readiness.
  • Platform product and experience management. Treat shared infrastructure capabilities as products with defined users, service levels, adoption targets, roadmaps, documentation, support models, and feedback loops. Enable self-service through service catalogs, APIs, golden paths, and internal developer portal capabilities that reduce friction without weakening controls.
  • Infrastructure as Code and CI/CD. Implement approved designs with domain owners through secure IaC delivery pipelines using version control, peer review, automated validation, policy checks, plan/review/apply controls, secrets handling, artifact traceability, release gates, rollback, and drift detection. Work across Terraform or OpenTofu, Ansible, PowerShell, Python, ARM/Bicep, and applicable platform-native automation.
  • Runbook-to-automation engineering. Identify high-volume, high-risk, and toil-heavy runbooks; decompose procedures into deterministic steps; define prerequisites, approvals, validations, error handling, rollback, evidence capture, and exception paths; then deliver production-ready orchestration.
  • Generative and Agentic AI for operations. Design and implement governed AI-assisted workflows that can interpret approved runbooks, assemble execution plans, invoke tools, preserve state, request approvals, produce evidence, and stop safely when confidence, policy, or environmental conditions are not met.
  • Prompt and context engineering. Create version-controlled system instructions, task prompts, tool descriptions, retrieval/context strategies, structured outputs, and prompt test suites. Manage prompt injection, data-boundary, hallucination, and tool-misuse risks through least privilege, allowlists, approvals, and validation.
  • Agent harnesses and orchestration. Engineer the runtime scaffolding around agents, including tool interfaces, session state, memory, planning, bounded loops, approval policies, observability, error recovery, and human-in-the-loop handoffs. Separate model reasoning from deterministic control logic and privileged execution.
  • Continuous agentic improvement. Establish bounded self-improvement loops in which production telemetry, failed cases, reviewer feedback, and test results propose changes to prompts, policies, tools, or workflows. Require evaluation, versioning, peer review, approval, and controlled rollout before promotion. Do not permit unreviewed self-modification in production.
  • Evaluation and quality engineering. Build offline and pre-production evaluation harnesses for task success, tool choice, policy compliance, grounding, security, latency, cost, failure recovery, and reproducibility. Maintain representative test cases, regression suites, red-team scenarios, and release thresholds.
  • Financial and capacity stewardship. Apply financial analysis and FinOps practices to platform planning and delivery, including demand and capacity forecasting, unit-cost and consumption visibility, cost allocation, vendor and licensing trade-offs, optimization opportunities, and benefit realization. Balance resilience, performance, technical debt, experience, and cost in recommendations.
  • Security, risk, and compliance by design. Coordinate with Security and Governance to embed approved identity, secrets, least privilege, MFA, logging, audit evidence, encryption, policy-as-code, vulnerability controls, and exception-management requirements within automation and agent solutions. Align AI lifecycle controls to approved enterprise risk practices and NIST-aligned governance.
  • Reliability and observability. Define SLIs, SLOs, telemetry, traces, dashboards, alerts, run histories, change evidence, and operational health measures for pipelines and agents. Lead troubleshooting and systemic correction for high-impact platform and automation failures.
  • Platform lifecycle and resilience. Set and enforce lifecycle practices for provisioning, configuration, patching, upgrades, currency, backup, restore, high availability, disaster recovery, decommissioning, documentation, and operational reporting. Ensure lifecycle risk and recovery readiness are visible in roadmaps and governance decisions.
  • Basis Engineering and Definition of Done. Translate domain specifications into pipelines, controls, scorecards, reference implementations, and acceptance criteria. A capability is not done until authoritative inventory and ownership are recorded; supported lifecycle and compatibility are verified; security and privileged-access controls pass; testing covers normal, failure, rollback, and recovery paths; monitoring, logging, SLI/SLO and alert routing are operational; ServiceNow CI, dependency, change and knowledge records are complete; runbooks and support handoffs are current; evidence is retained; and exceptions have accountable owners and expiry dates.
  • Technical leadership and organizational capability. Set the technical bar, lead implementation and operational-readiness reviews, coach senior platform leaders and engineers, and strengthen capability across internal and supplier teams. Independently verify completion claims using artifacts, telemetry, CMDB records, test results, monitoring coverage, recovery evidence, financial outcomes, and change results; communicate material risks, trade-offs, and recommendations to senior leadership.
Required Infrastructure Domain Breadth
Domain
Principal-level evidence expected
Windows and identity
Windows Server lifecycle, Active Directory/GPO, DNS, privileged access, patching, configuration baselines, hybrid identity dependencies, PowerShell/DSC and recovery.
AIX and Linux
AIX/POWER, HMC/PowerVM/NIM, RHEL or equivalent Linux, package and kernel lifecycle, SSH/PAM/sudo, Ansible, filesystem and multipath dependencies, patching and recovery.
Compute and virtualization
VMware vSphere/vCenter/ESXi, Nutanix AOS/Prism, Hyper-V or comparable platforms; capacity, firmware/compatibility, cluster resilience, migration, lifecycle and automation.
Storage, backup and SAN
Block/file/storage services, Fibre Channel zoning and multipath, SAN/NAS lifecycle, backup policy, immutable protection, capacity, replication and tested restore or failover.
Middleware and runtime
Application servers, web tiers, messaging, integration runtimes or comparable middleware; configuration, certificates, dependencies, patching, HA, logging and performance.
Networking and firewalls
IP/DNS/load-balancing dependencies, routing, segmentation, firewall policy, ports and protocols, network change validation, telemetry and rollback coordination.
Observability
Datadog, Nagios or comparable platforms; coverage, tagging, service checks, metrics/logs/traces, SLI/SLO, actionable alerts, dependency suppression, synthetic health and evidence retention.
Service management
ServiceNow CMDB/Discovery, CI identification and reconciliation, completeness/correctness/compliance, service mapping, catalog/workflow, change, incident, problem and knowledge integration.
Required Qualifications
  • 12+ years of progressive infrastructure, platform, site reliability, or infrastructure software engineering experience, including principal-level ownership of complex cross-platform outcomes in large enterprise production environments.
  • Demonstrated production experience across at least four hosting-domain groups-Windows/AD; AIX/Linux; VMware/Nutanix/Hyper-V; storage/backup/SAN; middleware; networking/firewalls; observability; and ServiceNow-with deep hands-on expertise in at least two.
  • Hands-on ability to design, code, test, deploy and operate production automation using Python and at least two of PowerShell, Ansible, Terraform/OpenTofu, ARM/Bicep, shell scripting or comparable technologies.
  • Demonstrated experience applying Generative AI to engineering or operational workflows, including prompt engineering, context management, structured outputs, retrieval grounding, tool calling, and evaluation.
  • Direct experience designing or operating agentic workflows or autonomous/semi-autonomous systems with tool integration, state/memory, human approvals, observability, bounded execution, and failure recovery.
  • Experience creating evaluation harnesses, regression suites, test datasets, red-team cases, and measurable release gates for AI-enabled workflows.
  • Expert ability to define platform strategy and roadmaps, establish governance, lead complex multi-team initiatives, and align senior business and technology stakeholders around priorities and measurable outcomes.
  • Working mastery of Lean-Agile or SAFe principles, platform product management, business analysis, experience design, user stories, estimation, planning, dependency management, and incremental delivery.
  • Demonstrated financial analysis and FinOps capability, including demand and capacity forecasting, unit-cost and consumption analysis, cost allocation, licensing and vendor trade-offs, optimization, and benefits realization.
  • Strong knowledge of platform resource provisioning and configuration, CMDB and service mapping, patch and update management, incident/change/problem management, monitoring and logging, lifecycle and capacity management, documentation and reporting, security administration, backup, disaster recovery, high availability, runbook engineering, and production support.
  • Ability to translate ambiguous operational problems and approved direction into reusable engineering products, measurable outcomes, implementation patterns, and clear technical standards.
  • Proven ability to influence senior engineers, suppliers, technical stakeholders, and leaders and to challenge design assumptions through implementation evidence.
  • Bachelor's degree in computer science, engineering, information systems, or a related field, or equivalent demonstrable experience.
  • Able to work in the office a minimum of 3 days per week including every Tuesday, and Wednesday (excluding holidays & paid time off).
Preferred Qualifications
  • Experience engineering and automating hybrid hosting estates spanning enterprise data centers, retail or edge compute, cloud, managed-service delivery, and multiple supplier boundaries.
  • Experience with Microsoft Agent Framework, Azure AI/Foundry agent services, Semantic Kernel, LangGraph, or comparable orchestration frameworks.
  • Experience integrating ServiceNow CMDB, Discovery, service mapping, catalog/workflow, change, incident and knowledge processes with source control, IaC, observability and operational automation.
  • Experience with container platforms, Kubernetes, GitOps, golden paths, platform engineering, software supply-chain security, and artifact provenance.
  • Familiarity with NIST AI RMF and the Generative AI Profile, CIS Benchmarks, NIST Cybersecurity Framework, zero trust, PCI-related controls, and regulated enterprise environments.
  • Experience building self-service infrastructure products used by multiple engineering teams.
  • Published technical writing, implementation guidance, open-source contributions, conference participation, or a sustained technical community presence.
  • Relevant advanced certifications or directly equivalent demonstrated depth in Microsoft/Windows, Red Hat or IBM AIX, VMware, Nutanix, storage/SAN, networking/security, ServiceNow, observability, cloud, DevOps, Kubernetes or Terraform.

Salary Range: $163,280 - $244,920
Actual compensation offered to a candidate may vary based on their unique qualifications and experience, internal equity, and market conditions. Final compensation decisions will be made in accordance with company policies and applicable laws.
#LI-Hybrid #LI-CW1
At Ahold Delhaize USA, we provide services to one of the largest portfolios of grocery companies in the nation, and we're actively seeking top talent.
Our team shares a common motivation to drive change, take ownership and enable our brands to better care for their customers. We thrive on supporting great local grocery brands and their strategies.
Our associates are the heartbeat of our organization. We are committed to offering a welcoming work environment where all associates can succeed and thrive. Guided by our values of courage, care, teamwork, integrity (and even a little humor), we are dedicated to being a great place to work.
We believe in collaboration, curiosity, and continuous learning in all that we think, create and do. While building a culture where personal and professional growth are just as important as business growth, we invest in our people, empowering them to learn, grow and deliver at all levels of the business.
#BI-Hybrid

Skills Required

  • 12+ years of progressive infrastructure, platform, site reliability, or infrastructure software engineering experience, including principal-level ownership in large enterprise production environments.
  • Production experience across at least four hosting domains, with deep hands-on expertise in at least two.
  • Production automation experience using Python and at least two of PowerShell, Ansible, Terraform or OpenTofu, ARM or Bicep, shell scripting, or comparable technologies.
  • Experience applying generative AI to engineering or operational workflows, including prompt engineering, context management, structured outputs, retrieval grounding, tool calling, and evaluation.
  • Experience designing or operating agentic workflows or autonomous/semi-autonomous systems with tool integration, state or memory, human approvals, observability, bounded execution, and failure recovery.
  • Experience creating AI evaluation harnesses, regression suites, test datasets, red-team cases, and measurable release gates.
  • Expertise defining platform strategy and roadmaps, establishing governance, leading multi-team initiatives, and aligning senior stakeholders.
  • Working mastery of Lean-Agile or SAFe principles, platform product management, business analysis, user stories, estimation, planning, dependency management, and incremental delivery.
  • Financial analysis and FinOps experience, including forecasting, unit-cost analysis, cost allocation, licensing and vendor trade-offs, optimization, and benefits realization.
  • Strong knowledge of infrastructure provisioning, CMDB, service mapping, patch management, IT service management, monitoring, lifecycle management, security administration, backup, disaster recovery, high availability, runbooks, and production support.
  • Ability to translate ambiguous operational problems into reusable engineering products, measurable outcomes, implementation patterns, and technical standards.
  • Ability to influence senior engineers, suppliers, technical stakeholders, and leaders and challenge assumptions through evidence.
  • Bachelor's degree in computer science, engineering, information systems, or related field, or equivalent demonstrable experience.
  • Ability to work in the Salisbury, North Carolina office at least three days per week, including Tuesdays and Wednesdays.
  • Experience automating hybrid hosting estates across data centers, retail or edge compute, cloud, managed services, and multiple suppliers.
  • Experience with Microsoft Agent Framework, Azure AI Foundry agent services, Semantic Kernel, LangGraph, or comparable orchestration frameworks.
  • Experience integrating ServiceNow with source control, IaC, observability, and operational automation.
  • Experience with container platforms, Kubernetes, GitOps, golden paths, platform engineering, software supply-chain security, and artifact provenance.
  • Familiarity with NIST AI RMF, NIST Cybersecurity Framework, CIS Benchmarks, zero trust, PCI controls, and regulated enterprise environments.
  • Experience building self-service infrastructure products used by multiple engineering teams.
  • Published technical writing, open-source contributions, conference participation, or sustained technical community presence.
  • Relevant advanced certifications or equivalent depth across infrastructure, cloud, DevOps, Kubernetes, Terraform, ServiceNow, observability, or related domains.

What the Team is Saying

Claire Peters

Ahold Delhaize USA Compensation & Benefits Highlights

  • Healthcare Strength Core medical, dental, and vision plans are administered centrally under formal welfare plan documents, pointing to a mature, comprehensive offering. Materials also highlight options like HSAs/FSAs, mental-health and EAP resources, and in some cases inclusive provisions such as abortion travel benefits.
  • Retirement Support Multiple brands promote a 401(k) savings plan with company match, with examples published by Food Lion for eligible associates. Unionized banners such as Stop & Shop additionally reference maintained pension benefits, reinforcing depth in long-term financial support.
  • Parental & Family Support Company and brand pages describe paid parental leave, adoption/surrogacy assistance, and guidance around life events such as childbirth and caregiving. Some banners also note inclusive family benefits like domestic-partner coverage and fertility support.

Ahold Delhaize USA Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Chicago, IL
10,000 Employees
Year Founded: 2018

What We Do

Ahold Delhaize USA, a division of global food retailer Ahold Delhaize, is part of the U.S. family of brands, which includes five leading omnichannel grocery brands – Food Lion, Giant Food, The GIANT Company, Hannaford and Stop & Shop. Our associates support the brands with a wide range of services, including Finance, Legal, Sustainability, Commercial, Digital and E-commerce, Technology and more. Our team includes some of the best and brightest talent from a variety of backgrounds, ranging from decades-long careers in retail to fresh perspectives from outside our industry. With a purpose-driven culture grounded in our values of courage, care, integrity, teamwork and humor, we are committed to fostering a culture of belonging where everyone is valued. Our team shares a common motivation to drive change, take ownership and enable the brands we support to nourish their customers and communities. We thrive on supporting great local grocery brands and their strategies. As part of the largest grocery retail group on the East Coast, we understand our vital role in enabling healthier people and a healthier planet and have an ongoing commitment to driving sustainable change that leads to a thriving food system, nourishes local communities, and creates a better world.

Why Work With Us

We love fresh perspectives, not just fresh produce. We believe that an inclusive workplace fosters creativity, accelerates innovation, and helps us create an even better product. At Ahold Delhaize USA, you’ll find coworkers who are caring and committed, and who focus on dreaming big and getting things done.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Ahold Delhaize USA Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: 3 days a week
HQChicago, IL
Carlisle, PA
Landover, MD
Mauldin, SC
Quincy, MA
Salisbury, NC
Scarborough, ME
Learn more

Similar Jobs

Ahold Delhaize USA Logo Ahold Delhaize USA

Digital Analytics Co-op

AdTech • eCommerce • Food • Marketing Tech • Retail
In-Office
Salisbury, NC, USA
10000 Employees
36K-37K Hourly

Ahold Delhaize USA Logo Ahold Delhaize USA

Software Engineer

AdTech • eCommerce • Food • Marketing Tech • Retail
In-Office
Salisbury, NC, USA
10000 Employees
36K-37K Hourly

Ahold Delhaize USA Logo Ahold Delhaize USA

Software Development Co-op

AdTech • eCommerce • Food • Marketing Tech • Retail
In-Office
Salisbury, NC, USA
10000 Employees
36K-37K Hourly

Ahold Delhaize USA Logo Ahold Delhaize USA

Project Management Analyst Co-op

AdTech • eCommerce • Food • Marketing Tech • Retail
In-Office
Salisbury, NC, USA
10000 Employees
31K-34K Hourly

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account