AWS AgentCore Platform Engineer - 67417

Posted Yesterday
Be an Early Applicant
Reading, PA, USA
In-Office
Senior level
Fintech • Information Technology • Logistics
The Role
Designs and operates enterprise AI platforms using AWS Bedrock, AgentCore, MCP servers, and cloud-native technologies. Responsibilities include observability, distributed tracing, logging, cost governance, monitoring, incident management, IAM and ABAC security, Terraform infrastructure, CI/CD automation, and platform standardization. The role collaborates with architecture, engineering, security, and business stakeholders to deliver reliable, secure, scalable, and cost-efficient AI infrastructure.
Summary Generated by Built In

Function

Cloud & Data EngineeringOur Company

We’re Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world’s potential. We’re people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what’s now to what’s next. We make it happen through the power of acceleration.

Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don’t expect you to ‘fit’ every requirement – your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.

Job descriptionMeet Our Team

Join a forward-thinking Engineering and AI Platform team focused on building the next generation of enterprise AI solutions. Our team is pioneering agentic AI ecosystems powered by AWS Bedrock, AgentCore, MCP servers, and modern cloud-native technologies.

As a Senior AWS AgentCore Platform Engineer, you'll work alongside Cloud Architects, AI Engineers, Platform Engineers, and Security specialists to establish scalable, secure, and observable AI platforms. You'll play a critical role in defining the operational foundation that enables enterprise teams to deploy AI agents with confidence, governance, and efficiency.

This is an exciting opportunity to shape enterprise AI infrastructure, drive innovation in LLMOps, and influence platform standards across multiple business units.

What You'll Be Doing

AI Platform Observability & Reliability

  • Design and implement enterprise-grade observability solutions for AI agent ecosystems built on AWS Bedrock, AgentCore, and MCP servers.
  • Assess and optimize CloudWatch, X-Ray, Bedrock logging, and AgentCore tracing capabilities against agentic workflow requirements.
  • Conduct gap analyses and implement observability solutions using Dynatrace and other monitoring platforms.
  • Develop distributed tracing frameworks for AI workloads, including:
    • LLM decision paths
    • Tool invocations
    • Sub-agent interactions
    • MCP server communications
  • Build structured logging frameworks to support troubleshooting, governance, and performance optimization.
  • Design post-deployment validation pipelines for AI agents and MCP servers, including deployment health monitoring and registration verification.

Cost Governance & Optimization

  • Architect cost visibility and governance frameworks across AI workloads.
  • Extend cloud tagging strategies to include agent runtimes, vector databases, MCP services, and Bedrock token consumption.
  • Develop cost allocation models to provide spending transparency by team, department, and application.
  • Build dashboards and reporting solutions for AI platform cost tracking and forecasting.
  • Configure AWS Budgets, automated alerts, anomaly detection, and optimization recommendations.
  • Deliver automated cost reporting through Microsoft Teams and email channels.

Monitoring & Incident Management

  • Define enterprise monitoring standards and alerting frameworks across AI platform services.
  • Create and manage P1-P4 alerting strategies covering:
    • Deployment failures
    • Runtime exceptions
    • Tool invocation errors
    • MCP connectivity issues
  • Integrate monitoring and notification workflows with Microsoft Teams and email.
  • Develop operational runbooks and self-service documentation within Confluence.
  • Evaluate AWS-native and third-party monitoring solutions and recommend target-state architectures.

Security & Platform Governance

  • Assess IAM architectures and multi-team access models for enterprise-scale AI environments.
  • Design Attribute-Based Access Control (ABAC) frameworks to support secure multi-team isolation.
  • Evaluate Cedar policy engine capabilities within AgentCore for fine-grained authorization models.
  • Develop reusable Terraform modules to enforce governance, security, and compliance standards.
  • Identify scalability risks and implement secure platform design patterns for enterprise AI adoption.

Platform Engineering & Automation

  • Build and maintain Infrastructure-as-Code solutions using Terraform.
  • Design and enhance CI/CD pipelines supporting AI platform deployments.
  • Collaborate with engineering, security, architecture, and business stakeholders in Agile environments.
  • Drive platform standardization, automation, and operational excellence initiatives.
What You'll Bring to the Team

Required Qualifications

  • 8+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering.
  • Strong expertise in AWS cloud services including:
    • IAM
    • CloudWatch
    • AWS Lambda
    • AWS Bedrock
    • Cloud-native monitoring and governance services
  • Hands-on experience implementing observability and distributed tracing solutions using tools such as:
    • Dynatrace
    • Jaeger
    • Honeycomb
    • OpenTelemetry
  • Experience designing and managing Infrastructure-as-Code using Terraform.
  • Strong background building and maintaining CI/CD pipelines in enterprise environments.
  • Experience working in Agile teams utilizing Microsoft Teams, Confluence, and modern collaboration tools.

Preferred Qualifications

  • Experience supporting AI, Generative AI, or LLM-based platforms.
  • Familiarity with AgentCore, LangChain, LangFuse, LiteLLM, MCP servers, or similar AI orchestration frameworks.
  • Understanding of LLM lifecycle management, prompt execution flows, token consumption tracking, and AI workload optimization.
  • Knowledge of cloud cost management, FinOps practices, and governance frameworks.
  • Experience designing enterprise-scale security architectures using ABAC and policy-based authorization models.
  • Strong analytical and problem-solving skills with the ability to translate complex technical challenges into scalable platform solutions.

Success Factors

  • Passion for emerging AI technologies and cloud-native engineering.
  • Ability to balance reliability, security, performance, and cost optimization.
  • Strong communication skills with the ability to influence technical and business stakeholders.
  • Proven ability to lead platform modernization initiatives and establish engineering best practices.

Location: Reading, PA (Hybrid – 2-3 days onsite per week)

About us

We’re a global, team of innovators. Together, we harness engineering excellence and passion to co-create meaningful solutions to complex challenges. We turn organizations into data-driven leaders that can make a positive impact on their industries and society. If you believe that innovation can bring a better tomorrow closer to today, this is the place for you.

Fostering innovation through diverse perspectives

Hitachi is a global company operating across a wide range of industries and regions. One of the things that sets Hitachi apart is the diversity of our business and people, which drives our innovation and growth.

We are committed to building an inclusive culture based on mutual respect and merit-based systems. We believe that when people feel valued, heard, and safe to express themselves, they do their best work.

How we look after you

We help take care of your today and tomorrow with industry-leading benefits, support, and services that look after your holistic health and wellbeing. We’re also champions of life balance and offer flexible arrangements that work for you (role and location dependent). We’re always looking for new ways of working that bring out our best, which leads to unexpected ideas. So here, you’ll experience a sense of belonging, and discover autonomy, freedom, and ownership as you work alongside talented people you enjoy sharing knowledge with.

We’re proud to say we’re an equal opportunity employer and welcome all applicants for employment without attention to race, colour, religion, sex, sexual orientation, gender identity, national origin, veteran, age, disability status or any other protected characteristic. Should you need reasonable accommodations during the recruitment process, please let us know so that we can do our best to set you up for success.

Skills Required

  • 8+ years of experience in Platform Engineering, DevOps, Site Reliability Engineering, or Cloud Infrastructure Engineering
  • Strong expertise in AWS services including IAM, CloudWatch, Lambda, Bedrock, and cloud-native monitoring and governance services
  • Hands-on experience implementing observability and distributed tracing with Dynatrace, Jaeger, Honeycomb, OpenTelemetry, or similar tools
  • Experience designing and managing Infrastructure-as-Code using Terraform
  • Strong experience building and maintaining enterprise CI/CD pipelines
  • Experience working in Agile teams using Microsoft Teams, Confluence, and modern collaboration tools
  • Experience supporting AI, generative AI, or LLM-based platforms
  • Familiarity with AgentCore, LangChain, LangFuse, LiteLLM, MCP servers, or similar AI orchestration frameworks
  • Understanding of LLM lifecycle management, prompt execution flows, token consumption tracking, and AI workload optimization
  • Knowledge of cloud cost management, FinOps practices, and governance frameworks
  • Experience designing enterprise-scale security architectures using ABAC and policy-based authorization models
  • Strong analytical and problem-solving skills for translating complex technical challenges into scalable platform solutions
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tokyo
33,676 Employees

What We Do

Since its founding in 1910, Hitachi has responded to the expectations of society and its customers through technology and innovation. Our mission is to “Contribute to society through the development of superior, original technology and products.” Over the past 100+ years this commitment has led us to work towards creating a more sustainable society through our “Social Innovation Business”. We work to apply our expertise in information technology (IT), operational technology (OT), and a wide variety of products to advance social infrastructure systems and improve quality of life across the world.

Similar Jobs

Zscaler Logo Zscaler

Program Manager

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
USA
8697 Employees
1-1 Annually

Zscaler Logo Zscaler

Principal Product Manager

Cloud • Information Technology • Security • Software • Cybersecurity
Easy Apply
Remote or Hybrid
USA
8697 Employees
172K-245K Annually

The Aerospace Corporation Logo The Aerospace Corporation

Senior Program Manager

Aerospace • Artificial Intelligence • Cloud • Machine Learning • Software • Cybersecurity • Defense
Hybrid
Lansdale, PA, USA
4600 Employees

Gradient AI Logo Gradient AI

Vice President Of Sales

Artificial Intelligence • Insurance • Machine Learning • Software • Analytics
Easy Apply
Remote or Hybrid
USA
130 Employees
500K-500K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account