AI Infrastructure Engineer

Posted One Month Ago
3 Locations
In-Office
Mid level
Artificial Intelligence • Information Technology • Software
The Role
Build and operate scalable, secure cloud infrastructure for real-time AI services. Improve reliability with SLOs, observability, capacity planning, and failure testing. Implement IaC, CI/CD, deployment safety, secrets and vulnerability management. Partner with product and AI teams, diagnose distributed failures, reduce cost, and automate platform patterns to enable faster, safer releases.
Summary Generated by Built In

Palona’s AI agents operate continuously in production, handle real-time guest interactions, integrate with restaurant systems, and face sharp traffic peaks. Infrastructure is therefore part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the guest and operator experience.

We are looking for an Infrastructure Engineer who combines cloud and reliability depth with strong software engineering judgment. You will build and operate the platform beneath Palona’s AI products, improve how engineers ship, and turn production signals into durable system improvements. This is not a ticket-driven IT or operations role. You will write production code, design systems, automate repetitive work, and own outcomes across the full service lifecycle.

Our current environment includes Python services, Docker, AWS and selected Azure services, ECS and Lambda workloads, API Gateway, load balancers, relational data systems, OpenTofu/Terraform, Datadog, and CI/CD automation. We value the ability to learn and make sound tradeoffs more than exact tool-for-tool matching.

What you will own:
  • Design, build, and evolve secure, scalable cloud infrastructure for real-time AI services and customer-facing applications.
  • Improve service reliability through clear SLOs, actionable observability, capacity planning, failure testing, and pragmatic incident prevention.
  • Build deployment and release systems that make production changes fast, repeatable, auditable, and safe.
  • Own infrastructure as code, environment consistency, and reusable platform patterns across development, staging, and production.
  • Partner with product and AI engineers on architecture, performance, data flows, and operational readiness for new capabilities.
  • Diagnose complex distributed-system failures across application, network, database, model-provider, and third-party integration boundaries.
  • Reduce infrastructure and model-serving cost without compromising customer experience or engineering velocity.
  • Strengthen secrets management, access controls, backup and recovery, vulnerability management, and other practical security foundations.
  • Build internal tooling and paved paths that let engineers ship and operate services with less manual work.
  • Participate in incident response and turn incidents into better systems, automation, documentation, and engineering judgment.

Requirements
  • 3+ years industrial experience in relevant technical domain.
  • Strong software engineering fundamentals and experience building or operating production distributed systems.
  • Hands-on experience with a major cloud platform; AWS experience is especially relevant.
  • Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging.
  • Ability to write reliable automation and services in Python or another modern programming language.
  • Sound judgment around availability, latency, scalability, security, and cost tradeoffs.
  • A track record of taking ambiguous operational problems from diagnosis through durable resolution.
  • Clear communication during architecture reviews, launches, and incidents.
  • AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems.

Benefits
  • Competitive Salary and Stock Option Plan.
  • Medical, dental, vision, retirement, leave, and disability benefits as applicable.
  • Family Leave
  • Short Term & Long Term Disability
  • Paid time off and company holidays.
  • Learning and development support.

Skills Required

  • 3+ years industrial experience in relevant technical domain.
  • Strong software engineering fundamentals and experience building or operating production distributed systems.
  • Hands-on experience with a major cloud platform; AWS experience is especially relevant.
  • Experience with containers, infrastructure as code, CI/CD, monitoring, alerting, and production debugging.
  • Ability to write reliable automation and services in Python or another modern programming language.
  • Sound judgment around availability, latency, scalability, security, and cost tradeoffs.
  • Track record of taking ambiguous operational problems from diagnosis through durable resolution.
  • Clear communication during architecture reviews, launches, and incidents.
  • AI-native working habits and curiosity about the operational behavior of LLM- and agent-powered systems.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Menlo Park, CA
26 Employees
Year Founded: 2024

What We Do

We are building a new, completely proprietary AI system with multi-agents, multimodal-to-action models, combined with unrivaled emotional intelligence language models to empower businesses to scale by delivering a high-touch customer experience, automation and efficiency. Our AI Agents will supercharge the core of your business.

Similar Jobs

Block Logo Block

Infrastructure Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
264K-395K Annually

Cash App Logo Cash App

Infrastructure Engineer

Blockchain • Fintech • Mobile • Payments • Software • Financial Services
Remote or Hybrid
8 Locations
3500 Employees
264K-395K Annually

AstraZeneca Logo AstraZeneca

Cloud Platform Engineer

Biotech • Pharmaceutical
Hybrid
Mississauga, ON, CAN
70000 Employees
135K-177K Annually
In-Office
Markham, ON, CAN
1770 Employees
58K-104K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account