Principal Platform Engineer

Reposted 15 Days Ago
Be an Early Applicant
Seattle, WA, USA
In-Office
Senior level
Artificial Intelligence • Software • Generative AI
The Role
Own and improve the reliability, scalability, and operational health of Gradial's production platform. Lead Kubernetes, CI/CD, GitOps, infrastructure-as-code, observability, monitoring, alerting, incident response, and operational tooling. Drive performance, cost, security, and platform readiness in partnership with engineering teams.
Summary Generated by Built In

Gradial is the marketing operations system of work that helps marketers and creatives move from idea to execution faster. Our platform orchestrates across martech stacks, workflows, and people to automate marketing execution, cutting execution time, so marketers focus on the work only humans can do: shaping brands, understanding customers, and creating the work that moves people.

Backed by leading investors, we’re building software that adapts to the user, not the other way around. We move with urgency, operate with ownership, and solve hard problems from first principles. If you want to do ambitious work, take real responsibility, and help define the future of AI-native content operations, you’ll do your best work here.

The Role

As a Principal Platform Engineer at Gradial, you will shape the foundation our platform runs on as we scale. You will work closely with the CTO and engineering team to make our systems faster, more resilient, and easier to operate in a high-growth environment. This is a hands-on individual contributor leadership role for someone who wants real ownership, high leverage, and the opportunity to define how platform reliability looks at an AI-native company.

What You’ll Own

  • Own the reliability, scalability, and operational health of Gradial’s production platform.
  • Lead the evolution of Kubernetes, CI/CD, observability, and infrastructure as code across the stack.
  • Set the standard for how we design, ship, and operate reliable systems.
  • Build the tooling and automation that help engineers move faster with more confidence.
  • Drive improvements in monitoring, alerting, incident response, and service readiness.
  • Partner with engineering to identify scaling risks early and solve them before they slow us down.
  • Influence the long-term direction of our platform across reliability, security, performance, and cost.

What We’re Looking For

  • 5+ years of experience in platform engineering, infrastructure, SRE, DevOps, or related roles with direct ownership of production systems.
  • Proven success designing and operating production-grade infrastructure in fast-moving, high-growth environments.
  • Deep expertise in Kubernetes, cloud-native architecture, and container orchestration.
  • Strong experience with infrastructure as code, GitOps, CI/CD workflows, and modern deployment practices.
  • Strong command of observability and reliability fundamentals across metrics, logging, tracing, alerting, and incident response.
  • A track record of leading through influence, making sound technical decisions, and raising the bar across engineering teams.

Nice to Have

  • Familiarity with AI or ML infrastructure, including GPU provisioning, model deployment, or compute-intensive workloads.
  • Experience supporting cloud or multi-cloud environments with a focus on resilience and scale.
  • Comfort with TypeScript or Python for internal tooling and operational automation.
What we offer
  • Competitive salary and meaningful equity
  • Comprehensive health, dental and vision coverage
  • Fast-paced environment with flexibility and ownership
  • Real impact, zero bureaucracy
  • A front-row seat to building category-defining AI infrastructure
You'll thrive here if you...
  • Learn quickly, actively seek out new challenges, and regularly reconsider “how it’s always been done.”
  • Have an innate drive for being 1% better than the day before, building towards greatness.
  • Embrace AI as a core tool for problem-solving, innovation, and scale.
  • Show customer-obsession (internal or external), high ownership/accountability, and bias for action.
  • Communicate clearly, directly, with curiosity, and assuming good intentions.
  • Thrive in fast-paced, hyper-growth environments where building better > maintaining status quo.
AI Literacy & Interviewing Tools

As an AI-first company, we prioritize AI literacy as a core competency in our hiring decisions. We’re excited by candidates who thoughtfully apply AI tools in their work, but during interviews we’re focused on you. This is your opportunity to show how you think, communicate, and solve problems. Over-reliance on AI-generated responses during the interview process (especially when it obscures your own voice) will result in disqualification. We want to understand your unique perspective and how you approach challenges, both with and without AI.

Gradial is dedicated to creating an environment where diverse perspectives are valued and all team members can grow. We offer competitive compensation, equity, flexible work hours, comprehensive benefits, and a collaborative culture focused on learning and impact.

Skills Required

  • 5+ years experience in platform engineering, infrastructure, SRE, DevOps, or related roles with direct ownership of production systems
  • Proven success designing and operating production-grade infrastructure in high-growth environments
  • Deep expertise in Kubernetes, cloud-native architecture, and container orchestration
  • Experience with infrastructure as code, GitOps, CI/CD workflows, and modern deployment practices
  • Strong command of observability and reliability fundamentals across metrics, logging, tracing, alerting, and incident response
  • Track record of leading through influence and making sound technical decisions across engineering teams
  • Familiarity with AI or ML infrastructure, including GPU provisioning, model deployment, or compute-intensive workloads
  • Experience supporting cloud or multi-cloud environments with focus on resilience and scale
  • Comfort with TypeScript or Python for internal tooling and operational automation
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Seattle, WA
6 Employees
Year Founded: 2023

What We Do

Gradial is the enterprise generative AI platform for transforming content management and customer journey intelligence. Gradial leverages frontier generative image and text models to enable personalized digital experiences at scale. We empower creative teams to bring their vision to life using words and truly differentiate their brand's voice in a sea of content.

Similar Jobs

CrowdStrike Logo CrowdStrike

Principal Software Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
4 Locations
11000 Employees
195K-290K Annually
Hybrid
4 Locations
289097 Employees

Snap Inc. Logo Snap Inc.

Principal Software Engineer

Artificial Intelligence • Cloud • Machine Learning • Mobile • Software • Virtual Reality • App development
Remote or Hybrid
6 Locations
5000 Employees
235K-414K Annually

Microsoft Logo Microsoft

Principal Software Engineer

Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
In-Office
Redmond, WA, USA
206870 Employees
143K-304K Annually

Similar Companies Hiring

Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
300 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account