Head of Platform

Posted 17 Days Ago
Be an Early Applicant
2 Locations
Hybrid
Senior level
Artificial Intelligence • Machine Learning • Software
The Role
Lead and build the Platform function (SRE, DevEx, Infra, Security enablement) to provide self-service foundations: developer workflows, IaC/GitOps, observability, incident practices, security-by-default, cost controls, and reliability. Set roadmap, hire and grow the team, partner with product and engineering, deliver measurable developer productivity and operational excellence, and be hands-on early to unblock and set standards.
Summary Generated by Built In

Note: Partly is headquartered in Austin, TX with offices in London, UK, Christchurch, NZ and Auckland, NZ. Wherever you're based, we'll connect you with your nearest office for onboarding, and fly you to join the full team for our quarterly "Season Openers" (we cover travel and accommodation). If you're relocating to join us, we can also assist with relocation costs.

🚀 Our story

Partly is connecting the world's parts, and we're doing that by building the AI infrastructure layer for the global repair industry, starting with the $2tn automotive market. Our frontier model, Interpreter, is the world's first AI purpose-built to understand vehicle damage and the parts needed to fix it. Thousands of businesses across the global repair supply chain already rely on it.

Founded by ex-Rocket Lab engineers, we've tripled in size in the last 18 months and have recently raised a $50m Series B led by DST Global (Anthropic, Airbnb, Meta, TikTok, Spotify) and including Blackbird Ventures (Canva, CultureAmp etc.), WNDR, Activant Capital, Icehouse Ventures, Square Peg, Airtree, and Ecliptic Venture Capital. We're headquartered in Austin, with offices in New Zealand and London.

We're continuing to build a world-class team ensuring Partly is a place where people can do the best work of their lives. We're proud of the culture we've built, and our values are lived throughout every experience.

Want to learn more about the problems we're solving and the culture we're building at Partly? Hear directly from our team here: https://shorturl.at/iAFUX

🖍️ This role

Head of Platform is responsible for building and leading the team that provides Partly’s foundation for shipping and operating software: the internal platform, infrastructure foundations, reliability practices, and security-by-default capabilities that let product teams move fast without breaking things. This now extends in two directions: the specialised infrastructure behind Partly's foundational ML/AI capabilities (GPU compute, model training and serving, inference cost), and a platform that natively supports agentic development so regionally-embedded product engineers (including forward-deployed engineers) can move exceptionally fast. You’ll treat platform as an internal product - setting strategy, driving adoption, and partnering closely with engineering and the business to improve delivery speed, uptime, and cost efficiency as we scale.

💻 What will you do
  • Platform Strategy & Team Leadership: Build and lead the Platform function (SRE, DevEx, Infrastructure, Security enablement as appropriate). Set a clear roadmap, establish ways of working, hire and grow the team, and manage prioritisation/trade-offs.

  • Developer Experience : Create the default path for engineers to build, test, deploy, and operate services (service templates, CI/CD, environment provisioning, secrets/config, deployment patterns, feature flags). Focus on adoption and measurable improvements in developer productivity.

  • Reliability & Operational Excellence: Own or drive (depending on org boundaries) our reliability foundations: observability (metrics/logs/traces), alerting standards, incident response, postmortems, SLO/error budget practices, rollout/rollback patterns, backups/DR, and reducing on-call toil.

  • Infrastructure Foundations: Ensure our cloud and Kubernetes foundations are scalable, secure, and maintainable. Use Infrastructure-as-Code and automation (Terraform for GCP, GitOps with ArgoCD, Python/Bash tooling, etc.) to run repeatable, auditable infrastructure.

  • ML & Foundation-Model Infrastructure: Build the platform beneath Partly's foundational models — accelerator (GPU/TPU) provisioning, scheduling and utilisation, training/fine-tuning orchestration, and scalable, low-latency model serving. Own the ML deployment lifecycle (model registry, experiment tracking, evaluation/observability) and treat inference economics as a first-class cost driver, since inference — not training — typically dominates AI infrastructure spend at scale.

  • Agentic & Forward-Deployed Enablement: Make the platform natively support agentic development. Design golden paths that are self-service and machine-consumable (discoverable, executable, safe) so AI coding agents and engineers alike can go from generated code to production without the platform team as a bottleneck — with isolated execution environments, guardrails, and observability built in. Ensure regional and forward-deployed engineers can move at maximum speed on a robust but extensible foundation.

  • Security Enablement: Partner with security/compliance to make secure-by-default the easiest path (IAM patterns, secrets management, vulnerability management, policy-as-code where appropriate, audit evidence automation).

  • Cost & Performance Ownership: Establish FinOps-style visibility and guardrails, track cost drivers, and deliver optimisations that improve unit economics without sacrificing reliability or developer velocity.

  • Cross-Functional Collaboration: Work closely with Product/Engineering leadership and stream-aligned teams to understand bottlenecks, influence architectural direction, and ensure platform work translates into real outcomes for customers and the business.

  • Hands-on Delivery (especially early): You’ll be technical enough to dive in, review designs, unblock incidents, prototype solutions, and set technical standards, while building a team that doesn’t rely on you as the single point of execution.

Want to learn more about the problems we're solving and the culture we're building at Partly? Hear directly from our team here: https://shorturl.at/iAFUX

🥷 Your skills
  • Platform Leadership: Proven experience leading a platform / infrastructure / SRE function, including roadmap ownership, stakeholder management, and building teams in a fast-moving environment. You know how to balance reliability, developer productivity, security, and cost.

  • SRE & Operations Expertise: Strong grounding in SRE practices (SLOs/error budgets, incident management, observability, capacity planning, resilience engineering) and a track record of improving uptime and reducing operational toil.

  • Cloud & Kubernetes Depth: Deep familiarity with running production workloads on a major cloud (GCP preferred) and Kubernetes. You can design scalable infrastructure, debug systems issues, and make pragmatic build vs buy decisions.

  • Infrastructure-as-Code & Automation: Hands-on expertise with IaC and GitOps workflows (Terraform, ArgoCD or equivalent) and the software engineering ability to build robust tooling (not just scripts).

  • Developer Experience Mindset: You treat platform as a product: you can define “golden paths,” simplify workflows, drive adoption through empathy and excellent docs, and measure impact (e.g., lead time, deploy frequency, MTTR, change failure rate).

  • Security-by-Default: Practical experience embedding security into platforms and SDLC (IAM, secrets, vulnerability management, supply chain hygiene). Bonus if you’ve helped achieve/maintain compliance (SOC2/ISO).

  • Strong Engineering Fundamentals: Solid CS and system engineering fundamentals (concurrency, networking, Linux internals, performance profiling, distributed systems, reliability patterns).

  • ML/AI Infrastructure Depth: Familiarity with the AI infrastructure stack — GPU/accelerator compute, training/fine-tuning orchestration, inference and model serving, MLOps (registries, experiment tracking, evals), and the cost/performance trade-offs of running models in production. You know where the real bottlenecks and spend live.

  • Agentic-Native Platform Thinking: You design platforms for a world where AI agents are first-class users, not just humans — machine-consumable golden paths, safe autonomous execution, and self-service that scales to far higher workload volume. Bonus if you’ve enabled forward-deployed / embedded engineers to customise and ship rapidly without sacrificing robustness.

  • Communication & Influence: Excellent written/verbal communication. You can align senior stakeholders, explain trade-offs to non-specialists, and coach engineers across the org.

  • Ownership & Bias for Action: You create clarity in ambiguity, deliver outcomes, and don’t wait to be told what to do. You’re comfortable being accountable for foundational systems.

  • Bonus Points:

    • Experience scaling platform practices through rapid growth (multiple teams/services).

    • Familiarity with our stack (GCP, ArgoCD, GitLab CI, Kafka, Postgres).

    • Experience building internal developer platforms, service frameworks, or multi-tenant platform capabilities.

Please note: if you don't have all the skills/experience listed above but believe you could be outstanding in this role, please still consider applying. Many folks, especially those from underrepresented or marginalised groups, often count themselves out. Please allow us to learn more about you and why you're exceptional!

🪅 Benefits
  • High trust, low process and no bureaucracy. We hire exceptional people whose judgment we trust. This means we proactively remove any process or rules that slow us down (for example, our expense policy is simply the “red face test”).

  • Competitive base salary + equity. We offer competitive salaries and generous equity options for all full-time employees, ensuring everyone shares in the financial upside when we win.

  • Flexible working hours. Choose when to work based on what time you’re most effective (no mandatory or set hours). We combine flexibility with an office-first approach (in cities where we have critical mass, i.e. London, Christchurch, Auckland).

  • Focus Days. Two days per week, with zero meetings, dedicated solely to uninterrupted deep work

  • Take time when you need it. We don’t ask questions or care if people have a negative leave balance. We work extremely hard and trust our team to take the time they need to recharge.

  • Offices in Christchurch CBD and on Auckland’s Drake Street. We invest heavily in our offices (standing desks, healthy snacks, quality coffee, drinks on tap) to ensure they’re places people are excited by, where they build relationships and get their best work done.

  • Learn from the best. Whether it’s during a ‘Lunch n Learn’ or hearing from a unicorn CEO at a Fireside chat, you’ll have the opportunity to constantly learn from the world’s best.

  • Quarterly season openers & annual global offsite. Connect regularly at the nearest centralised location for a week of collaboration, big-picture planning and team events.

  • Team connection. Monthly team lunches, celebrating our wins, happy hours and more!

  • Parental leave and flexible return to work. Do what works for you. Primary carers can return with 4-day weeks (on 100% pay for the first 12 weeks). Secondary carers get 10 days full pay.

  • Payroll Giving: We encourage generous giving and donate to the high-impact charities you support

🛬 Relocation
  • If you are relocating from overseas or domestically to Partly HQ, we offer a generous relocation allowance to support your move

Skills Required

  • Proven experience leading a platform/infrastructure/SRE function, owning roadmap and building teams
  • Strong grounding in SRE practices (SLOs, incident management, observability, resilience engineering)
  • Experience running production workloads on cloud and Kubernetes (GCP preferred)
  • Hands-on expertise with Infrastructure-as-Code and GitOps workflows (Terraform, ArgoCD)
  • Software engineering ability to build robust tooling (Python/Bash familiarity)
  • Developer experience mindset: define golden paths, drive adoption, measure lead time/MTTR/deploy frequency
  • Practical experience embedding security into platforms and SDLC (IAM, secrets, vulnerability management)
  • Strong engineering fundamentals (concurrency, networking, Linux internals, distributed systems)
  • Excellent communication and stakeholder management; ability to align senior leadership
  • Experience scaling platform practices across rapid-growth organisations
  • Familiarity with stack: GitLab CI, Kafka, Postgres
  • Experience with compliance programs (SOC2/ISO)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: London
81 Employees

What We Do

Partly’s mission is to connect the world's parts. We leverage advancements in machine learning and AI to transform the predominantly offline $1.9 trillion auto parts industry. Partly Infrastructure is chosen by leading enterprises to build their auto parts procurement platform, enables repairers and suppliers to transact in real time, providing supply chain visibility across OEM, aftermarket, and recycled suppliers in one convenient place. Partly is backed by industry-leading investors including Octopus Ventures, Blackbird, Squarepeg, I2BF, Ten13, Hillfarrance, Shasta Ventures, Icehouse Ventures, Peter Beck (Rocket Lab), Randy Reddig (Square), Dylan Field (Figma), and Akshay Kothari (Notion). We’re continuing to build a world-class team and ensuring Partly is a place where people can do the best work of their lives. We’re proud of the culture we’ve built at Partly, and our values are lived through every experience. Partly is ISO27001 Certified

Similar Jobs

Halter Logo Halter

Senior Software Engineer

Greentech • Hardware • Internet of Things • Machine Learning • Software • Business Intelligence • Agriculture
In-Office
Auckland, NZL
350 Employees

Halter Logo Halter

Lead React Native Engineer

Greentech • Hardware • Internet of Things • Machine Learning • Software • Business Intelligence • Agriculture
In-Office
Auckland, NZL
350 Employees

Mastercard Logo Mastercard

Senior Analyst, Account Management - Pacific Islands

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Auckland, NZL
38800 Employees

Xero Logo Xero

Analytics Manager

Cloud • Fintech • Information Technology • Machine Learning • Software
Hybrid
Auckland, NZL
4500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account