Technical Program Manager

Posted 2 Days Ago
Be an Early Applicant
Palo Alto, CA, USA
In-Office
160K-300K Annually
Mid level
Artificial Intelligence • Information Technology • Machine Learning • Software
The Role
Drive complex, cross-functional AI infrastructure programs across inference, training, hardware enablement, and engineering teams. Own milestones, dependencies, risks, release management, technical coordination, incident response, operating cadence, metrics, and stakeholder communication. Partner with product, engineering, research, hardware vendors, AI labs, and enterprise customers to deliver scalable infrastructure programs.
Summary Generated by Built In
About the Role
As a Technical Program Manager at RadixArk, you'll drive the execution of complex, cross-functional programs across our inference and training infrastructure. You'll partner closely with Product Management, Research, and Engineering to turn ambitious technical roadmaps into shipped reality, coordinating across kernel teams, distributed systems engineers, and external partners to deliver infrastructure that serves billions of tokens daily and coordinates 10,000+ GPU training runs.
This role is for someone who thrives at the intersection of deep technical understanding and rigorous program execution. You'll own the "how" and "when" of our most critical initiatives.
Key ResponsibilitiesProgram Execution & Delivery
  • Drive end-to-end execution of large-scale, cross-functional programs spanning inference engines (e.g., SGLang), training frameworks (e.g., Miles), and hardware integration efforts.
  • Define program structure, including milestones, dependencies, critical paths, risks, and success criteria. Maintain a clear source of truth for status across all stakeholders.
  • Run design reviews, sprint planning, release readiness reviews, and post-mortems. Ensure decisions are documented and follow-ups are closed out.
  • Identify and unblock cross-team dependencies across kernel, runtime, scheduler, networking, and model teams before they become release blockers.
  • Drive release management for major versions, including changelog ownership, compatibility validation, partner rollout sequencing, and rollback planning.
Technical Coordination
  • Partner with Product Management to translate roadmap priorities into executable program plans, with clear scope, staffing, and timelines.
  • Work shoulder-to-shoulder with engineering leads on technical trade-off decisions; understand the architecture deeply enough to ask the right questions and surface hidden risks.
  • Coordinate hardware enablement programs with partners like Nvidia, Google, and AWS, including new accelerator bring-up, kernel co-development, and benchmark validation.
  • Manage integration programs with frontier AI labs and early adopters, ensuring technical requirements, SLAs, and feedback loops are well-defined.
Operational Excellence
  • Build and improve the engineering operating cadence, including standups, planning rituals, OKR tracking, dashboards, and reporting to leadership.
  • Establish metrics and instrumentation for program health such as velocity, defect rates, benchmark regressions, and customer-reported issues, and drive accountability against them.
  • Lead incident response coordination for production issues affecting partners; own root-cause review and corrective-action tracking.
  • Improve developer productivity by identifying and removing systemic friction in our build, test, and release pipelines.
Stakeholder Communication
  • Serve as the connective tissue between engineering, product, GTM, and external partners, ensuring everyone has the right information at the right altitude.
  • Produce clear, concise written updates for leadership and partners. Translate engineering progress into business-relevant signals.
  • Represent program status honestly, including risks and slips, with concrete mitigation plans.
QualificationsMinimum Requirements
  • Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
  • 4+ years of direct experience in Technical Program Management, Engineering Management, or a senior engineering role with significant program ownership, in a software or infrastructure company.
  • Strong technical fluency in systems software, distributed systems, or AI/ML infrastructure; able to read code, follow architecture discussions, and challenge technical assumptions productively.
  • Demonstrated track record shipping complex, multi-team programs on time, including managing dependencies, risks, and scope changes.
  • Excellent written and verbal communication skills; able to drive alignment across engineers, executives, and external partners.
Preferred (Bonus) Qualifications
  • Direct experience shipping AI/ML infrastructure such as inference engines, training frameworks, GPU kernels, distributed schedulers, or model serving platforms.
  • Hands-on coding background (Python, C++, CUDA) and comfort working in engineering codebases, including reading PRs, running benchmarks, and reproducing issues.
  • Experience coordinating with hardware vendors (Nvidia, AMD, Google TPU, AWS Trainium/Inferentia) on enablement or co-engineering programs.
  • Experience driving open-source release programs or working in OSS communities, including issue triage, RFC processes, and contributor coordination.
  • Familiarity with release engineering, CI/CD systems, and observability tooling for large-scale distributed systems.
  • Experience supporting B2B or developer-facing products with enterprise SLAs.
About RadixArk

RadixArk is an infrastructure-first company built by engineers who've shipped production AI systems, created SGLang (30K+ GitHub stars, the fastest open LLM serving engine), and developed Miles (our large-scale RL framework). Founded by AI infrastructure veterans from xAI and NVIDIA, we're on a mission to democratize frontier-level AI infrastructure by building world-class open systems for inference and training. Our team has optimized kernels serving billions of tokens daily, designed distributed training systems coordinating 10,000+ GPUs, and contributed to infrastructure that powers leading AI companies and research labs.

Compensation

Depending on background, skills, and experience, the expected annual salary range for this position is $160,000 - $300,000 USD + equity.

Equal Opportunity

RadixArk is an Equal Opportunity Employer and is proud to offer equal employment opportunity to everyone regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity, veteran status, and more.


Skills Required

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field
  • 4+ years of direct experience in Technical Program Management, Engineering Management, or senior engineering with significant program ownership
  • Experience in a software or infrastructure company
  • Strong technical fluency in systems software, distributed systems, or AI/ML infrastructure
  • Ability to read code, follow architecture discussions, and challenge technical assumptions productively
  • Track record of shipping complex, multi-team programs on time while managing dependencies, risks, and scope changes
  • Excellent written and verbal communication skills across engineers, executives, and external partners
  • Experience shipping AI/ML infrastructure such as inference engines, training frameworks, GPU kernels, distributed schedulers, or model serving platforms
  • Hands-on coding background in Python, C++, or CUDA
  • Experience coordinating with hardware vendors on enablement or co-engineering programs
  • Experience driving open-source release programs or working in OSS communities
  • Familiarity with release engineering, CI/CD systems, and observability tooling for large-scale distributed systems
  • Experience supporting B2B or developer-facing products with enterprise SLAs
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
59 Employees
Year Founded: 2025

What We Do

RadixArk is an infrastructure-first company building large-scale inference and training systems for the AI community. It invests in SGLang, an open engine for serving modern models, and develops Miles, an open-source framework for large-scale reinforcement-learning post-training. The company also provides managed infrastructure and tooling for developers, startups, enterprises, and research labs, with a mission to make frontier-level AI infrastructure open and accessible to everyone.

Similar Jobs

Hybrid
2 Locations
289097 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Technical Program Manager

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
4 Locations
85422 Employees
106K-243K Annually

Capital One Logo Capital One

Technical Program Manager

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
210K-287K Annually

Capital One Logo Capital One

Technical Program Manager

Fintech • Machine Learning • Payments • Software • Financial Services
Hybrid
4 Locations
55000 Employees
210K-287K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account