Software Engineer, ML Platform

Posted 4 Days Ago
2 Locations
In-Office
Entry level
Artificial Intelligence • Generative AI
The Role
Build and operate ML platform infrastructure supporting researchers and product engineers. Responsibilities include designing distributed systems, data pipelines, observability tools, scheduling and orchestration infrastructure, and GPU fleet systems. The role partners closely with ML research, owns reliability and performance, improves developer experience, and ships platform primitives in a high-ownership environment.
Summary Generated by Built In

Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.

About the role

As a Software Engineer on ML Platform at Cursor, you'll build the infrastructure that turns real product usage into better models — and keeps research moving fast on large GPU fleets. ML Platform is organized into four teams. Depending on your background, you may join any of them:

  • Telemetry — Own the collection and serving path that turns real product use into a record research can trust; without slowing the product, and under a small, explicit policy. Client-side or high-volume ingestion experience is a plus.

  • ML Data Platform — Build the shared environments and pipeline substrate researchers extend, so new experiments don’t fork their own stack.

  • Observability — Make it easy for researchers to start, watch, and debug their own runs.

  • ML DevX and Systems — Shorten the path from idea to a trusted run on the research fleet.

We're looking for strong distributed-systems and infrastructure engineers who want to sit next to research and ship platform primitives that move the product.

We're in-person with cozy offices in North Beach, San Francisco, Palo Alto, and Manhattan, New York, complete with well-stocked libraries.

What you’ll do
  • Design, build, and operate core platform systems used daily by ML researchers and product engineers

  • Partner closely with research to turn recurring pain into durable infrastructure

  • Own reliability, performance, and developer experience for the systems in your lane

  • Ship iteratively in a flat, high-ownership environment. Measure impact, then raise the bar

You may be a fit if
  • You have a strong background in systems / infrastructure software engineering and enjoy building platforms other engineers depend on

  • You've owned production distributed systems at meaningful scale (ingestion, data pipelines, scheduling/orchestration, or similar)

  • You're comfortable across Linux, cloud and/or bare metal, and modern orchestration (Kubernetes, Ray, or equivalent)

  • You like working closely with ML researchers and product engineers

  • You thrive where ownership is high and the feedback loop is short

Especially strong backgrounds by team
  • Telemetry: event ingestion, product analytics pipelines, OpenTelemetry / tracing, reliable data APIs

  • Product Data Platform: data frameworks, Spark / Flink / Ray, ML dataset and training-data infrastructure

  • Observability: experiment / run monitoring, debug and eval tooling, agent-friendly observability UX

  • ML DevX and Systems: GPU / cluster scheduling, job queues, node health, research compute developer experience

Applying

If there appears to be a fit, we'll reach out to schedule 2-3 short technicals. After, we'll schedule an onsite in our office, where you'll work on a small project, discuss ideas, and meet the team.

Skills Required

  • Strong background in systems or infrastructure software engineering
  • Experience building platforms that other engineers depend on
  • Experience owning production distributed systems at meaningful scale, such as ingestion, data pipelines, scheduling, or orchestration systems
  • Comfort working across Linux, cloud or bare-metal infrastructure, and modern orchestration systems such as Kubernetes or Ray
  • Ability to work closely with ML researchers and product engineers
  • Event ingestion, product analytics pipelines, OpenTelemetry, tracing, or reliable data APIs
  • Data frameworks, Spark, Flink, Ray, or ML dataset and training-data infrastructure
  • Experiment monitoring, run monitoring, debugging, evaluation tooling, or observability user experience
  • GPU or cluster scheduling, job queues, node health, or research-compute developer experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
San Francisco, CA
300 Employees
Year Founded: 2022

What We Do

We'd like to automate coding. To advance that mission, we're building Cursor. Our work includes training the world’s most widely used coding models, creating infrastructure that supports billions of requests per day, and building better ways for humans and AIs to work together.

Similar Jobs

Hybrid
4 Locations
289097 Employees

Gusto Logo Gusto

Software Engineer

Fintech • HR Tech
Easy Apply
Hybrid
3 Locations
4405 Employees
160K-240K Annually
Hybrid
4 Locations
289097 Employees
Hybrid
New York, NY, USA
61 Employees
200K-250K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
LTX Thumbnail
Robotics • Conversational AI • Generative AI
Jerusalem, Israel
200 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account