Staff Software Engineer - AI Research Infrastructure

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office
199K-270K Annually
Senior level
Big Data • Machine Learning • Software • Analytics • Big Data Analytics
The Role
Develop and manage the research infrastructure for AI at Databricks, focusing on large-scale training, orchestration, and improving developer productivity.
Summary Generated by Built In
Staff Software Engineer - AI Research Infrastructure

P-1215

At Databricks, we are obsessed with enabling data teams to solve the world’s toughest problems, from security threat detection to cancer drug development. We do this by building and running the world’s best data and AI platform so our customers can focus on the high-value challenges that are central to their own missions.

The Databricks AI Research organization enables companies to develop AI models and agents using their own data, with technologies ranging from post-training open source LLMs to developing advanced multi-agent architectures. Databricks AI does so by producing novel science and putting it into production. Databricks AI is committed to the belief that a company’s AI models and agents are just as valuable as any other core IP, and that high-quality AI should be available to all.

Job Description

As a Staff Software Engineer, AI Research Infrastructure, you will be developing and running the research stack that powers Databricks AI Research. You will design and build services that schedule, orchestrate, and observe large‑scale training and inference experiment workloads across thousands of GPUs, improve our dev tooling and ensure that researchers can iterate quickly without sacrificing reliability, efficiency, or security.

You’ll partner closely with research scientists, ML engineers, and platform teams to turn experimental workloads into robust, repeatable pipelines, and to push the limits of what our infrastructure can support.

The Impact you will have

As a Staff Software Engineer on the AI Research Infra Team at Databricks, you will: 

  • Design and implement infrastructure that supports large‑scale experiments, data processing, and model training (e.g., HPC clusters, GPU fleets, or cloud‑based systems)
  • Enable researchers to go from idea to large‑scale experiment in minutes, not days, by building powerful abstractions for job submission, scheduling, and monitoring.
  • Create tooling that improves research developer productivity, such as experiment management systems, CI/testing infrastructure for research code, and workflows that reduce iteration time.
  • Influence the long‑term roadmap for research computation, shaping how Databricks AI Research train, evaluate, and ship models to customers.
  • Serve as a technical mentor and force multiplier for other engineers working on compute, infra, and AI systems.

What We Look for

  • BS/MS or PhD in Computer Science or related field
  • 5+ years of software engineering experience, including substantial time working on large‑scale distributed systems or infrastructure.
  • Have deep experience with building and operating distributed systems, data pipelines, or large‑scale backend services, ideally involving GPUs, clusters, or major cloud providers.
  • Are proficient in one or more systems programming languages (e.g., C++, Rust, Go, Java, Scala) and can design, implement, and debug complex services.
  • Have built or significantly contributed to cluster schedulers, resource managers, or large‑scale job orchestration systems (e.g., Kubernetes, Slurm, Ray, custom internal systems).
  • Understand modern ML training and inference workflows (e.g., distributed training, model parallelism, fine‑tuning, evaluation), even if you’re not primarily a research scientist.
  • Can move fast and be pragmatic in getting things done, while caring about operational excellence. Have driven complex systems from prototype to stable, well‑owned services.
  • Communicate clearly with both researchers and engineers, and enjoy translating between research needs and infra realities.

Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.  Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.


Local Pay Range
$199,000$270,000 USD

About Databricks

Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.
Benefits
At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.

Our Commitment to Diversity and Inclusion

At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.

Compliance

If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.

Skills Required

  • BS/MS or PhD in Computer Science or related field
  • 5+ years of software engineering experience
  • Experience with large-scale distributed systems
  • Proficiency in systems programming languages such as C++, Rust, Go, Java, Scala
  • Experience with cluster schedulers or job orchestration systems
  • Understanding of modern ML training and inference workflows
  • Ability to communicate between research needs and infrastructure realities

Databricks Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Databricks and has not been reviewed or approved by Databricks.

  • Equity Value & Accessibility Equity grants and RSUs are a major part of total compensation and are highlighted for meaningful upside potential. Stock-based awards and refreshers contribute to strong overall pay positioning across senior technical and go-to-market roles.
  • Healthcare Strength Medical, dental, and vision coverage are complemented by mental-health resources, an EAP, and wellness reimbursements. Health benefits are consistently framed as comprehensive and competitive.
  • Parental & Family Support Paid parental leave for all parents, fertility support, and backup care options provide tangible assistance for family needs. Hybrid work norms and team-day structure further ease coordination for caregivers.

Databricks Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
New York, NY
2,200 Employees
Year Founded: 2013

What We Do

As the leader in Unified Data Analytics, Databricks helps organizations make all their data ready for analytics, empower data science and data-driven decisions across the organization, and rapidly adopt machine learning to outpace the competition. By providing data teams with the ability to process massive amounts of data in the Cloud and power AI with that data, Databricks helps organizations innovate faster and tackle challenges like treating chronic disease through faster drug discovery, improving energy efficiency, and protecting financial markets.

Similar Jobs

Databricks Logo Databricks

Staff Software Engineer

Big Data • Machine Learning • Software • Analytics • Big Data Analytics
In-Office
2 Locations
2200 Employees
190K-270K Annually

Brigit Logo Brigit

Vice President Of Finance

Fintech • Mobile • Social Impact • Financial Services
Hybrid
New York, NY, USA
132 Employees
200K-250K Annually

Alloy Logo Alloy

Senior Solutions Architect

Fintech • Information Technology • Software • Financial Services
Easy Apply
Hybrid
New York City, NY, USA
315 Employees
120K-167K Annually

Optimum Logo Optimum

Director Technology Program Delivery

AdTech • Digital Media • Internet of Things • Marketing Tech • Mobile • Retail • Software
Hybrid
New York, NY, USA
9000 Employees
141K-202K Annually

Similar Companies Hiring

Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York City, NY
100 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account