Software Engineer - GenAI inference

Reposted 21 Days Ago
San Francisco, CA, USA
In-Office
142K-205K Annually
Mid level
Big Data • Machine Learning • Software • Analytics • Big Data Analytics
The Role
As a Software Engineer for GenAI inference, you will design and optimize the inference engine for Databricks' Foundation Model API, focusing on performance and efficiency across various systems.
Summary Generated by Built In

P-1284

About This Role

As a software engineer for GenAI inference, you will help design, develop, and optimize the inference engine that powers Databricks’ Foundation Model API. You’ll work at the intersection of research and production, ensuring our large language model (LLM) serving systems are fast, scalable, and efficient. Your work will touch the full GenAI inference stack — from kernels and runtimes to orchestration and memory management.

What You Will Do
  • Contribute to the design and implementation of the inference engine, and collaborate on model-serving stack optimized for large-scale LLMs inference
  • Collaborate with researchers to bring new model architectures or features (sparsity, activation compression, mixture-of-experts) into the engine
  • Optimize for latency, throughput, memory efficiency, and hardware utilization across GPUs, and accelerators
  • Build and maintain instrumentation, profiling, and tracing tooling to uncover bottlenecks and guide optimizations
  • Develop and enhance scalable routing, batching, scheduling, memory management, and dynamic loading mechanisms for inference workloads
  • Support reliability, reproducibility, and fault tolerance in the inference pipelines, including A/B launches, rollback, and model versioning
  • Integrate with federated, distributed inference infrastructure – orchestrate across nodes, balance load, handle communication overhead
  • Collaborate cross-functionally: with platform engineers, cloud infrastructure, and security/compliance teams
  • Document and share learnings, contributing to internal best practices and open-source efforts when possible
What We Look For
  • BS/MS/PhD in Computer Science, or a related field
  • Strong software engineering background (3+ years or equivalent) in performance-critical systems
  • Solid understanding of ML inference internals: attention, MLPs, recurrent modules, quantization, sparse operations, etc.
  • Hands-on experience with CUDA, GPU programming, and key libraries (cuBLAS, cuDNN, NCCL, etc.)
  • Comfortable designing and operating distributed systems, including RPC frameworks, queuing, RPC batching, sharding, memory partitioning
  • Demonstrated ability to uncover and solve performance bottlenecks across layers (kernel, memory, networking, scheduler)
  • Experience building instrumentation, tracing, and profiling tools for ML models
  • Ability to work closely with ML researchers, translate novel model ideas into production systems
  • Ownership mindset and eagerness to dive deep into complex system challenges
  • Bonus: published research or open-source contributions in ML systems, inference optimization, or model serving


Pay Range Transparency

Databricks is committed to fair and equitable compensation practices. The pay range(s) for this role is listed below and represents the expected salary range for non-commissionable roles or on-target earnings for commissionable roles.  Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to job-related skills, depth of experience, relevant certifications and training, and specific work location. Based on the factors above, Databricks anticipates utilizing the full width of the range. The total compensation package for this position may also include eligibility for annual performance bonus, equity, and the benefits listed above. For more information regarding which range your location is in visit our page here.


Local Pay Range
$142,200$204,600 USD

About Databricks

Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.
Benefits
At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.

Our Commitment to Diversity and Inclusion

At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.

Compliance

If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.

Skills Required

  • BS/MS/PhD in Computer Science or related field
  • 3+ years of experience in performance-critical systems
  • Hands-on experience with CUDA and GPU programming
  • Solid understanding of ML inference internals
  • Experience with distributed systems design and operation
  • Ability to uncover and solve performance bottlenecks
  • Experience in building profiling tools for ML models
  • Ownership mindset towards complex system challenges
  • Published research or open source contributions in ML systems (bonus)

Databricks Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Databricks and has not been reviewed or approved by Databricks.

  • Equity Value & Accessibility Equity grants and RSUs are a major part of total compensation and are highlighted for meaningful upside potential. Stock-based awards and refreshers contribute to strong overall pay positioning across senior technical and go-to-market roles.
  • Healthcare Strength Medical, dental, and vision coverage are complemented by mental-health resources, an EAP, and wellness reimbursements. Health benefits are consistently framed as comprehensive and competitive.
  • Parental & Family Support Paid parental leave for all parents, fertility support, and backup care options provide tangible assistance for family needs. Hybrid work norms and team-day structure further ease coordination for caregivers.

Databricks Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
New York, NY
2,200 Employees
Year Founded: 2013

What We Do

As the leader in Unified Data Analytics, Databricks helps organizations make all their data ready for analytics, empower data science and data-driven decisions across the organization, and rapidly adopt machine learning to outpace the competition. By providing data teams with the ability to process massive amounts of data in the Cloud and power AI with that data, Databricks helps organizations innovate faster and tackle challenges like treating chronic disease through faster drug discovery, improving energy efficiency, and protecting financial markets.

Similar Jobs

Databricks Logo Databricks

Staff Software Engineer

Big Data • Machine Learning • Software • Analytics • Big Data Analytics
In-Office
San Francisco, CA, USA
2200 Employees
191K-233K Annually

PNC Bank Logo PNC Bank

Software Engineer

Machine Learning • Payments • Security • Software • Financial Services
Remote or Hybrid
USA
55000 Employees
45K-138K Annually

CrowdStrike Logo CrowdStrike

Sr. Director, AI Program Management (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
USA
10000 Employees
210K-300K Annually

MetLife Logo MetLife

MIM - Loan Asset Management Associate

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Los Angeles, CA, USA
43000 Employees
130K-150K Annually

Similar Companies Hiring

Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account