Site Reliability Engineer

Posted 6 Days Ago
Be an Early Applicant
Hiring Remotely in Costa Rica
Remote
Senior level
Big Data • Machine Learning • Software • Analytics • Big Data Analytics
The Role
As a Site Reliability Engineer, you will automate, optimize, and manage cloud infrastructure while ensuring system reliability and collaborating cross-functionally.
Summary Generated by Built In

GAQ127R40 

Team: IT Infrastructure and Operations

About the Role

At Databricks Information Technology, we are a product-led organization transforming how we work—from the ease of using our IT services to the applications we develop to scale seamlessly during rapid growth.

As a Site Reliability Engineer (SRE), you will bridge the gap between software engineering and systems architecture. You will be a core contributor to the IT Infrastructure team, owning the evolution of core infrastructure and observability platforms. This role requires a strong software engineering mindset and deep technical breadth to deliver high-quality, scalable solutions for "immature" system problems. Your focus will be on building resilient, automated infrastructure that empowers development teams and ensures our cloud environment is cost-optimized, secure, and highly available.

The Impact You Will Have
  • Architect and Automate: Design and deploy production-grade infrastructure on cloud platforms (AWS/Azure) using Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Reliability and Performance Engineering:Optimize system performance, architecture, and scaling to ensure maximum uptime and minimal latency for critical IT services.
  • CI/CD Excellence: Architect robust deployment pipelines (e.g., GitHub Actions), managing both hosted and self-hosted runners for specialized build requirements.
  • Observable by Default: Create underlying infrastructure to ensure new internal applications are secure and have logging, metrics and alerts enabled by default.
  • Agentic ToolingI: Build internal AI plugins, and automation scripts to streamline developer workflows and enhance operational efficiency.
  • Incident Response: Focus on subsequent data usage, incident management workflows, and creating necessary dashboards to maintain service health. Participate in a shared on-call rotation, leading rapid incident response and technical troubleshooting for production outages.Facilitate blameless post-mortems to identify root causes and implement permanent preventive engineering solutions.
  • Partner Cross-Functionally: Collaborate with Security, Engineering, and Support teams to deliver real business outcomes.
What We Look For
  • Software Engineering Expertise: 5+ years of production-level experience with strong proficiency in Python (non-negotiable).
  • Infrastructure as Code (IaC): Expert-level proficiency in Terraform (modules, state management) or Pulumi.
  • Cloud & Containers: Hands-on experience with AWS, Azure, or GCP, along with Kubernetes, Docker, and containerization concepts.
  • Observability Mindset: Deep understanding of observability pillars (logging, metrics, tracing) and experience with tools such as Datadog, Prometheus, or ELK.
  • Distributed Systems: Proficiency in running systems using concepts like Kafka or messaging queues.
  • CI/CD Proficiency: Advanced knowledge of GitHub Actions and GitHub Runners.
  • Independent Execution: Ability to take ownership of ambiguous projects, follow a vision set by tech leads, and execute independently with minimal guidance.

About Databricks

Databricks is the data and AI company. More than 10,000 organizations worldwide — including Comcast, Condé Nast, Grammarly, and over 50% of the Fortune 500 — rely on the Databricks Data Intelligence Platform to unify and democratize data, analytics and AI. Databricks is headquartered in San Francisco, with offices around the globe and was founded by the original creators of Lakehouse, Apache Spark™, Delta Lake and MLflow. To learn more, follow Databricks on Twitter, LinkedIn and Facebook.
Benefits
At Databricks, we strive to provide comprehensive benefits and perks that meet the needs of all of our employees. For specific details on the benefits offered in your region click here.

Our Commitment to Diversity and Inclusion

At Databricks, we are committed to fostering a diverse and inclusive culture where everyone can excel. We take great care to ensure that our hiring practices are inclusive and meet equal employment opportunity standards. Individuals looking for employment at Databricks are considered without regard to age, color, disability, ethnicity, family or marital status, gender identity or expression, language, national origin, physical and mental ability, political affiliation, race, religion, sexual orientation, socio-economic status, veteran status, and other protected characteristics.

Compliance

If access to export-controlled technology or source code is required for performance of job duties, it is within Employer's discretion whether to apply for a U.S. government license for such positions, and Employer may decline to proceed with an applicant on this basis alone.

Skills Required

  • 5+ years of production-level experience with strong proficiency in Python
  • Expert-level proficiency in Terraform or Pulumi
  • Hands-on experience with AWS, Azure, or GCP, Kubernetes, and Docker
  • Deep understanding of observability with tools like Datadog, Prometheus, or ELK
  • Proficiency in running distributed systems using Kafka or messaging queues
  • Advanced knowledge of GitHub Actions and Runners
  • Ability to take ownership of ambiguous projects

Databricks Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Databricks and has not been reviewed or approved by Databricks.

  • Equity Value & Accessibility Equity grants and RSUs are a major part of total compensation and are highlighted for meaningful upside potential. Stock-based awards and refreshers contribute to strong overall pay positioning across senior technical and go-to-market roles.
  • Healthcare Strength Medical, dental, and vision coverage are complemented by mental-health resources, an EAP, and wellness reimbursements. Health benefits are consistently framed as comprehensive and competitive.
  • Parental & Family Support Paid parental leave for all parents, fertility support, and backup care options provide tangible assistance for family needs. Hybrid work norms and team-day structure further ease coordination for caregivers.

Databricks Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
New York, NY
2,200 Employees
Year Founded: 2013

What We Do

As the leader in Unified Data Analytics, Databricks helps organizations make all their data ready for analytics, empower data science and data-driven decisions across the organization, and rapidly adopt machine learning to outpace the competition. By providing data teams with the ability to process massive amounts of data in the Cloud and power AI with that data, Databricks helps organizations innovate faster and tackle challenges like treating chronic disease through faster drug discovery, improving energy efficiency, and protecting financial markets.

Similar Jobs

Akamai Technologies Logo Akamai Technologies

Site Reliability Engineer

Cloud • Security • Software • Cybersecurity
In-Office or Remote
2 Locations
10285 Employees
15M-32M Annually

Lambda Logo Lambda

Senior Site Reliability Engineer

Artificial Intelligence • Cloud • Machine Learning • Infrastructure as a Service (IaaS)
Remote or Hybrid
4 Locations
750 Employees
240K-356K Annually

TransUnion Logo TransUnion

Consultant

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Remote or Hybrid
Heredia, Ulloa, Lagunilla, CRI
13000 Employees

TransUnion Logo TransUnion

Advisor, Data Science and Analytics

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Remote or Hybrid
Heredia, Ulloa, Lagunilla, CRI
13000 Employees

Similar Companies Hiring

Bellagent Thumbnail
Artificial Intelligence • Machine Learning • Business Intelligence • Generative AI
Chicago, IL
20 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account