Staff Data Engineer

Posted 3 Days Ago
Hiring Remotely in Canada
Remote
Senior level
Software • Cybersecurity
Sonatype is the software supply chain management company.
The Role
Design, build, and maintain scalable data pipelines, optimize Databricks/Spark and Delta Lake architectures, deliver trusted datasets for analytics and ML, implement observability and data quality, and drive platform architecture and team mentorship.
Summary Generated by Built In

Sonatype is the software supply chain security company. We provide the world’s best end-to-end software supply chain security solution, combining the only proactive protection against malicious open source, the only enterprise grade SBOM management and the leading open source dependency management platform. This empowers enterprises to create and maintain secure, quality, and innovative software at scale.

As founders of Nexus Repository and stewards of Maven Central, the world’s largest repository of Java open-source software, we are software pioneers and our open source expertise is unmatched. We empower innovation with an unparalleled commitment to build faster, safer software and harness AI and data intelligence to mitigate risk, maximize efficiencies, and drive powerful software development.

More than 2,000 organizations, including 70% of the Fortune 100 and 15 million software developers, rely on Sonatype to optimize their software supply chains.

About the role:

    We’re looking for a Staff Data Engineer to join our growing Data Platform team. You’ll play a key role in designing and scaling the infrastructure and pipelines that power analytics, machine learning, and business intelligence across Sonatype.You’ll work closely with stakeholders across product, engineering, and business teams to ensure data is reliable, accessible, and actionable. This role is ideal for someone who thrives on solving complex data challenges at scale and enjoys building high-quality, maintainable systems.

    At Sonatype, we:
  • Use data with purpose: you'll get the chance to work on problems that directly impact how the world builds secure software
  • Use modern tooling: you'll get the chance to leverage the best of open-source and cloud-native technologies
  • Have a deep collaborative culture: you'll be joining a passionate team that values learning, autonomy, and impact

What you'll do:

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes

  • Architect and optimize data models and storage solutions for analytics and operational use

  • Collaborate with data scientists, analysts, and engineers to deliver trusted, high-quality datasets

  • Own and evolve parts of our data platform using Databricks and Spark

  • Implement observability, alerting, and data quality monitoring for critical pipelines

  • Drive best practices in data engineering, including documentation, testing, and CI/CD

  • As a Staff Engineer you will help drive long-term architectural vision and mentor the team on engineering best practices, while partnering with stakeholders to ensure data solutions support business outcomes.

  • Contribute to the design and evolution of our next-generation data lakehouse architecture

What you bring:

  • 8+ years of experience as a Data Engineer or in a similar backend engineering role

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field

  • Databricks Optimization: Tune Spark jobs, optimize join performance, and manage Delta Lake architecture for batch and streaming data.

  • Experience leveraging AI-assisted development tools and AI/ML technologies to improve data engineering workflows, developer productivity, data quality and ops.

  • Strong programming skills in Python, Scala, or Java

  • Hands-on experience with distributed data systems like Spark or Kafka

  • Proficient in writing complex SQL and NoSQL queries and optimizing queries for performance

  • Experience building and maintaining robust ETL/ELT pipelines in production

  • Understanding of data modeling techniques (star schema, dimensional modeling, etc.)

It’d be great if you also had:

  • Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data

  • A track record of improving data platform reliability, scalability, performance, and cost efficiency

  • Familiarity with workflow orchestration tools (Airflow, Dagster, or similar)

  • Hands-on experience with cloud data platforms, particularly AWS

  • Familiarity with modern table formats such as Delta Lake, Apache Iceberg, or Apache Hudi

  • Experience implementing data observability, lineage, governance, and automated data quality frameworks

  • Experience designing real-time or streaming data architectures using data lake technologies

Things we are proud of:

  • 2026 Gartner® Magic Quadrant™ Leader for Software Supply Chain Security
  • 2026 Celebrating 15 Years of Sonatype Research Labs – Industry-leading software supply chain and open source security research
  • 2026 Founding Member of the Linux Foundation Initiative for Open Source Sustainability
  • 2026 State of the Software Supply Chain® Report – Continuing industry leadership in software supply chain security and AI security research
  • 2025 Visionary in Gartner® Magic Quadrant™ for Application Security Testing!
  • 2025 AI Compliance Solution of the Year - AI Breakthrough Awards
  • 2025 DEVIES Award to our SBOM Manager for a new product for its innovation and impact in developer technology
  • 2024 Industry Leader in Forrester-Wave for Software Composition Analysis (2024 Q4 report)
  • Constellation AST Shortlist: Sonatype has been listed on the Constellation ShortList™ for Application Security Testing for 2024
  • Data Breakthrough Awards: Sonatype was announced as a 2024 winner in the "Open Source Data Solution of the Year."
  • SD Times: Best in Show Security
  • Fast Company Best Workplaces for Innovators 2024
  • The Herd Top 100 Private Software Companies 2024
  • Diversity & Inclusion Working Groups
  • Parental Leave Policy
  • Paid Volunteer Time Off (VTO)

At Sonatype, we value diversity and inclusivity. We offer perks such as parental leave, diversity and inclusion working groups, and flexible working practices to allow our employees to show up as their whole selves. We are an equal-opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. If you have a disability or special need that requires accommodation, please do not hesitate to let us know.
 

Skills Required

  • 8+ years of experience as a Data Engineer or similar backend engineering role
  • Bachelor's degree in Computer Science, Engineering, or related technical field
  • Databricks optimization: tune Spark jobs, optimize joins, manage Delta Lake architecture for batch and streaming
  • Experience leveraging AI-assisted development tools and AI/ML technologies in data engineering workflows
  • Strong programming skills in Python, Scala, or Java
  • Hands-on experience with distributed data systems such as Spark or Kafka
  • Proficient in writing and optimizing complex SQL and NoSQL queries
  • Experience building and maintaining robust ETL/ELT pipelines in production
  • Understanding of data modeling techniques (star schema, dimensional modeling)
  • Familiarity with workflow orchestration tools (Airflow, Dagster, or similar)
  • Hands-on experience with cloud data platforms, particularly AWS
  • Familiarity with modern table formats such as Apache Iceberg or Apache Hudi
  • Experience implementing data observability, lineage, governance, and automated data quality frameworks
  • Experience designing real-time or streaming data architectures using data lake technologies
  • Familiarity with software supply chain, cybersecurity, or large-scale software ecosystem data
  • Track record of improving data platform reliability, scalability, performance, and cost efficiency
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Fulton, MD
600 Employees
Year Founded: 2008

What We Do

The Sonatype journey started almost 15 years ago, just as the concept of “open source” software development was gaining steam. From our humble beginning as core contributors to Apache Maven, to supporting the world’s largest repository of open source components (Central), to distributing the world's most popular repository manager (Nexus), we’ve played a meaningful role in helping the world embrace the power of open innovation. We empower developers and security professionals with intelligent tools to innovate more securely at scale. Our platform addresses every element of an organization’s entire software development life cycle, including third-party open source code, first-party source code, and containerized code. Sonatype identifies critical security vulnerabilities and code quality issues and reports results directly to developers when they can most effectively fix them. This helps organizations develop consistently high-quality, secure software which fully meets their business needs and those of their end-customers and partners. More than 2,000 organizations, including 70% of the Fortune 100, and 15 million software developers rely on our tools and guidance to help them deliver and maintain exceptional and secure software.

Why Work With Us

We're on a mission to change how the world innovates by making software development easier. Already used by 15 million developers, we have lofty goals for our technology to be in the hands of every engineering team. And, we need you to do that. Join us!

Gallery

Gallery

Similar Jobs

Remote
Québec, QC, CAN
1926 Employees
Remote
Canada
404 Employees
152K-238K Annually

Walmart Global Tech Logo Walmart Global Tech

Data Engineer

Big Data • Cloud • Logistics • Machine Learning • Retail
Remote or Hybrid
8 Locations
578950 Employees
110K-286K Annually
Remote
Canada
322 Employees
136K-170K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account