What You’ll Do
- Design, build, and maintain scalable data pipelines and ETL/ELT processes
- Architect and optimize data models and storage solutions for analytics and operational use
- Collaborate with data scientists, analysts, and engineers to deliver trusted, high-quality datasets
- Own and evolve parts of our data platform (e.g., Airflow, dbt, Spark, Redshift, or Snowflake)Implement observability, alerting, and data quality monitoring for critical pipelines
- Drive best practices in data engineering, including documentation, testing, and CI/CD
- Contribute to the design and evolution of our next-generation data lakehouse architecture
What We’re Looking For
- 8+ years of experience as a Data Engineer or in a similar backend engineering role
- Strong programming skills in Python, Scala, or Java Hands-on experience with HBase or similar NoSQL columnar stores
- Hands-on experience with distributed data systems like Spark, Kafka, or Flink
- Proficient in writing complex SQL and optimizing queries for performance
- Experience building and maintaining robust ETL/ELT pipelines in production
- Familiarity with workflow orchestration tools (Airflow, Dagster, or similar)
- Understanding of data modeling techniques (star schema, dimensional modeling, etc.)
Why You’ll Love Working Here
- Data with purpose: Work on problems that directly impact how the world builds secure software
- Modern tooling: Leverage the best of open-source and cloud-native technologies
- Collaborative culture: Join a passionate team that values learning, autonomy, and impact
Skills Required
- 8+ years of experience as a Data Engineer or in a similar backend engineering role
- Strong programming skills in Python, Scala, or Java
- Hands-on experience with HBase or similar NoSQL columnar stores
- Hands-on experience with distributed data systems such as Spark, Kafka, or Flink
- Proficiency writing complex SQL and optimizing queries for performance
- Experience building and maintaining robust ETL/ELT pipelines in production
- Familiarity with workflow orchestration tools such as Airflow or Dagster
- Understanding of data modeling techniques, including star schema and dimensional modeling
What We Do
The Sonatype journey started almost 15 years ago, just as the concept of “open source” software development was gaining steam. From our humble beginning as core contributors to Apache Maven, to supporting the world’s largest repository of open source components (Central), to distributing the world's most popular repository manager (Nexus), we’ve played a meaningful role in helping the world embrace the power of open innovation. We empower developers and security professionals with intelligent tools to innovate more securely at scale. Our platform addresses every element of an organization’s entire software development life cycle, including third-party open source code, first-party source code, and containerized code. Sonatype identifies critical security vulnerabilities and code quality issues and reports results directly to developers when they can most effectively fix them. This helps organizations develop consistently high-quality, secure software which fully meets their business needs and those of their end-customers and partners. More than 2,000 organizations, including 70% of the Fortune 100, and 15 million software developers rely on our tools and guidance to help them deliver and maintain exceptional and secure software.
Why Work With Us
We're on a mission to change how the world innovates by making software development easier. Already used by 15 million developers, we have lofty goals for our technology to be in the hands of every engineering team. And, we need you to do that. Join us!
Gallery






