Data Scientist

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office or Remote
Senior level
Software • Cybersecurity
Sonatype is the software supply chain management company.
The Role
Lead applied AI projects from prototype to production, building and validating ML and GenAI models (LLMs, embeddings, RAG, agentic workflows) for malicious-behavior, anomaly, and fraud detection. Advise product and engineering teams, design experiments and evaluation pipelines, translate research into scalable APIs/workflows, and ensure data governance, privacy, and ethical practices.
Summary Generated by Built In

Sonatype is the software supply chain management company that invented componentized software development and pioneered the software supply chain category. As leaders in the open-source community and the DevSecOps industry, we run the world’s largest repository of Java open-source components—Maven Central.

Our groundbreaking, full-spectrum platform empowers customers to rapidly create, deploy, and maintain innovative software at scale, all while aligning directly to their business needs. Trusted by more than 2,000 organizations—including 70% of the Fortune 100—and over 15 million software developers, Sonatype’s tools and guidance help deliver exceptional, secure software.

From inventing modern artifact management with Nexus Repository to introducing the world’s only solution that halts malicious open-source malware in its tracks, we’re committed to constant innovation. We leverage AI/ML to give our clients, developers, and the industry complete confidence in the quality, automation, and security of their software. 

Learn more at www.sonatype.com

The Role

    We're looking for a Data Scientist to join our growing AI & Data Science team. You'll operate as an internal AI consultant and technical lead, helping multiple teams across Sonatype apply machine learning and generative AI to real-world problems — from malicious-behavior and anomaly detection in our security data, to developer- and analyst-facing GenAI experiences.

    You'll explore complex datasets, design experiments, build and validate models, and collaborate closely with product, engineering, and security experts to turn research ideas into practical, scalable solutions. We have a mature data engineering team, so you can focus on doing what you do best — building and shipping models.

    This role is ideal for someone who thrives on autonomy, loves translating ambiguous ideas into working systems, and enjoys working across boundaries rather than staying in a single product lane.

What you'll do:

  • Lead applied AI projects from concept to impact — prototype, validate, and help teams deploy practical ML and GenAI solutions.

  • Act as an internal consultant across product, engineering, security, and research teams: scope problems, evaluate approaches, and advise on ML/AI best practices and productive use of generative technologies.

  • Lead the research, development, and deployment of models for use cases such as malicious behavior detection, anomaly detection, and fraud analysis — using techniques ranging from classical ML to LLMs, embeddings, retrieval-augmented generation, and agentic workflows.

  • Design robust experiments and establish evaluation pipelines for model reliability, accuracy, and business impact (cross-validation, drift monitoring, ground-truth evaluation).

  • Bridge research and production: translate research insights into scalable APIs, tools, or workflows that enable other teams to adopt AI effectively.

  • Explore new techniques (LLMs, embeddings models, RAG, agentic workflows) to enhance developer and security experiences.

  • Communicate technical concepts, tradeoffs, and recommendations clearly to both technical and non-technical stakeholders through presentations, documentation, and collaboration; mentor peers and help elevate the organization's AI literacy and capabilities.

  • Partner with our data governance team to ensure compliance with data-privacy regulations and ethical considerations when working with customer data.

What you bring:

  • 5+ years of hands-on experience in applied data science, machine learning, AI engineering, or AI research.

  • Computer Science or equivalent technical degree strongly preferred

  • Strong Python skills and practical experience with data and AI libraries/platforms such as Databricks, and LLM APIs, scikit-learn

  • Experience building and shipping ML or GenAI applications—from early prototype through usable internal or customer-facing workflows.

  • Deep familiarity with modern LLM ecosystems, including OpenAI, Anthropic/Claude, Hugging Face, and open-weight models.

  • Ability to select models and design effective LLM applications using prompting, context management, structured outputs, retrieval, and tool use.

  • Experience building agentic or multi-step AI workflows with LangGraph, LangChain, Semantic Kernel, or similar orchestration frameworks.

  • Strong evaluation mindset: defining useful quality metrics, building representative evaluation datasets, assessing reliability, and making data-driven tradeoffs.

  • Comfortable working with large, messy, structured, and unstructured data to produce features, insights, and clear visualizations.

  • Proficiency with Git, testing, code review, and collaborative software-development practices.

  • Practical, balanced judgment: comfortable exploring emerging AI capabilities while building maintainable, secure, dependable systems.

  • Proactive and accountable, with strong written and verbal communication skills across technical and non-technical partners.

It'd be great if you had:

  • Strong MLOps experience, including MLflow or comparable tooling, experiment tracking, reproducible pipelines, model/application versioning, CI/CD, serving, and production monitoring.

  • Experience operating ML or GenAI systems at scale, including observability, tracing, incident response, and data or model-drift detection.

  • Experience with Databricks ML, AWS SageMaker, Azure ML, or similar managed ML platforms.

  • Familiarity with MCP, agent-tool integrations, LLM guardrails, and production safety practices.

  • Experience with AI-assisted development tools such as Copilot, Claude Code, or Codex.

  • Exposure to cybersecurity, fraud detection, anomaly detection, code analysis, or software supply-chain security.

  • Experience with PySpark and production data pipelines.

  • Experience working within a software product company or SaaS.

Things we're proud of:

  • 2026 Gartner® Magic Quadrant™ Leader for Software Supply Chain Security
  • 2026 Celebrating 15 Years of Sonatype Research Labs – Industry-leading software supply chain and open source security research
  • 2026 Founding Member of the Linux Foundation Initiative for Open Source Sustainability
  • 2026 State of the Software Supply Chain® Report – Continuing industry leadership in software supply chain security and AI security research
  • 2025 Visionary in Gartner® Magic Quadrant™ for Application Security Testing!
  • 2025 AI Compliance Solution of the Year - AI Breakthrough Awards
  • 2025 DEVIES Award to our SBOM Manager for a new product for its innovation and impact in developer technology
  • 2024 Industry Leader in Forrester-Wave for Software Composition Analysis (2024 Q4 report)
  • Constellation AST Shortlist: Sonatype has been listed on the Constellation ShortList™ for Application Security Testing for 2024
  • Data Breakthrough Awards: Sonatype was announced as a 2024 winner in the "Open Source Data Solution of the Year."
  • SD Times: Best in Show Security
  • Fast Company Best Workplaces for Innovators 2024
  • The Herd Top 100 Private Software Companies 2024
  • Diversity & Inclusion Working Groups
  • Parental Leave Policy
  • Paid Volunteer Time Off (VTO)

At Sonatype, we value diversity and inclusivity. We offer perks such as parental leave, diversity and inclusion working groups, and flexible working practices to allow our employees to show up as their whole selves. We are an equal-opportunity employer, and we do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. If you have a disability or special need that requires accommodation, please do not hesitate to let us know.

Skills Required

  • 5+ years of hands-on experience in applied data science, machine learning, AI engineering, or AI research.
  • Strong Python skills.
  • Practical experience with data and AI libraries/platforms such as Databricks and scikit-learn.
  • Practical experience with LLM APIs and modern LLM ecosystems (OpenAI, Anthropic/Claude, Hugging Face, open-weight models).
  • Experience building and shipping ML or GenAI applications from prototype to usable internal or customer-facing workflows.
  • Experience designing LLM applications using prompting, context management, structured outputs, retrieval, and tool use.
  • Experience building agentic or multi-step AI workflows with LangGraph, LangChain, Semantic Kernel, or similar.
  • Strong evaluation mindset: defining metrics, building evaluation datasets, assessing reliability, and monitoring drift.
  • Comfortable working with large, messy, structured and unstructured data to produce features, insights, and visualizations.
  • Proficiency with Git, testing, code review, and collaborative software-development practices.
  • Computer Science or equivalent technical degree.
  • Strong MLOps experience (MLflow or comparable tooling, experiment tracking, reproducible pipelines, model/versioning, CI/CD, serving, monitoring).
  • Experience operating ML or GenAI systems at scale, including observability, tracing, incident response, and drift detection.
  • Experience with Databricks ML, AWS SageMaker, Azure ML, or similar managed ML platforms.
  • Familiarity with agent-tool integrations, LLM guardrails, and production safety practices.
  • Experience with AI-assisted development tools (Copilot, Claude Code, Codex).
  • Exposure to cybersecurity, fraud detection, anomaly detection, code analysis, or software supply-chain security.
  • Experience with PySpark and production data pipelines.
  • Experience working within a software product company or SaaS.
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Fulton, MD
600 Employees
Year Founded: 2008

What We Do

The Sonatype journey started almost 15 years ago, just as the concept of “open source” software development was gaining steam. From our humble beginning as core contributors to Apache Maven, to supporting the world’s largest repository of open source components (Central), to distributing the world's most popular repository manager (Nexus), we’ve played a meaningful role in helping the world embrace the power of open innovation. We empower developers and security professionals with intelligent tools to innovate more securely at scale. Our platform addresses every element of an organization’s entire software development life cycle, including third-party open source code, first-party source code, and containerized code. Sonatype identifies critical security vulnerabilities and code quality issues and reports results directly to developers when they can most effectively fix them. This helps organizations develop consistently high-quality, secure software which fully meets their business needs and those of their end-customers and partners. More than 2,000 organizations, including 70% of the Fortune 100, and 15 million software developers rely on our tools and guidance to help them deliver and maintain exceptional and secure software.

Why Work With Us

We're on a mission to change how the world innovates by making software development easier. Already used by 15 million developers, we have lofty goals for our technology to be in the hands of every engineering team. And, we need you to do that. Join us!

Gallery

Gallery

Similar Jobs

Square Logo Square

Data Scientist

eCommerce • Fintech • Hardware • Payments • Software • Financial Services
Remote or Hybrid
8 Locations
12000 Employees
240K-359K Annually

Block Logo Block

Data Scientist

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
240K-359K Annually

Quora Logo Quora

Data Scientist

Artificial Intelligence • Consumer Web • Digital Media • Machine Learning • Software
In-Office or Remote
4 Locations
240 Employees
122K-178K Annually

Toast Logo Toast

Data Scientist

Cloud • Fintech • Food • Information Technology • Software • Hospitality
Remote
Canada
5000 Employees
127K-203K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account