Senior Applied Data Scientist | NDA

Posted Yesterday
Be an Early Applicant
Warsaw, Warszawa, Mazowieckie, POL
Hybrid
Senior level
Information Technology
The Role
Develop and evaluate machine learning, embedding, and LLM-based methods for large-scale entity resolution. Improve matching of complex business records using scoring, ranking, NLP, classification, clustering, and similarity techniques. Define benchmarks, metrics, experiments, and error analyses to measure quality improvements. Partner with engineering teams to productionize successful models while balancing accuracy, scalability, cost, latency, and explainability.
Summary Generated by Built In

GT was founded in 2019 by a former Apple, Nest, and Google executive. GT’s mission is to connect the world’s best talent with product careers offered by high-growth companies in the UK, USA, Canada, Germany, and the Netherlands.

On behalf of our client, GT is looking for a Senior Applied Data Scientist interested in developing and testing new ML, embedding, and LLM-based approaches to solve complex data matching problems at scale.

About the Client

Our client is a leading global management consultancy known for tackling some of the world’s most complex business challenges. With a focus on strategy, transformation, and performance improvement, the firm partners with major organizations across industries to drive lasting impact.

About the Role

We are looking for a Senior Applied Data Scientist to improve how entity resolution is performed at scale.

You will develop and test new ML, embedding, and LLM-based approaches for matching complex business records across multiple data sources.

The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches.

A key part of the role is exploring how newer foundation-model techniques can improve matching quality while remaining practical and scalable for very large datasets.

Responsibilities:

Develop better ways to match company records

  • Build new ML, embedding, and LLM-based approaches for matching entities

  • Improve how the system handles messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies.

  • Develop scoring and ranking approaches to distinguish accurate matches from duplicates, similar-looking records, and unrelated entities.

  • Evaluate and implement AI and machine learning techniques to improve matching quality while considering accuracy, scalability, and cost.

  • Design approaches that can operate efficiently at scale, taking model usage and computational cost into consideration.

Improve evaluation, experimentation, and match quality

  • Define and improve methods for evaluating match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review effort.

  • Assist in building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout.

  • Explore LLM-assisted review and validation to assess matching performance and benchmark more scalable approaches.

  • Turn ambiguous matching problems into clear hypotheses, experiments, metrics, and recommendations.

Partner with engineering to bring successful ideas into production

  • Work closely with data engineering and software engineering teams to turn promising prototypes into production-ready matching logic.

  • Provide engineering partners with clear model specifications, evaluation results, expected behavior, edge cases, and rollout requirements.

  • Help determine the most appropriate matching techniques based on data characteristics, confidence levels, and cost considerations.

  • Continuously evaluate matching performance, investigate regressions, and recommend improvements to models and matching logic.

  • Clearly communicate technical tradeoffs related to matching performance, scalability, cost, latency, explainability, and operational considerations.

Essential knowledge, skills & experience:
  • 5–8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role.

  • Strong applied ML fundamentals, with hands-on experience building and evaluating models on real data.

  • Excellent Python and SQL skills.

  • Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques.

  • Hands-on experience training supervised and unsupervised models, including classification and NLP tasks.

  • Working knowledge of neural network and transformer architectures.

  • Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret.

  • Experience retraining a taxonomy classifier or maintaining classification models in production.

  • Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved.

  • Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners.

Nice-to-have:
  • Experience with entity resolution, record linkage, deduplication, or similar matching problems.

  • Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation.

  • Experience applying LLMs or embeddings to business problems where cost and scale matter.

  • Exposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQuery.

  • Familiarity with company, domain, website, firmographic, or other business-entity data.

Interview Steps:
  1. GT interview with Recruiter

  2. Technical interview

  3. Final interview

Skills Required

  • 5-8 years of relevant experience in Data Science, Applied Data Science, Applied Machine Learning, or a similar role
  • Strong applied machine learning fundamentals and hands-on experience building and evaluating models on real data
  • Excellent Python and SQL skills
  • Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques
  • Hands-on experience training supervised and unsupervised models, including classification and NLP tasks
  • Working knowledge of neural network and transformer architectures
  • Proficiency with TensorFlow, PyTorch, and PyCaret
  • Experience retraining a taxonomy classifier or maintaining classification models in production
  • Ability to define baselines, metrics, test sets, and error analysis to evaluate quality improvements
  • Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners
  • Experience with entity resolution, record linkage, deduplication, or similar matching problems
  • Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation
  • Experience applying LLMs or embeddings to business problems where cost and scale matter
  • Exposure to Spark, Snowflake, Databricks, or BigQuery
  • Familiarity with company, domain, website, firmographic, or other business-entity data
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
686 Employees
Year Founded: 2019

What We Do

GT was founded in 2019 by a former executive from Apple, Nest, and Google. GT’s mission is to build teams and products that address some of the more challenging requirements from fast-growth clients in Europe and North America.

Similar Jobs

JPMorganChase Logo JPMorganChase

Controller

Financial Services
Hybrid
2 Locations
289097 Employees

Sprout Social Logo Sprout Social

Staff Software Engineer

Marketing Tech • Social Media • Software • Analytics • Business Intelligence
Easy Apply
Remote or Hybrid
Poland
1400 Employees
29K-44K Annually

Samsara Logo Samsara

Senior Software Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
Poland
4000 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account