Senior Data Scientist (NLP and Unstructured Data Analytics)

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Washington, DC, USA
In-Office or Remote
Senior level
Artificial Intelligence • Software
The Role
Design, implement, and scale NLP and ML solutions to detect loan fraud and improper payments. Clean and analyze large unstructured corpora, build supervised and unsupervised models, support criminal investigations, document methodology to meet evidentiary standards, produce dashboards and reports, and coordinate with data engineering for production pipelines.
Summary Generated by Built In

Senior Data Scientist (NLP and Unstructured Data Analytics)

Location: Herndon, VA (Remote Work)

Must have a Public Trust Clearance

KEY RESPONSIBILITIES

  • Integrate and scale natural language processing methods to parse, clean, and analyze large corpora of unstructured and semi structured text, using optical character recognition, semantic similarity algorithms, and large language models as needed.
  • Design, develop, test, calibrate, and implement statistical and machine learning models targeting financial fraud, improper payments, and non compliance within SBA programs.
  • Build and refine supervised and unsupervised models, including regression, Bayesian, clustering, and ensemble approaches.
  • Review, maintain, and support all existing loan fraud indicators developed by TSD.
  • Perform data quality analysis on source tables and develop repeatable processes for combining and analyzing large data sources.
  • Collaborate directly with criminal investigators to determine and execute analytic strategies supporting loan fraud cases, and adhere closely to the federal rules of criminal procedure governing protected information, including Rule 6(e).
  • Develop case leads for SBA OIG investigations from model outcomes.
  • Document all methodology, test models, and production models in a form that satisfies criminal evidentiary requirements.
  • Build visualizations and dashboards conveying methodological choices, outcomes, and predictive capability, iterated on end user feedback.
  • Deliver findings in multiple registers: data summaries and visualizations for investigative staff, executive summaries for OIG leadership.
  • Coordinate with the data engineering seat so the architecture supports machine learning and text processing pipelines efficiently.
  • Create programming and automation techniques using SharePoint, Python, Excel, Power BI, Power Apps, and similar tools.
  • Identify new business questions that expand the scope of analysis and reporting.

Requirements

Required

Education

Master's, Ph.D., or doctorate level equivalent degree in data science, machine learning, computer science, mathematics, or a related field. Alternatively, ten years of applied work experience in any of the same fields.

  • 5+ years Designing, implementing, and maintaining advanced AI systems and predictive models, including both supervised and unsupervised models.
  • 5+ years Developing analytic rules and models using leading edge analytic tools and best practices.
  • 5+ years Developing regression, classification, and other statistical models to identify anomalies, patterns, and predictive variables.
  • 3+ years Providing data support for criminal investigations into financial fraud or abuse of government funds.
  • 3+ years Manipulating data in Python. Pandas is required.
  • 3+ years Working in a modern cloud environment: Azure, AWS, or GCP. Certifications preferred.
  • 2+ years Conducting advanced data analysis in SQL, specifically SQL Server and PostgreSQL.
  • 2+ years Developing and scaling natural language processing solutions.
  • 2+ years Presenting methods and findings to technical and non technical stakeholders, both orally and in written products and visualizations.

PREFERRED QUALIFICATIONS

  • Production experience with named entity recognition and entity resolution across messy document corpora.
  • Retrieval augmented generation, vector stores, embeddings, and semantic search at scale.
  • Large language model integration under federal security constraints, including boundary controlled deployment and prompt versioning.
  • Optical character recognition pipelines applied to scanned or low quality source documents.
  • Topic modeling, document classification, or clustering applied to audit, legal, or investigative text.
  • Cloud certification in Azure, AWS, or GCP.

Benefits

We are proud to offer competitive compensation and benefits packages to include

  • Medical 
  • Dental
  • Vision
  • Basic Life 
  • Health Saving Account
  • 401K matching
  • Three weeks of PTO/Sick
  • 11 Paid Holidays
  • Pre-Approved Online Training

Skills Required

  • Must have an active Public Trust clearance
  • Master's, Ph.D., or equivalent advanced degree in data science, ML, CS, mathematics, or 10 years applied experience
  • 5+ years designing, implementing, and maintaining advanced AI systems and predictive models
  • 5+ years developing analytic rules and models using leading-edge analytic tools
  • 5+ years developing regression, classification, and other statistical models
  • 3+ years providing data support for criminal investigations into financial fraud or abuse of government funds
  • 3+ years manipulating data in Python (Pandas required)
  • 3+ years working in a modern cloud environment (Azure, AWS, or GCP); certifications preferred
  • 2+ years conducting advanced data analysis in SQL, specifically SQL Server and PostgreSQL
  • 2+ years developing and scaling natural language processing solutions
  • 2+ years presenting methods and findings to technical and non-technical stakeholders, orally and in writing
  • Experience creating automation and programs using SharePoint, Excel, Power BI, Power Apps
  • Preferred: production experience with named entity recognition and entity resolution across messy document corpora
  • Preferred: retrieval augmented generation, vector stores, embeddings, and semantic search at scale
  • Preferred: large language model integration under federal security constraints, including boundary controlled deployment
  • Preferred: OCR pipelines applied to scanned or low quality source documents
  • Preferred: topic modeling, document classification, or clustering applied to audit, legal, or investigative text
  • Preferred: cloud certification in Azure, AWS, or GCP
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Leesburg, VA
11 Employees
Year Founded: 2017

What We Do

We Love Building Innovative Digital Solutions using Automation and A.I./ML to Solve Complex Problems for our Customer's Mission

Similar Jobs

MetLife Logo MetLife

Consultant

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
100K-130K Annually

MetLife Logo MetLife

Customer Care Advocate - Virtual 8.3.26 - 18516

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
42K-42K Annually

MetLife Logo MetLife

AVP, Disability & Absence Claim Operations

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Remote or Hybrid
United States
43000 Employees
200K-235K Annually

Shield AI Logo Shield AI

Manager, GTM Sales Enablement (R5468)

Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Remote
USA
100K-150K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account