Senior Data Scientist

Posted 4 Days Ago
5 Locations
In-Office or Remote
Senior level
Information Technology
The Role
Build and deploy production machine learning and AI solutions on Databricks and Azure. Responsibilities include data discovery, feature engineering, model development, experiment tracking, model serving, monitoring, retraining, LLM and RAG use cases, responsible AI documentation, and modernization of legacy R, Stata, and SAS workloads. The role also involves client consultation, cloud adoption, training, mentoring, and communicating results to technical and executive audiences.
Summary Generated by Built In

BlueFlag is looking for a Senior Data Scientist to support a premier data and AI platform for the Department of Veterans Affairs. You'll work directly with internal clients to turn operational problems into production machine learning and AI solutions on Databricks in Azure, and help their teams adopt modern cloud tooling. This role is hands-on. We want someone who has taken models from a notebook to production and kept them running.

What You'll Do

  • Consult with internal clients to frame business problems as analytical or ML problems, define success metrics, and set realistic scope
  • Build data products and workflows that support critical operations, from source data discovery through production deployment
  • Migrate and modernize client workloads from legacy environments (including R, Stata, and SAS) to Python and Databricks, with measurable gains in runtime, cost, or reliability
  • Develop, train, and validate ML models (classification, regression, forecasting, clustering, anomaly detection) for complex business problems
  • Engineer features and build reusable, governed feature pipelines on large datasets with PySpark and SQL
  • Track experiments, register models, and manage model versions and promotion through MLflow
  • Deploy models for batch scoring and real-time serving, and automate retraining with scheduled jobs and CI/CD pipelines
  • Monitor models in production for performance, data drift, and data quality, and set thresholds and alerts that trigger review or retraining
  • Lead AI adoption with clients, including LLM and agentic use cases such as retrieval-augmented generation (RAG), document summarization, and classification
  • Evaluate LLM and agent outputs for accuracy, groundedness, and safety before and after release
  • Document models (purpose, data, assumptions, limitations, evaluation results) so they hold up to governance and responsible AI review
  • Build reference use cases with tutorials, reference code, and training to drive adoption of cloud tooling for concrete business results
  • Host office hours and pair with client analysts and data scientists to upskill them
  • Present findings and model results clearly to technical and non-technical audiences, including leadership

Why Join BlueFlag

At BlueFlag, we're passionate about leveraging cutting-edge technology to make a real difference. You'll be at the forefront of cloud innovation, working on projects that directly impact people's lives. We offer a high-growth, entrepreneurial environment that values fresh ideas and authentic teamwork.

If you're ready to take your data engineering career to new heights and contribute to meaningful projects that push the boundaries of technology, we want to hear from you. Join BlueFlag and be part of a team that's shaping the future of AI solutions!


Requirements
  • Bachelor's degree in Engineering, Computer Science, Statistics, Mathematics, Systems, Business, or a related scientific or technical discipline, and 15+ years of experience (or commensurate experience)
  • Proficient in Python (pandas, NumPy, SciPy, scikit-learn) and advanced SQL (window functions, CTEs, query tuning) for data analysis
  • 2+ years of hands-on work on a leading cloud data platform such as Databricks, Azure, AWS, or GCP
  • Experience across the end-to-end data science workflow, from finding and assessing datasets to production deployment, including experiment tracking and model management with MLflow or similar
  • Solid grounding in traditional machine learning: supervised and unsupervised methods, gradient-boosted trees (XGBoost, LightGBM), model selection, cross-validation, and hyperparameter tuning
  • Strong applied statistics: hypothesis testing, regression, sampling, and experimental design
  • Experience working with large datasets in a distributed environment (Spark/PySpark)
  • Sound model evaluation practice: picking the right metrics, handling class imbalance, avoiding leakage, and explaining model behavior (for example SHAP or feature importance)
  • Working knowledge of large language models (LLMs) and agentic AI workflows, including prompt design and RAG patterns
  • Version control with Git and collaborative development practices (code review, branching, testing)
  • Ability to explain technical work to non-technical stakeholders and turn ambiguous requests into defined deliverables
  • US Citizen: Must be a citizen of the United States
  • Security Clearance: Must be able to obtain a public trust clearance.  Must be eligible to work in the United States.

Desired

  • 5+ years as a data scientist, shipping multiple products that run in operation
  • Working experience with Databricks in Azure, including Unity Catalog, Delta Lake, Databricks Jobs, and Databricks notebooks/Repos
  • MLOps experience across the lifecycle, such as:
    • Model Serving endpoints for real-time inference
    • Feature engineering and feature tables governed in Unity Catalog
    • CI/CD for ML with Azure DevOps or GitHub Actions, and Databricks Asset Bundles
    • Production monitoring for drift and model quality (for example Lakehouse Monitoring)
    • Champion/challenger or A/B testing of models in production
    • Automated retraining and model lineage and auditability
  • Experience with Azure AI services (Azure OpenAI, Azure Machine Learning) or Mosaic AI (Vector Search, Agent Framework)
  • Prior experience shipping products that use LLMs or AI agents, including evaluation and guardrails
  • 2+ years building visual insights (dashboards and reports) that support operational needs, with Power BI, Databricks AI/BI dashboards, or Tableau
  • Prior experience with R, Stata, or SAS
  • Experience refactoring R, Stata, or SAS codebases to Python, including validating that results match
  • Experience with VA or federal healthcare data, and handling PHI/PII under federal privacy and security requirements
  • Familiarity with federal AI governance and responsible AI practices (bias testing, model documentation, human oversight)
  • Experience training or mentoring analysts and data scientists
  • Master's or PhD in a quantitative field

Benefits
  • Competitive salary
  • Generous annual leave and paid holidays
  • Comprehensive group health and dental plans
  • 401(k) with company match
  • Life insurance and AD&D coverage
  • Ongoing training and professional development opportunities

Skills Required

  • Bachelor's degree in Engineering, Computer Science, Statistics, Mathematics, Systems, Business, or a related scientific or technical discipline
  • 15+ years of experience or commensurate experience
  • Proficiency in Python, pandas, NumPy, SciPy, scikit-learn, and advanced SQL
  • At least 2 years of hands-on experience with a leading cloud data platform such as Databricks, Azure, AWS, or GCP
  • End-to-end data science experience, including dataset assessment, production deployment, experiment tracking, and model management with MLflow or similar
  • Knowledge of supervised and unsupervised machine learning, gradient-boosted trees, model selection, cross-validation, and hyperparameter tuning
  • Applied statistics experience including hypothesis testing, regression, sampling, and experimental design
  • Experience working with large datasets in distributed environments using Spark or PySpark
  • Experience with model evaluation, class imbalance, leakage prevention, and model explainability using SHAP or feature importance
  • Working knowledge of large language models, agentic AI workflows, prompt design, and retrieval-augmented generation
  • Version control and collaborative development practices using Git, code review, branching, and testing
  • Ability to explain technical work to non-technical stakeholders and define deliverables from ambiguous requests
  • United States citizenship
  • Ability to obtain a Public Trust security clearance
  • Eligibility to work in the United States
  • At least 5 years as a data scientist with multiple products operating in production
  • Experience with Databricks in Azure, including Unity Catalog, Delta Lake, Databricks Jobs, and Databricks notebooks or Repos
  • MLOps experience including model serving, governed feature tables, CI/CD, production monitoring, A/B testing, retraining, lineage, and auditability
  • Experience with Azure AI services or Mosaic AI
  • Experience shipping LLM or AI-agent products with evaluation and guardrails
  • At least 2 years building operational dashboards and reports with Power BI, Databricks AI/BI dashboards, or Tableau
  • Experience with R, Stata, or SAS
  • Experience refactoring R, Stata, or SAS codebases to Python and validating matching results
  • Experience with VA or federal healthcare data and PHI/PII privacy and security requirements
  • Familiarity with federal AI governance and responsible AI practices
  • Experience training or mentoring analysts and data scientists
  • Master's or PhD in a quantitative field
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Melbourne
4 Employees

What We Do

BlueFlag is a partnership between industry experts who believe we provide a better service to our customers and better home to our team members through cooperation. Culture Our first motivation to start BlueFlag was to create a culture that we wanted to work in. Visit blueflag.co/culture to learn more. We think you'll want to be here. Services We're really good at everything project management and talent. We've got some unique approaches to both, so please visit our website to learn more: www.blueflag.co/projects and www.blueflag.co/talent Veteran-owned We care more about doing the right thing than, well, anything. It's something we learned serving our country. We'll do the right thing by you also.

Similar Jobs

Vantor Logo Vantor

Senior Data Scientist

Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
Remote
United States
2500 Employees
170K-200K Annually

Vantor Logo Vantor

Senior Data Scientist

Aerospace • Artificial Intelligence • Computer Vision • Software • Analytics • Defense • Big Data Analytics
Remote
United States
2500 Employees
170K-200K Annually

DuckDuckGo Logo DuckDuckGo

Senior Data Scientist

Information Technology
Remote
13 Locations
393 Employees
179K-179K Annually

Automation Anywhere Logo Automation Anywhere

Senior Data Scientist

Artificial Intelligence • Cloud • Robotics • Software
Remote
USA
6564 Employees

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account