Database Administrator

Posted Yesterday
Be an Early Applicant
Hiring Remotely in Reston, VA, USA
In-Office or Remote
Senior level
Information Technology • Consulting
The Role
Designs and maintains ETL pipelines for IRS data warehousing using SQL, Python, shell scripting, Sybase IQ, and PostgreSQL. Applies statistical and machine learning methods to clean, analyze, and model large datasets. Responsibilities include data quality, pipeline optimization, troubleshooting, documentation, metadata management, stakeholder collaboration, and deployment in Linux, Windows, OpenShift, and Kubernetes environments.
Summary Generated by Built In

Brillient Corporation is seeking a Data Scientist / ETL Engineer to support a mission-critical IRS data warehousing initiative. In this hybrid role, you will design, build, and maintain ETL processes that move data into the Compliance Data Warehouse (CDW). You will also apply statistical and machine learning techniques to that data to produce predictive models and actionable insights. You'll work closely with data architects, business analysts, data quality specialists, and mission stakeholders to make sure data is accurate, consistent, and analytically useful.

This role supports CDW operations within RAAS. That work includes data analysis, process improvement recommendations, and research and evaluation of emerging technologies to improve RAAS data availability, analytics, and value to stakeholders. The CDW is a non-IT data warehouse containing all of the IRS's return, entity, information return, regulatory, enforcement, web, and security data. The team provides technical guidance and executes work across ETL, database administration, SQL and Bash development, SAS administration, system security, COTS ETL development, metadata, data quality, customer service, web design, and intergovernmental data exchanges.

Key Responsibilities

ETL Design and Development

  • Extract data and tables from Unix/Linux systems, and transform and load them into Sybase IQ and Postgres.
  • Refactor existing ETL jobs and build new solutions as needed. This includes building an object-oriented, Unix-based framework of scripts and stored procedures to run and monitor multiple ETL and statistics processes.
  • Develop and run shell, SQL, and Python scripts to support data processing, automation, and analytical workflows.
  • Build repeatable, scalable data pipelines and analytical processes.
  • Troubleshoot errors in shell and SQL, and configure and operate SSH clients across servers.

Data Science and Analytics

  • Prepare, clean, transform, and explore data, and engineer features, using Pandas and NumPy.
  • Analyze large, complex datasets and tables to find trends, patterns, relationships, and actionable insights.
  • Develop, implement, evaluate, and maintain statistical and machine learning models with scikit-learn or comparable frameworks.
  • Validate model results, assess performance, and find ways to improve accuracy and reliability.
  • Conduct exploratory and statistical analysis to support data-driven decision-making.
  • Develop and document analytical workflows and models in Jupyter Notebook.
  • Work with structured and unstructured data from multiple sources and formats.

Optimization, Performance, and Data Quality

  • Optimize ETL processes for efficient, timely data processing.
  • Monitor ETL jobs and resolve performance bottlenecks and failures promptly.
  • Keep data accurate and intact through validation, cleansing, and auditing.
  • Work with the Data Quality team to resolve data inconsistencies and issues.

Collaboration and Documentation

  • Work with Data Architects to design data models and schemas that meet business needs.
  • Translate business and mission requirements from analysts, SMEs, and stakeholders into ETL and analytical solutions.
  • Document ETL processes, workflows, data dictionaries, methodologies, models, assumptions, and results.
  • Help users move data into Sybase IQ and provide metadata for the metadata repository.
  • Present complex technical findings clearly to both technical and non-technical audiences.
  • Take part in Scrum and client meetings, and keep Kanban board cards up to date.

Environment and Continuous Improvement

  • Work in Linux/Unix and Windows environments, including containerized and distributed platforms (OpenShift, Kubernetes).
  • Keep up with emerging ETL, data science, ML, and AI technologies, and recommend improvements to processes and infrastructure.
QualificationsRequired Education and Qualifications
  • Bachelor's degree in Data Science, Computer Science, Statistics, Mathematics, Engineering, Information Systems, or a related technical field.
  • US citizenship.
  • Ability to obtain and maintain a Public Trust security clearance.
  • At least eight (8) years of combined professional experience in ETL development and data science, analytics, or machine learning.
  • Strong SQL ETL experience, focused on load, select, and update commands.
  • High proficiency in Linux/Unix, Sybase IQ, and Sybase SQL scripting.
  • Working experience with Postgres and Postgres SQL scripting.
  • Strong hands-on Python experience for ETL, data analysis, and machine learning, including Pandas, NumPy, and Jupyter Notebook.
  • Demonstrated experience building and evaluating ML models with scikit-learn or comparable frameworks.
  • Experience with data cleaning, transformation, EDA, feature engineering, and model evaluation.
  • Shell scripting for automation and data processing (KSH, Bash, SSH, sed, awk).
  • Experience working in Linux and Windows environments.
  • Experience with OpenShift and/or Kubernetes.
  • Strong communication skills, and the ability to work both independently and collaboratively in a fast-paced environment.
Preferred Qualifications
  • Experience with AI, LLMs, Retrieval-Augmented Generation (RAG), or model fine-tuning.
  • Experience with Apache Airflow and automated ETL pipelines.
  • Experience with SAP Data Services or other COTS ETL tools.
  • Experience with Git/GitLab, DBeaver, VS Code, SecureCRT/SecureFX, Wiki, JSON, and Perl.
  • Experience deploying ML models in containerized or cloud environments.
  • Experience with MLOps: model deployment, monitoring, and lifecycle management.

DISCLAIMER: The above statements are intended to describe the general nature and level of work performed. They are not intended to be an exhaustive list of all responsibilities, duties, skills, efforts, requirements, or working conditions. Management reserves the right to revise the job or to require that other or different tasks be performed as assigned in accordance with business demands and/or contractual requirements.

Brillient is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to any status protected under applicable federal, state, or local law.

Skills Required

  • Bachelor's degree in Data Science, Computer Science, Statistics, Mathematics, Engineering, Information Systems, or a related technical field
  • US citizenship
  • Ability to obtain and maintain a Public Trust security clearance
  • At least eight years of combined professional experience in ETL development and data science, analytics, or machine learning
  • Strong SQL ETL experience involving load, select, and update commands
  • High proficiency with Linux or Unix, Sybase IQ, and Sybase SQL scripting
  • Working experience with PostgreSQL and PostgreSQL scripting
  • Hands-on Python experience for ETL, data analysis, and machine learning, including Pandas, NumPy, and Jupyter Notebook
  • Experience building and evaluating machine learning models with scikit-learn or comparable frameworks
  • Experience with data cleaning, transformation, exploratory data analysis, feature engineering, and model evaluation
  • Shell scripting experience using KSH, Bash, SSH, sed, and awk
  • Experience working in Linux and Windows environments
  • Experience with OpenShift and/or Kubernetes
  • Strong communication skills and ability to work independently and collaboratively
  • Experience with AI, LLMs, Retrieval-Augmented Generation, or model fine-tuning
  • Experience with Apache Airflow and automated ETL pipelines
  • Experience with SAP Data Services or other commercial ETL tools
  • Experience with Git or GitLab, DBeaver, VS Code, SecureCRT, SecureFX, Wiki, JSON, and Perl
  • Experience deploying machine learning models in containerized or cloud environments
  • Experience with MLOps, including model deployment, monitoring, and lifecycle management
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Reston, VA
404 Employees
Year Founded: 2006

What We Do

Brillient is an award-winning Full Spectrum Digital Transformation company enabling clients to transform through the continuum of analog, to digital, to analytics leading to insight-driven decision making and mission execution. We help clients achieve better efficiencies and lower costs in their digital government and IT modernization initiatives enabling friction-free interaction with citizens and business. We’ve worked with more than 22 federal agency clients over the past decade to increase efficiencies, combat waste, reduce costs, and improve citizen/customer satisfaction. In 2017 and 2019, Brillient was awarded Small Business of the Year by the Department of Homeland Security (DHS). Brillient has experienced stupendous growth making the Inc. 5000 list of fastest-growing private businesses in America, six times. Brillient's commitment to exemplary quality is evidenced by our CMMI Level 3, ISO 9001/20000/27001 certifications.

Similar Jobs

Horizon Industries Logo Horizon Industries

Database Administrator

Information Technology • Security • Business Intelligence • Consulting
Remote
USA
127 Employees

MetroStar Logo MetroStar

Database Administrator

Information Technology • Consulting
Remote
USA
250 Employees
100K-120K Annually
Remote
US
36 Employees

General Dynamics Information Technology Logo General Dynamics Information Technology

Database Administrator

Aerospace • Information Technology • Professional Services • Security • Software
Remote
United States
21625 Employees
98K-132K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account