Senior Data Engineer

Posted 3 Days Ago
Raleigh, NC, USA
In-Office
158K-180K Annually
Senior level
Cloud • Information Technology • Internet of Things • Software • Consulting • Infrastructure as a Service (IaaS) • Automation
Creating better technology the open source way
The Role
Design, build, and operate high-volume data pipelines between Snowflake and Databricks using PySpark and dbt; orchestrate workflows with Airflow; administer Databricks and OpenShift environments; implement retrieval and ML systems (NER, XGBoost, Keras, forecasting); operationalize MLOps with MLflow; manage CI/CD, container builds, and OpenShift deployments; lead security/compliance assessments and remediation with SAST and vulnerability scanning.
Summary Generated by Built In

*Telecommuting role to be performed anywhere in the U.S.

Architect and implement complex, high-volume data pipelines between Snowflake and Databricks utilizing PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality tests and validation frameworks.

What You Will Do:

  • Orchestrate pipeline scheduling, dependency management, and automated failure recovery using Apache Airflow to deliver B2B marketing attribution and multi-touch targeting analytics.
  • Administer the enterprise Databricks platform by configuring IAM roles for secure Amazon S3 bucket access, managing application credentials and secrets through Databricks' built-in vault system and OpenShift secrets, and establishing workspace governance policies and cluster configurations for cross-functional data science and engineering teams.
  • Design and deploy intelligent retrieval architecture and AI-driven workflows using vector-based search methods and enterprise data platforms, building marketing retrieval and decision-automation applications that integrate multiple data sources and APIs.
  • Operationalize MLOps methodologies using MLflow for experiment tracking and model registry management, and Lakehouse monitoring for automated post-production model performance tracking to optimize predictive accuracy and increase marketing return on investment.
  • Implement end-to-end machine learning models and deliver stakeholder-facing analytical outputs by building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise systems analysis, developing predictive models using XGBoost and Scikit-learn, constructing deep learning architectures using Keras, and designing time-series forecasting models for event-based user adoption prediction.
  • Manage CI/CD pipelines using Git and Tekton to ensure reliable, repeatable code delivery for production applications.
  • Build and manage container images using buildah and skopeo, pushing to internal container registries for deployment.
  • Lead the deployment and maintenance of containerized data science models and enterprise applications on Red Hat OpenShift (Kubernetes), managing network routes, TLS termination, and container orchestration for highavailability services.
  • Lead application security initiatives by completing comprehensive enterprise security compliance assessments encompassing 20+ security controls across the full technology stack, aligned with industry frameworks such as NIST and CIS Controls.
  • Perform static application security testing (SAST) using SonarQube, execute vulnerability scanning using Qualys and pip-audit, complete Privacy Impact Assessments (PIA), and conduct STRIDE-based threat modeling.
  • Collaborate with enterprise information security teams to remediate identified vulnerabilities, navigate compliance audits, and maintain centralized logging and monitoring through Splunk.

What You Will Bring:

  • Master's degree (U.S. or foreign equivalent) in Computer Science or related field and three (3) years of experience in the job offered or related role OR Bachelor's degree (U.S. or foreign equivalent) in Computer Science or related field and five (5) years of experience in the job offered or related role.
  • Must have three (3) years of experience with: architecting and implementing high-volume data pipelines between cloud data warehouse (Snowflake) and lakehouse (Databricks) platforms using PySpark for distributed data processing and dbt for SQL-based data transformation, including developing macro-driven data quality test frameworks and validation logic; orchestrating and scheduling data pipeline workflows using Apache Airflow, including configuring DAG-based dependency management, automated failure recovery, and pipeline monitoring for enterprise analytics workloads; administering enterprise Databricks environments, including configuring IAM roles for secure cloud object storage (Amazon S3) access, managing application secrets through platform vault systems and OpenShift secrets, and establishing workspace governance and cluster policies for cross-functional teams; implementing end-to-end machine learning models by: 1) building Named Entity Recognition (NER) systems using TextBlob, gensim, and fastText for enterprise text analysis; 2) developing predictive models using gradient boosting frameworks (XGBoost) and Scikit-learn; 3) constructing deep learning architectures using Keras; and 4) designing time-series forecasting models for event-based prediction; delivering full-scale information retrieval systems for enterprise data by researching, evaluating, and implementing Transformer architectures and Transfer Learning methodologies using deep learning frameworks for semantic search, text classification, and vector-based clustering; operationalizing MLOps methodologies using MLflow for experiment tracking and model registry management, and implementing automated post-production model monitoring to track performance degradation and optimize predictive accuracy; managing CI/CD pipelines using Git and Tekton, building and publishing container images using buildah and skopeo to internal container registries, and deploying containerized applications on Red Hat OpenShift (Kubernetes) with network route management, TLS termination, and high-availability configurations; and leading enterprise security compliance assessments, including performing static application security testing (SAST) using SonarQube, executing vulnerability scans using Qualys, completing Privacy Impact Assessments (PIA), and conducting STRIDE-based threat modeling.
     

#LI-DNI

The salary range for this position is $158,309 - $180,000/year. Actual offer will be based on your qualifications.

Pay Transparency

Red Hat determines compensation based on several factors including but not limited to job location, experience, applicable skills and training, external market value, and internal pay equity. Annual salary is one component of Red Hat’s compensation package. This position may also be eligible for bonus, commission, and/or equity. For positions with Remote-US locations, the actual salary range for the position may differ based on location but will be commensurate with job duties and relevant work experience.


About Red Hat

Red Hat is the world’s leading provider of enterprise open source software solutions, using a community-powered approach to deliver high-performing Linux, cloud, container, and Kubernetes technologies. Spread across 40+ countries, our associates work flexibly across work environments, from in-office, to office-flex, to fully remote, depending on the requirements of their role. Red Hatters are encouraged to bring their best ideas, no matter their title or tenure. We're a leader in open source because of our open and inclusive environment. We hire creative, passionate people ready to contribute their ideas, help solve complex problems, and make an impact.

Inclusion at Red Hat
Red Hat’s culture is built on the open source principles of transparency, collaboration, and inclusion, where the best ideas can come from anywhere and anyone. When this is realized, it empowers people from different backgrounds, perspectives, and experiences to come together to share ideas, challenge the status quo, and drive innovation. Our aspiration is that everyone experiences this culture with equal opportunity and access, and that all voices are not only heard but also celebrated. We hope you will join our celebration, and we welcome and encourage applicants from all the beautiful dimensions that compose our global village.

Equal Opportunity Policy (EEO)
Red Hat is proud to be an equal opportunity workplace and an affirmative action employer. We review applications for employment without regard to their race, color, religion, sex, sexual orientation, gender identity, national origin, ancestry, citizenship, age, veteran status, genetic information, physical or mental disability, medical condition, marital status, or any other basis prohibited by law.


Red Hat does not seek or accept unsolicited resumes or CVs from recruitment agencies. We are not responsible for, and will not pay, any fees, commissions, or any other payment related to unsolicited resumes or CVs except as required in a written contract between Red Hat and the recruitment agency or party requesting payment of a fee.
Red Hat supports individuals with disabilities and provides reasonable accommodations to job applicants. If you need assistance completing our online job application, email [email protected]. General inquiries, such as those regarding the status of a job application, will not receive a reply. 

Skills Required

  • Master's degree in Computer Science or related field plus 3 years' experience OR Bachelor's degree plus 5 years' experience.
  • 3 years experience architecting and implementing high-volume data pipelines between Snowflake and Databricks using PySpark and dbt, including macro-driven data quality tests and validation frameworks.
  • 3 years experience orchestrating and scheduling data pipeline workflows using Apache Airflow, configuring DAG dependencies, automated failure recovery, and pipeline monitoring.
  • 3 years administering enterprise Databricks environments, configuring IAM for Amazon S3 access, managing secrets (Databricks vault/OpenShift), and establishing workspace governance and cluster policies.
  • 3 years implementing end-to-end machine learning models: NER using TextBlob/gensim/fastText, predictive models with XGBoost and scikit-learn, deep learning with Keras, and time-series forecasting.
  • 3 years delivering information retrieval systems using Transformer architectures and transfer learning for semantic search, text classification, and vector-based clustering.
  • 3 years operationalizing MLOps using MLflow for experiment tracking and model registry and implementing post-production model monitoring.
  • 3 years managing CI/CD pipelines using Git and Tekton, and building/publishing container images using buildah and skopeo to internal registries.
  • 3 years deploying and maintaining containerized applications and models on Red Hat OpenShift (Kubernetes), including network routes, TLS termination, and high-availability configurations.
  • 3 years leading enterprise security compliance assessments, performing SAST with SonarQube, vulnerability scanning with Qualys and pip-audit, completing PIAs, and conducting STRIDE threat modeling.

Red Hat Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Red Hat and has not been reviewed or approved by Red Hat.

  • Healthcare Strength Healthcare coverage is presented as comprehensive, spanning medical, dental, and vision along with life and disability coverage. Access to HSA/FSA options and broadly positive reception of health benefits support the view that healthcare is a core strength.
  • Leave & Time Off Breadth Time-off offerings are described as generous, with substantial PTO for new hires plus additional recharge days and an end-of-year shutdown for many non-critical roles. Paid volunteer time, holidays, sick days, and supportive expectations around taking time off reinforce the breadth of leave benefits.
  • Strong & Reliable Incentives The rewards package includes performance bonuses and a recurring quarterly bonus program tied to company and individual performance. Availability of ESPP participation further adds to incentive pathways beyond base pay.

Red Hat Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Raleigh, NC
20,000 Employees
Year Founded: 1993

What We Do

At Red Hat, we connect an innovative community of customers, partners, and contributors to deliver an open source stack of trusted, high-performing solutions. We offer cloud, Linux, middleware, storage, and virtualization technologies, together with award-winning global customer support, consulting, and implementation services. Red Hat is a rapidly growing company supporting more than 90% of Fortune 500 companies.

Why Work With Us

Red Hatters freely exchange different viewpoints, contribute ideas, and solve problems together. Our love of collaboration, accountability, a sense of community, and a measure of autonomy combine to create a powerful force that fosters innovation and makes Red Hat a great place to work.

Gallery

Gallery

Similar Jobs

PwC Logo PwC

Senior Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
46 Locations
370000 Employees
124K-280K Annually

PwC Logo PwC

Senior Data Engineer

Artificial Intelligence • Professional Services • Business Intelligence • Consulting • Cybersecurity • Generative AI
Hybrid
40 Locations
370000 Employees
99K-232K Annually

GRAIL Logo GRAIL

Senior Data Engineer

Artificial Intelligence • Big Data • Healthtech • Machine Learning • Software • Biotech
Hybrid
Durham, NC, USA
918 Employees
86K-106K Annually

MetLife Logo MetLife

Senior Data Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Cary, NC, USA
43000 Employees
115K-135K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account