Data Engineer

Posted 4 Days Ago
Hiring Remotely in New York City, NY, USA
In-Office or Remote
140K-180K Annually
Senior level
Artificial Intelligence • Cybersecurity
The Role
Build and operate large-scale data pipelines (batch and streaming) for multi-terabyte audio/video datasets. Deploy and scale containerized workloads on Kubernetes/AWS, maintain Spark/Ray jobs, run workflow orchestration (Airflow), tune enterprise databases, and collaborate with ML teams on training and MLOps workflows.
Summary Generated by Built In
Who we are.

Reality Defender is an award-winning cybersecurity company helping enterprises and governments detect deepfakes and AI-generated media. Utilizing a patented multi-model approach, Reality Defender is robust against the bleeding edge of generative platforms producing video, audio, imagery, and text media. Reality Defender's API-first deepfake detection platform empowers teams and developers alike to identify fraud, disinformation campaigns, and harmful deepfakes in real time.

Backed by world class investors including DCVC, Illuminate Financial, Y Combinator, Booz Allen Hamilton, IBM, Accenture, Rackhouse, and Argon VC, Reality Defender works with leading enterprise clients, financial institutions, and governments in order to ensure AI-generated media is not used for malicious purposes.

Youtube: Reality Defender Wins RSA Most Innovative Startup

The Data Engineer Role.

We're looking for a Data Engineer who can build and scale the infrastructure powering our data platform, with a strong foundation in distributed systems and cloud-native tooling. You'll design and operate the pipelines that move, process, and prepare multi-terabyte and streaming datasets — including audio and video — for Reality Defender's detection models, and you'll work closely with ML engineers and researchers to keep that data flowing reliably at scale. Responsibilities include:

  • Design, build, and operate large-scale data processing pipelines handling multi-terabyte and streaming datasets, including audio/video transcoding, feature extraction, and preprocessing workflows.

  • Deploy, scale, and troubleshoot containerized workloads on Kubernetes and AWS in production environments.

  • Build and maintain distributed data processing jobs using frameworks such as Spark and Ray.

  • Design and operate workflow orchestration systems (e.g., Airflow) with dependency management, retries, monitoring, and alerting for production pipelines.

  • Administer and tune enterprise databases, including performance tuning, backup/recovery, access control, and scaling strategies.

  • Partner with ML engineers and researchers to support training pipelines, model retraining triggers, feature stores, and other MLOps workflows.

Who you are.
  • Hands-on experience with Kubernetes and AWS, including deploying, scaling, and troubleshooting containerized workloads in production environments.

  • Proficiency with high-performance/distributed computing frameworks such as Spark and Ray for processing large-scale datasets.

  • Experience with workflow orchestration tools such as Airflow (or comparable systems like Dagster, Prefect, or Luigi) to schedule and manage complex data pipelines.

  • Strong programming skills in Python and SQL; experience with Golang is a plus.

  • Demonstrated track record building and operating large-scale data processing pipelines, ideally handling multi-terabyte or streaming datasets.

  • Experience working with audio or video data at scale is a strong plus (e.g., transcoding, feature extraction, or preprocessing pipelines).

  • Familiarity with common data transformation patterns applied to large datasets (ETL/ELT, batch and stream processing, data validation and quality checks).

  • Experience designing and maintaining job orchestration systems, including dependency management, retries, monitoring, and alerting for production pipelines.

  • Bonus: experience orchestrating machine learning workflows (training pipelines, model retraining triggers, feature stores, or MLOps tooling).

What we offer.

Reality Defender offers the following benefits to all our employees, regardless of location:

  • Healthcare plans with 100% premium coverage for employees and partial coverage available for dependents

  • Dental and Vision plans with 100% premium coverage for employees and their dependents

  • Short/Long-term disability and life insurance plans with 100% premium coverage for employees

  • FSA/HSA and 401k programs

  • Equity compensation

  • 20 days of PTO per year

  • 12 weeks of Parental Leave

  • Learning and Development budget

  • Monthly wellness benefits

  • Annual company-sponsored offsite

For employees working from Reality Defender’s HQ in NYC, we offer the following benefits:

  • Daily in-office lunch through UberEats

  • Commuter benefits

  • Remote Fridays

  • Happy Hours and other local events

Skills Required

  • Hands-on experience with Kubernetes
  • Hands-on experience with AWS
  • Proficiency with Spark for large-scale data processing
  • Proficiency with Ray for distributed processing
  • Experience with workflow orchestration tools (Airflow, Dagster, Prefect, or Luigi)
  • Strong programming skills in Python
  • Strong SQL skills
  • Experience building and operating large-scale batch and streaming data pipelines
  • Administering and tuning enterprise databases (performance tuning, backup/recovery, access control, scaling)
  • Experience with audio or video data processing at scale (transcoding, feature extraction, preprocessing)
  • Familiarity with ETL/ELT, batch and stream processing, and data validation/quality checks
  • Experience designing and maintaining job orchestration systems with retries, monitoring, and alerting
  • Golang experience
  • Experience orchestrating ML workflows, model retraining triggers, or feature stores
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
30 Employees
Year Founded: 2018

What We Do

Stop deepfakes before they become a problem with Reality Defender’s proactive AI-generated media detection platform. Detect dangerous AI-generated and manipulated content across audio, video, images, and text with our enterprise-grade API and web app.

Similar Jobs

Applied Systems Logo Applied Systems

Data Engineer

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Remote or Hybrid
United States
3079 Employees
70K-120K Annually

Dragos Logo Dragos

Data Engineer

Security • Cybersecurity
Remote
United States
295 Employees
225K-225K Annually

Coinbase Logo Coinbase

Senior Software Engineer

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
USA
4700 Employees
186K-219K Annually

Collectors Logo Collectors

Staff Data Engineer

Consumer Web • eCommerce • Machine Learning • Software • Sports • Analytics
In-Office or Remote
2 Locations
2246 Employees
165K-250K Annually

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account