AI Data Engineer

Posted 4 Days Ago
Hiring Remotely in Sítio Israel, Ibiúna, São Paulo, BRA
Remote
Mid level
Big Data • Security • Cybersecurity
I can't believe it's not SIEM
The Role
Build and operate scalable batch and streaming data pipelines powering AI systems. Own ingestion, storage, transformation, serving, monitoring, validation, and continuous improvement across the data lifecycle. Partner with AI engineers and researchers to translate model requirements into reliable production infrastructure, using data lakes, warehouses, relational, graph, and vector databases. Define data strategy and ensure high-quality, observable, timely data for AI models and agents.
Summary Generated by Built In
Description

Vega is one of the fastest-growing startups in cybersecurity, redefining security analytics and operations with an AI-native platform for the SOC. We are building the next-generation operating system for security teams. Vega is already delivering real impact at some of the world’s largest organizations - improving detection, unlocking the value of their security data, and reducing cost and complexity. With HQs in New York and TLV, we're looking for people who want to be a part of the next rocket-ship in cyber.

We’re looking for a talented Data Engineer to join our AI team and build the data foundations that power our AI systems at scale. As a key member of the AI team, you’ll be responsible for designing, building, and operating the data layer that enables our AI models and agents to perform in production. This includes data ingestion, storage, transformation, monitoring, and defining the data strategy that ensures our AI applications receive high-quality, reliable, and timely data.

We’re seeking an ambitious Data Engineer who is passionate about building robust data infrastructure, working with large-scale and complex datasets, and partnering closely with AI engineers and researchers to translate model needs into production-grade data systems.

WHAT YOU WILL DO

  • Design, build, and operate scalable data pipelines (batch + streaming) that power Vega’s AI systems in production.
  • Own the end-to-end data lifecycle: ingestion, storage, transformation, serving, and continuous improvement.
  • Build and maintain the data layer from raw data to semantically accessible data that enables AI models and agents to perform reliably at scale.
  • Ensure high data quality and high-standard operation through monitoring, alerting, and validation checks.
  • Work with modern storage systems (data lakes, warehouses, relational, graph, and vector databases) to support diverse AI workloads.
  • Partner closely with AI engineers and researchers to translate model requirements into production-grade data infrastructure.
  • Define and evolve the data strategy to for AI applications.
Requirements

WHAT YOU WILL BRING

  • 4+ years of professional experience in data engineering or ML infrastructure roles.
  • Strong experience designing and building data pipelines (batch and streaming) for large-scale production systems.
  • Hands-on experience with data storage systems such as data lakes, data warehouses, relational databases and graph databases.
  • Proven ability to build reliable, observable, and scalable data infrastructure, including monitoring, alerting, and data quality checks.
  • Demonstrated ownership across the full data lifecycle from ingestion and modeling, to serving, monitoring, and continuous improvement.
  • Ability to work independently, managing priorities effectively in a fast-paced, product-driven environment, according to a dynamic data strategy.
  • Experience with modern data and infrastructure technologies such as DuckDB, dbt, Temporal, Trino, Spark, PostgreSQL, PGVector Neo4j, Datadog, Python, Go, Docker, and Kubernetes.
  • Experience working with ML models and framework as part of data pipelines (e.g. text embedding models, vector databases, semantic search algorithms)

NICE TO HAVE

  • Experience building data infrastructure to support ML/AI systems, including feature extraction pipelines for downstream ML models and inference-time data access.
  • Background in working with high-volume or complex data sources such as logs, events, telemetry, or security data.
  • Familiarity with modern cloud platforms, preferably AWS, and cloud-native data tools.
  • Experience collaborating closely with ML/AI engineers to translate model and research requirements into scalable data solutions.

Skills Required

  • 4+ years of professional experience in data engineering or ML infrastructure roles
  • Experience designing and building batch and streaming data pipelines for large-scale production systems
  • Hands-on experience with data lakes, data warehouses, relational databases, and graph databases
  • Ability to build reliable, observable, and scalable data infrastructure with monitoring, alerting, and data quality checks
  • Experience owning the full data lifecycle from ingestion and modeling through serving, monitoring, and continuous improvement
  • Ability to work independently and manage priorities in a fast-paced, product-driven environment
  • Experience with DuckDB, dbt, Temporal, Trino, Spark, PostgreSQL, PGVector, Neo4j, Datadog, Python, Go, Docker, and Kubernetes
  • Experience working with ML models and frameworks in data pipelines, including embeddings, vector databases, or semantic search
  • Experience building ML/AI data infrastructure, including feature extraction pipelines and inference-time data access
  • Experience with high-volume or complex data such as logs, events, telemetry, or security data
  • Familiarity with AWS and cloud-native data tools
  • Experience collaborating with ML/AI engineers to translate model and research requirements into scalable data solutions

Vega (vega.io) Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Vega (vega.io) and has not been reviewed or approved by Vega (vega.io).

  • Fair & Transparent Compensation Publicly available information indicates the company is early-stage and well-funded, which can support the ability to offer competitive packages, but no direct pay-satisfaction content is provided.
  • Equity Value & Accessibility The data repeatedly frames compensation at this stage as a mix of cash and equity/tokens, implying equity could be a meaningful component, though terms and employee outcomes are not disclosed.

Vega (vega.io) Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, New York
134 Employees
Year Founded: 2024

What We Do

We're redefining the boundaries of Security Operations by eliminating the limits and compromises of the past. Founded in 2024, Vega is on a mission to help organizations harness the power of all of their data. Wherever it is. Whatever it is. Without any of the taxes that have plagued SIEM and Data Lakes for the past 20 years. Backed by Cyberstarts, Accel, Redpoint and CRV, Vega offers a lightweight Security Analytics fabric that introduces a new, AI-native, approach to interacting with security data wherever it sits, giving analysts complete visibility and detection coverage, without a single migration, replacement or compromise.

Similar Jobs

Rimini Street Logo Rimini Street

Data Engineer

Information Technology • Software
Remote
Brazil
1600 Employees

RevStar Consulting Logo RevStar Consulting

Artificial Intelligence Engineer

Information Technology • Business Intelligence • Consulting
Remote
6 Locations
53 Employees

Power Digital Marketing Logo Power Digital Marketing

Data Engineer

Agency • Marketing Tech
Remote
Brazil
775 Employees

Blue Orange Digital Logo Blue Orange Digital

Senior Data & LLM Engineer (AI‑Ready Data & Agentic Systems)

Artificial Intelligence • Machine Learning • Database
In-Office or Remote
8 Locations
75 Employees

Similar Companies Hiring

Credal.ai Thumbnail
Software • Security • Productivity • Machine Learning • Artificial Intelligence
Brooklyn, NY
Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account