Data Engineer

Posted 3 Days Ago
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Junior
Information Technology • Software
The Role
Build and maintain batch and streaming data pipelines, backend services, REST APIs, Airflow workflows, Spark jobs, and SQL transformations. Own monitoring, data quality, infrastructure stability, production debugging, documentation, and testing. Contribute to identity graph development for fraud detection and collaborate on data models, schemas, and real-time decisioning systems. The role requires strong programming, SQL, database, distributed systems, cloud, and software engineering fundamentals.
Summary Generated by Built In
About Bureau

Bureau is a unified risk decisioning platform for Compliance, Fraud, and Transaction risks. Our platform is a single decision-making engine, powered by a 1 billion+ identity knowledge graph. Over 150 Banks, fintechs, retailers, and digital platforms use Bureau to verify identities faster and stop fraud earlier globally.

Bureau has raised $50M+ from renowned Silicon Valley and global investors including Sorenson Capital and PayPal Ventures and is expanding rapidly from APAC to Americas, Europe, and beyond.

Why Bureau?

Bureau is building the infrastructure that makes digital identities and transactions safe and trustworthy for billions of people. The mission is big, the problems are complex, and the impact is real.

We hire people who want that level of responsibility. People who move fast, build systems from scratch, and care deeply about turning strategy into execution. If you want predictability or narrow scope, this won't be your place. If you want to shape how a scaling global company operates—keep reading.

What You'll Do

  • Build and maintain batch and streaming pipelines that ingest, clean, and transform data from device SDKs, internal services, partner APIs, and third-party data providers

  • Write and own backend services and RESTful APIs that serve data to internal teams and to real-time decisioning paths

  • Develop Airflow DAGs powering reporting, compliance analytics, model training, and feature pipelines, and keep them healthy day to day

  • Write Spark jobs and SQL transformations against our data lake, and tune them when they get slow or expensive

  • Add monitoring, alerting, and data quality checks to the pipelines you own so problems surface before a stakeholder notices

  • Debug production issues across the stack: a Kafka consumer lagging, a schema change breaking downstream, a query that got 10x slower after a data volume jump

  • Work with the infrastructure your pipelines and services run on — containers, deployments, cluster configs — and help keep it stable and cost-sane

  • Contribute to our identity graph work, helping model relationships between entities to surface fraud rings and hidden linkages

  • Write documentation and tests, participate in design reviews, and help keep our schemas and contracts sane as the system grows

What You'll Bring

Must have

  • 1–3 years of professional software engineering experience, with meaningful exposure to data-intensive systems

  • Strong programming skills in Python, Java, or Scala, and the ability to write production-quality, tested code

  • Strong SQL: joins, window functions, aggregations, and enough of a mental model of query execution to know why something is slow

  • Working understanding of databases, including the difference between OLTP and OLAP systems and when each is appropriate

  • Hands-on experience with at least one distributed data processing framework (Spark preferred) or a genuine willingness to ramp up quickly

  • Experience building or maintaining backend services and REST APIs

  • Familiarity with a major cloud platform (AWS preferred) and core services like S3, EC2, and managed databases

  • Solid computer science fundamentals: data structures, concurrency, and the basics of distributed systems

  • Comfort with Git, code review, and CI/CD

Nice to have

  • Infrastructure and systems knowledge — Docker, Kubernetes, Terraform or similar IaC, and a working sense of how services get deployed, scaled, and monitored in production

  • Experience running or tuning distributed workloads: cluster sizing, resource tuning, or tracking down a bottleneck between compute and storage

  • Exposure to Kafka or MSK, or any streaming/event-driven system

  • Experience with Airflow or a similar orchestration tool

  • Exposure to EMR, Athena, ClickHouse, Databricks, or Snowflake

  • Awareness of lakehouse table formats such as Iceberg, Delta Lake, or Hudi

  • Familiarity with observability tooling — Prometheus, Grafana, Datadog, or equivalent

  • Any experience with graph databases (Neo4j, TigerGraph, Amazon Neptune)

  • Interest in or exposure to fraud, risk, identity, fintech, or other high-stakes real-time domains

  • Exposure to ML workflows: feature pipelines, training data preparation, or model serving

What We Look For

  • You debug rather than guess, and you can explain what actually went wrong

  • You ask what the data is for before deciding how to model it

  • You're curious about the layer below the one you work in — how your code actually runs, where it fails, what it costs

  • You're comfortable being new to a tool and getting productive in it quickly

  • You care that numbers are right, because at Bureau a wrong number is a wrong risk decision

Our Culture
  • We hire self-motivated people and get out of their way

  • We value performance, not hours worked

  • Speed, ownership, and impact matter most

Compensation
  • Competitive salary + potential equity

  • Health benefits, flexible PTO, learning budget

Skills Required

  • 1-3 years of professional software engineering experience with meaningful exposure to data-intensive systems
  • Strong programming skills in Python, Java, or Scala
  • Ability to write production-quality, tested code
  • Strong SQL skills, including joins, window functions, aggregations, and query execution concepts
  • Working knowledge of OLTP and OLAP databases and their appropriate use cases
  • Hands-on experience with a distributed data processing framework, preferably Spark, or willingness to ramp up quickly
  • Experience building or maintaining backend services and REST APIs
  • Familiarity with a major cloud platform, preferably AWS, including S3, EC2, and managed databases
  • Computer science fundamentals including data structures, concurrency, and distributed systems
  • Comfort with Git, code review, and CI/CD
  • Infrastructure and systems knowledge including Docker, Kubernetes, Terraform, or similar infrastructure as code
  • Experience running or tuning distributed workloads
  • Exposure to Kafka, MSK, or another streaming/event-driven system
  • Experience with Airflow or a similar orchestration tool
  • Exposure to EMR, Athena, ClickHouse, Databricks, or Snowflake
  • Awareness of Iceberg, Delta Lake, Hudi, or similar lakehouse table formats
  • Familiarity with Prometheus, Grafana, Datadog, or equivalent observability tooling
  • Experience with graph databases such as Neo4j, TigerGraph, or Amazon Neptune
  • Experience or interest in fraud, risk, identity, fintech, or other high-stakes real-time domains
  • Exposure to machine learning workflows, including feature pipelines, training data preparation, or model serving
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
187 Employees
Year Founded: 2020

What We Do

Bureau is a no-code, identity decisioning platform that offers businesses the complete range of risk, compliance and ongoing fraud monitoring solutions innovated with AI. Using next-generation orchestration, Bureau provides businesses contextualized insights that lead to absolute decisions about digital identity trustworthiness. That makes it easy for consumers to transact and prevents bad actors from slipping by. Bureau helps businesses achieve growth, operate efficiently, and curtail fraud attacks as well as regulatory risks, ultimately saving businesses millions of dollars. The company is ISO 27001:2013 certified.

Similar Jobs

Accuris Logo Accuris

Data Engineer

Information Technology • Machine Learning • Software • Conversational AI • Generative AI • Manufacturing
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
1000 Employees

CrowdStrike Logo CrowdStrike

Data Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees

JPMorganChase Logo JPMorganChase

Data Engineer

Financial Services
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
289097 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Data Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
85422 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account