VN Technology Data Engineer - Analysis Platform

Posted 15 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Artificial Intelligence • Other • Software • Industrial • Manufacturing
The Role
Build and operate CADDi’s data platform, including batch and streaming pipelines, BigQuery warehouse modeling, data quality controls, governance, monitoring, and analytics enablement. Collaborate with ML engineers, product managers, and analysts to improve data reliability, annotation quality, metric definitions, and machine learning validation workflows. The role involves production ETL/ELT ownership, cloud data engineering, cost and performance optimization, and platform development in a hybrid Vietnam-based environment.
Summary Generated by Built In

RECRUITMENT BACKGROUND & EXPECTED ROLE

CADDi is on a mission to "Unleash the potential of manufacturing."

We operate "CADDi DRAWER," a cloud-based system that supports digital transformation centered on the use of drawings, which are the most essential data in the manufacturing industry.

One of the issues we are currently facing in the development of this rapidly growing product is developer productivity. The inefficiencies in the development environment have become more noticeable, and they are becoming a hindrance to the development of features and a degradation of the developer experience. As an activity to resolve these issues, the need for Platform Engineering is increasing.

There are various interpretations and approaches to Platform Engineering in the world, but here, we will be working on providing tools and infrastructure to increase the development velocity by “separating concerns.”

EXAMPLES OF ANTICIPATED TASKS

You will primarily work on the following (not limited to):

Data pipelines & warehouse

  • Build and operate batch and streaming pipelines that bring data from our databases, event streams, and SaaS tools into BigQuery.
  • Model the warehouse so business concepts are defined once and reused everywhere.
  • Keep pipelines and queries fast, reliable, and cost-efficient.

Data quality & governance

  • Set up data contracts, automated testing, and freshness monitoring so breakages are caught before stakeholders find them.
  • Establish ownership, documentation, and access control for sensitive customer data.
  • Improve data and annotation quality to directly raise ML model performance, together with the ML team.

Analytics enablement

  • Work with PMs and analysts to turn business questions into datasets they can self-serve.
  • Own the core metric definitions and event taxonomy so everyone reports the same number.
  • Shorten the time between "we have a hypothesis" and "we have the data to validate it."

* Besides the team we are recruiting for this time, you may be assigned to other teams depending on your experience and preferences. (In that case, we would be happy to discuss this with you at the interview.)

* After joining the company, your role may change due to organizational growth or an individual’s career perspective.

INTEREST AND EXPERIENCE GAINED FROM THIS POSITION

  • Ownership of a data platform from the ground up, rather than maintaining someone else’s — your design decisions will shape how the company reads its own product for years.
  • Close collaboration with ML engineers and product management members, allowing you to expand your responsibilities depending on your interests and initiative.
  • The fun of integrating a complex domain into a system.
  • Experience in solving difficult problems with highly motivated team members.
  • Experience in contributing to the scale of a product with technical skills.
  • Experience in developing products that are deployed globally.
  • Experience in providing value to society through the development of products that change an industry.

ORGANIZATION

Currently, the Analysis Platform team focuses on supporting modeling, ML serving, and operational workflows. Looking ahead, we aim not only to reduce the cognitive and operational burden around machine learning, but also to build a robust foundation for ML verification — ultimately shortening the lead time needed to validate business value.

To achieve this, the Analysis Platform team will expand its scope beyond traditional MLOps into data engineering proper: reliable pipelines, a well-governed warehouse, improved annotation efficiency, and mechanisms that enable faster and more reliable value validation. This role sits at the center of that expansion, working alongside ML engineers, platform engineers, and product teams.


Requirements

MUST-HAVE REQUIREMENTS

  • 5+ years of professional experience as a Data Engineer, or as a software engineer whose work was primarily building data pipelines and data platforms.
  • Strong SQL — able to write, debug, and optimize complex analytical queries, and to model data for analytical workloads (not just query existing tables).
  • Proficiency in Python for data processing, pipeline development, and automation.
  • Hands-on experience designing and operating ETL/ELT pipelines in production, including orchestration, scheduling, backfills, and failure handling (e.g. Airflow, dbt, Argo Workflows, Dagster, Spark, or equivalent).
  • Experience with a cloud data warehouse or large-scale data processing platform (BigQuery, Redshift, Snowflake, Databricks, Spark/Hadoop, or equivalent), including an understanding of cost and performance trade-offs.
  • Experience in development using public cloud platforms such as Google Cloud, AWS, etc.
  • A sense of ownership over data correctness — you treat a broken dashboard or a silently wrong number as your problem, and you build the checks that prevent it next time.
  • Fluent business communication skills in English, able to complete daily tasks in English, including text communication and meetings (CEFR B1 or higher).
  • Must currently reside in Vietnam or have plans to relocate. Foreign nationals must also hold a valid Vietnam work permit or be legally eligible to work in Vietnam.

NICE-TO-HAVE REQUIREMENTS

  • Experience with dbt for warehouse modeling, testing, and documentation.
  • Experience with streaming or event-driven data processing (Cloud Pub/Sub, Kafka, Apache Beam / Dataflow, Spark Structured Streaming).
  • Experience building and operating Data Lakes, Lakehouses, or Feature Stores.
  • Experience implementing initiatives to improve data quality for data-centric ML model improvement.
  • Experience with data quality / observability tooling and practices — data contracts, testing frameworks, lineage, anomaly detection.
  • Experience planning and driving data utilization initiatives — internally or externally — using tools such as BigQuery, Redash, Looker, or Metabase.
  • Hands-on experience with a statically typed language (Scala, Java/Kotlin, Go, TypeScript, Rust, etc.).
  • Experience with infrastructure as code and CI/CD (Terraform, GitHub Actions, Kubernetes).
  • Experience developing machine learning pipelines using tools such as Vertex AI Pipelines, Kubeflow, Apache Beam, or Spark.
  • Experience collaborating with ML engineers to continuously improve and deliver machine learning and data science models.
  • Familiarity with at least one ML/AI framework such as scikit-learn, PyTorch, or TensorFlow.
  • Basic knowledge of statistics, linear algebra, and core computer science concepts behind AI (vector spaces, embeddings, inference).
  • Experience leading design, development, and operations as a team lead or project driver, regardless of project size.
  • Experience working with Scrum or Agile methodologies.
  • Conversational-level Japanese proficiency (JLPT N2 or above is a guideline).

WE ARE LOOKING FOR THIS KIND OF PERSON

  • Those who can sympathize with CADDi’s mission “Unleashing the Potential of the Manufacturing Industry”.
  • Those who have a T-shaped ambition mindset to maximize their expertise by not only focusing on back-end and infrastructure, but also catching up on peripheral knowledge as needed.
  • Those who are able to face essential issues and take action to solve them with a sense of ownership.
  • Able to work through positive attitude and constructive discussions in fast-changing and uncertain situations.
  • Able to communicate and discuss with an attitude of respect for others, taking into consideration their context and resolution.

PRODUCT DEVELOPING ENVIRONMENT

  • Front-end: TypeScript, React, Next.js
  • Backend: Rust (axum), TypeScript, Node.js (Express, Fastify, NestJS), Python (FastAPI)
  • Machine Learning/Algorithms: Rust, Python, OpenCV, PyTorch, TorchServe, Elasticsearch, Vertex AI
  • Infrastructure: Google Cloud, Google Kubernetes Engine, Anthos Service Mesh, Istio, Cloudflare, Argo Workflows
  • Event Bus: Cloud Pub/Sub
  • DevOps: GitHub, GitHub Actions, ArgoCD, Kustomize, Helm, Terraform, Datadog, MixPanel, Sentry
  • Data: CloudSQL (PostgreSQL), AlloyDB, BigQuery, dbt, trocco
  • API: GraphQL, REST, gRPC
  • Authentication: Auth0
  • Development tools: GitHub Copilot, Figma, Storybook
  • Communication Tools: Slack, Discord, JIRA, Miro, Confluence

RECRUITING STEPS

  1. CV screening
  2. Technical assignment. We place more importance on whether you can imagine that you can work together with us to develop a product, rather than on your knowledge of algorithms or the speed of your answers.
  3. HR Interview (online)
  4. Technical interview (with engineer)
  5. Final interview (with CTO)
  6. Offer meeting
  • Please note that, depending on the situation, additional interviews or discussions may be proposed.
  • If desired, we can arrange casual interviews with employees even during the selection process. Please feel free to consult with us.
  • The average time from application to offer is about one month, but if you are in a hurry, please let us know. We will do our best to adjust the schedule to fit your job search timeline.

Benefits

APPLICATION GUIDELINES & BENEFITS

1. Working style:

  • Hybrid (come to Office at least once a week)

2. Office address: 

  • HCMC: 7F, Gia Loc Building, No. 27-29 Nguyen Cuu Van Street, Ward 17, Binh Thanh District, HCMC
  • Hanoi: Unit 9.03, 9F, The West Building, 265 Cau Giay Street, Cau Giay Ward, Hanoi

3. Employment type: 

  • Official full-time employee
  • Probation period: 2 months

4. Holidays and leave:

  • Annual paid leave: 12 days
  • National holidays
  • Year-end holidays (December 31 to January 3)
  • Tet holidays
  • Others (following Labor Regulations)

5. Benefits:

  • 13th month salary
  • Salary review: twice a year
  • 100% monthly basic salary and mandatory social insurances in 2-month probation
  • Premium Health Insurance
  • Social insurance, health insurance, unemployment insurance, workers’ accident compensation insurance
  • Annual health check-up
  • Allowances such as: child-care allowance, commuting allowance, life event congratulatory gift, etc
  • Growth support such as subsidy for server fee, support for attending external training courses
  • Intensive training program (external or internal training courses, workshop etc)
  • Devices: PC and display of desired specifications
  • Awards: Company awards, every 6 month MVP awards
  • Activities: Year-end-party, team building, etc

Skills Required

  • 5+ years of professional experience as a Data Engineer or software engineer primarily building data pipelines and data platforms
  • Strong SQL skills, including writing, debugging, optimizing analytical queries, and modeling data for analytical workloads
  • Proficiency in Python for data processing, pipeline development, and automation
  • Production experience designing and operating ETL/ELT pipelines, including orchestration, scheduling, backfills, and failure handling
  • Experience with Airflow, dbt, Argo Workflows, Dagster, Spark, or equivalent pipeline technologies
  • Experience with BigQuery, Redshift, Snowflake, Databricks, Spark/Hadoop, or an equivalent cloud data warehouse or large-scale processing platform
  • Experience developing with public cloud platforms such as Google Cloud or AWS
  • Strong ownership of data correctness and data quality
  • Fluent business communication in English, including written communication and meetings; CEFR B1 or higher
  • Currently residing in Vietnam or planning to relocate; foreign nationals must hold or be legally eligible for a Vietnam work permit
  • Experience with dbt for warehouse modeling, testing, and documentation
  • Experience with streaming or event-driven data processing, such as Cloud Pub/Sub, Kafka, Apache Beam, Dataflow, or Spark Structured Streaming
  • Experience building and operating data lakes, lakehouses, or feature stores
  • Experience improving data quality for data-centric machine learning models
  • Experience with data quality or observability practices, including data contracts, testing, lineage, or anomaly detection
  • Experience planning and driving data utilization initiatives using BigQuery, Redash, Looker, or Metabase
  • Hands-on experience with a statically typed language such as Scala, Java, Kotlin, Go, TypeScript, or Rust
  • Experience with infrastructure as code and CI/CD using Terraform, GitHub Actions, or Kubernetes
  • Experience developing machine learning pipelines using Vertex AI Pipelines, Kubeflow, Apache Beam, or Spark
  • Experience collaborating with ML engineers to improve and deliver machine learning or data science models
  • Familiarity with scikit-learn, PyTorch, TensorFlow, or another ML/AI framework
  • Basic knowledge of statistics, linear algebra, and AI-related computer science concepts
  • Experience leading design, development, and operations as a team lead or project driver
  • Experience working with Scrum or Agile methodologies
  • Conversational Japanese proficiency, approximately JLPT N2 or above

CADDi Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about CADDi and has not been reviewed or approved by CADDi.

  • Healthcare Strength — Employer-paid medical, dental, and vision coverage is described as being fully covered for employees in some U.S. postings, which stands out as a strong core benefit. The package is also described as comprehensive across health and wellness, including mental health benefits in some listings.
  • Retirement Support — A 401(k) with company match is repeatedly listed, including references to a day-one match and a specific match percentage in at least one posting. This suggests meaningful support for long-term savings compared with many early-stage employers.
  • Equity Value & Accessibility — Stock options or equity are frequently included as a standard component of total rewards, indicating participation in company upside is part of the package. Some descriptions frame equity alongside structured reviews, implying an emphasis on longer-term value rather than only cash pay.

CADDi Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Tokyo
367 Employees
Year Founded: 2017

What We Do

CADDi is a global supply chain company on a mission to "unleash the potential of manufacturing". The company strives to transform the manufacturing industry through its primary offering "CADDi Manufacturing", a one-stop service for procurement and manufacturing that utilizes original technologies to optimize quality, cost, and delivery within its supply chain infrastructure. In mid-2022, CADDi launched "CADDi Drawer," a cloud-based data utilization system to further digital transformation in the manufacturing industry.

Similar Jobs

Takeda Logo Takeda

Account Executive

Healthtech • Software • Analytics • Biotech • Pharmaceutical • Manufacturing
Remote or Hybrid
VNM
50000 Employees

Takeda Logo Takeda

Account Executive

Healthtech • Software • Analytics • Biotech • Pharmaceutical • Manufacturing
Remote or Hybrid
VNM
50000 Employees

Mastercard Logo Mastercard

Manager, Account Management, MCDS Vietnam

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Hà Nội, VNM
38800 Employees

Mastercard Logo Mastercard

Solutions Architect

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
Hà Nội, VNM
38800 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account