Data Foundry Engineer

Posted 14 Days Ago
Be an Early Applicant
Hiring Remotely in São Paulo, BRA
In-Office or Remote
Mid level
Artificial Intelligence • Hardware • Internet of Things • Industrial
The Role
Build and maintain data enrichment pipelines, curate and label datasets, resolve data quality issues, prototype APIs/tools, collaborate with ML/product teams, and document reproducible workflows for large, messy industrial data.
Summary Generated by Built In
Why join us
TRACTIAN is transforming the industrial world by empowering frontline maintenance workers to achieve more. We’ve fused cutting-edge hardware with innovative software into one powerful platform, disrupting legacy systems and delivering smarter, faster solutions for our clients.

Data Science at TRACTIAN
The Data Science team at TRACTIAN focuses on extracting valuable insights from vast amounts of industrial data. Using advanced statistical methods, algorithms, and data visualization techniques, this team transforms raw data into actionable intelligence that drives decision-making across engineering, product development, and operational strategies. The team constantly works on optimizing prediction models, identifying trends, and providing data-driven solutions that directly enhance the company’s operational efficiency and the quality of its products.

What you'll do

We’re hiring Data Foundry Engineers to join Tractian’s Machine Learning Engineering team.

This role focuses on building and improving datasets used across the company. The team works across data engineering, back-end, front-end, product, labeling, and AI engineering to make information more structured, reliable, and useful for internal systems, clients, and AI applications.

Tractian processes more than 20 million industrial data samples per day. We are looking for someone who is comfortable working close to the data: understanding source quality, identifying inconsistencies, improving processing logic, and helping turn fragmented information into cohesive datasets.

Responsibilities
  • Improve and maintain data enrichment pipelines built in Python
  • Work on data curation tasks such as labeling, deduplication, normalization, and inconsistency resolution
  • Investigate data quality issues by querying and combining data from multiple databases
  • Prototype tools, APIs, and workflows to support data operations
  • Use and create AI-based tools responsibly to support data tasks, including reviewing automated outputs when needed
  • Collaborate with product, data, and machine learning teams on data quality and usability
  • Document workflows and methodology to keep processes reproducible and auditable

Requirements
  • 3+ years of experience in data engineering, software engineering, machine learning engineering, or similar roles
  • Strong Python skills for data processing and automation
  • Solid experience with Pandas and tabular data manipulation
  • Experience with data curation, enrichment, labeling, or quality control workflows
  • Experience working with large or messy datasets
  • Ability to evaluate whether data outputs are correct, consistent, and useful
  • Comfortable working across new data domains and switching context between different datasets and workflows
  • Organized and methodical approach to technical work

Technical Skills
  • Experience with APIs and service integration
  • Familiarity with FastAPI or similar frameworks for internal APIs
  • Familiarity with SQL and/or NoSQL databases
  • Familiarity with workflow orchestration tools such as Airflow or Temporal
  • Familiarity with AI-assisted workflows, including using LLMs or agents for data tasks and reviewing automated outputs
  • Experience with spreadsheets and tabular analysis
  • Experience with web scraping or extracting data from semi-structured sources

Nice to Have
  • Experience with Streamlit or similar tools for internal prototypes
  • Experience with Selenium, Playwright, or similar browser automation tools
  • Familiarity with industrial, maintenance, or operational data
  • Familiarity with gRPC is a plus
  • Experience with annotation or labeling workflows
  • Background in scientific or highly reproducible technical work

Skills Required

  • 3+ years of experience in data engineering, software engineering, machine learning engineering, or similar roles
  • Strong Python skills for data processing and automation
  • Solid experience with Pandas and tabular data manipulation
  • Experience with data curation, enrichment, labeling, or quality control workflows
  • Experience working with large or messy datasets
  • Ability to evaluate whether data outputs are correct, consistent, and useful
  • Comfortable working across new data domains and switching context between different datasets and workflows
  • Organized and methodical approach to technical work
  • Experience with APIs and service integration
  • Familiarity with FastAPI or similar frameworks for internal APIs
  • Familiarity with SQL and/or NoSQL databases
  • Familiarity with workflow orchestration tools such as Airflow or Temporal
  • Familiarity with AI-assisted workflows, including using LLMs or agents for data tasks and reviewing automated outputs
  • Experience with spreadsheets and tabular analysis
  • Experience with web scraping or extracting data from semi-structured sources
  • Experience with Streamlit or similar tools for internal prototypes
  • Experience with Selenium, Playwright, or similar browser automation tools
  • Familiarity with industrial, maintenance, or operational data
  • Familiarity with gRPC
  • Experience with annotation or labeling workflows
  • Background in scientific or highly reproducible technical work
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
103 Employees

What We Do

Tractian is a machine-intelligence company delivering integrated hardware, cloud software and AI to prevent machine failures and boost industrial uptime. Their offering combines vibration and condition sensors, TracOS maintenance-management software, and AI-driven analytics to enable predictive maintenance, energy optimization and operational visibility for factories and asset-heavy operations globally.

Similar Jobs

TRACTIAN Logo TRACTIAN

Data Engineer

Artificial Intelligence • Machine Learning • Software
In-Office or Remote
São Paulo, BRA
103 Employees

CrowdStrike Logo CrowdStrike

Alliance Partner Executive, MSSP (Remote, BRA)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
Brazil
11000 Employees

Deepgram Logo Deepgram

Research Staff, LLMs

Artificial Intelligence • Machine Learning • Natural Language Processing • Software • Conversational AI
In-Office or Remote
49 Locations
150 Employees
150K-250K Annually

Circle Logo Circle

Director, Brazil Compliance Officer & MLRO

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
In-Office or Remote
São Paulo, BRA
1050 Employees

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account