The Role
Design, develop, and deploy production LLM-powered applications and APIs, implement RAG and document-processing workflows (OCR + LLM), build conversational interfaces, ensure reliability and fallbacks, and collaborate with product teams.
Summary Generated by Built In
Why join us
AI Engineering at Tractian
What you'll do
Responsibilities
Requirements
Preferred Qualifications
Technical Skills
Compensation & Benefits
Tractian is transforming the industrial world by empowering frontline maintenance workers to achieve more. We’ve fused cutting-edge hardware with innovative software into one powerful platform, disrupting legacy systems and delivering smarter, faster solutions for our clients.
The AI Engineering team at Tractian focuses on extracting actionable intelligence from vast amounts of industrial assets, telemetry, and maintenance workflows. We bridge the gap between state-of-the-art machine learning models and mission-critical industrial applications. By bringing rigorous software engineering principles to modern AI, we build scalable, event-driven, and resilient systems that frontline engineers rely on 24/7, directly optimizing frontline productivity and asset reliability.
We're seeking a Senior AI Engineer with a solid software engineering foundation to design, develop, and deploy production-ready AI and agentic systems. In this role, you will bring strong product vision and high agency, proactively exploring the problem space and driving technical solutions end-to-end. You'll build robust APIs, architect event-driven data streaming pipelines, implement intelligent information retrieval systems, and orchestrate custom agentic workflows that connect AI capabilities with real-world industrial applications.
Responsibilities
- Design and develop production-grade applications and custom agentic workflows powered by LLMs.
- Build intelligent data processing workflows using OCR, PDF layout analysis, audio and image models.
- Architect and scale event-driven data streaming pipelines using Apache Kafka to power AI workloads.
- Create and maintain high-performance APIs that integrate AI capabilities with our platform.
- Implement evaluation frameworks (evals) to benchmark and track model quality, accuracy, and regressions.
- Partner proactively with product teams to define technical requirements.
- Mentor engineers, drive architecture discussions, and ensure best practices through code reviews.
Requirements
- 5+ years of experience in software engineering with a strong backend/distributed systems foundation.
- 2+ years of experience building and shipping production AI/LLM systems.
- Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent technical field.
- Demonstrated experience building production-grade applications with custom agentic systems, document processing pipelines, and/or information-retrieval systems.
- Hands-on experience with event-driven architectures and streaming using Apache Kafka.
- Working knowledge of AI evaluation methods (evals), observability/tracing, security guardrails, cost optimization, latency reduction, and semantic caching.
- Strong proficiency in Python and/or Golang, with experience in frameworks like FastAPI or Flask.
- Strong proficiency in designing, building, and maintaining production REST APIs and databases.
- High ownership and strong product sense: able to take ambiguous problems, collect technical requirements, prioritize impact, and execute independently.
Preferred Qualifications
- Experience with LLM fine-tuning and model optimization techniques.
- Knowledge of containerization (Docker, Kubernetes) and microservices architecture.
Technical Skills
- Programming: Python and/or Golang.
- Frameworks: FastAPI, Flask.
- Agentic Systems: Custom state machines, tool-calling pipelines, execution loops.
- Streaming & Messaging: Apache Kafka, event-driven architecture.
- Document Processing: OCR engines, PDF parsing, multimodal LLMs.
- AI Reliability: Evals (LLM-as-a-judge, benchmarks), observability, tracing, guardrails.
- API Development: REST, OpenAPI, gRPC.
- Databases: PostgreSQL, Redis, vector databases.
- LLM Integration: OpenAI, Anthropic, and open-source models.
- Competitive Compensation
- 30 days of paid annual leave
- Education and courses stipend
- Earn a trip anywhere in the world every 4 years
- R$1.035/month for meals allowance
- Health plan with national coverage and without coparticipation
- Dental Insurance: we help you with dental treatment for a better quality of life.
- Wellhub Membership: Access a wide range of gyms and training programs.
Skills Required
- 1-4 years of experience in software engineering with a focus on backend development
- Bachelor's degree in Computer Science, Computer Engineering, Information Systems, or equivalent technical field
- Demonstrated experience building applications that utilize large language models
- Strong proficiency in Python
- Experience with web frameworks like FastAPI, Flask, or Django
- Experience designing, building, and maintaining production REST APIs
- Working knowledge of relational databases and SQL
- Familiarity with LLM orchestration tools such as LangChain or LangGraph
- Experience with vector embeddings and information retrieval systems
- History of shipping code to production environments
- Experience with LLM fine-tuning and model optimization techniques
- Knowledge of containerization and microservices architecture
- Experience with OCR and document processing technologies (PDF parsing)
- Understanding of performance optimization for AI applications
- Proficiency with Golang
- Experience with PostgreSQL, Redis, and vector databases
- Experience integrating with LLM providers such as OpenAI and Anthropic
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Tractian is a machine-intelligence company delivering integrated hardware, cloud software and AI to prevent machine failures and boost industrial uptime. Their offering combines vibration and condition sensors, TracOS maintenance-management software, and AI-driven analytics to enable predictive maintenance, energy optimization and operational visibility for factories and asset-heavy operations globally.








