Technical Lead

Posted 8 Days Ago
Be an Early Applicant
Hiring Remotely in Wierzchy, Osie, Świecki, Kujawsko-pomorskie, POL
In-Office or Remote
Senior level
Software
The Role
Found and lead the engineering team building Veris EvalOps from zero to paying clients. Own the AI evaluation engine, LLM observability and tracing, RAG knowledge health pipeline, multi-tenant SaaS platform, integrations, and architecture. Hire and manage 5–7 engineers across platform, release-gating, and knowledge-health streams while working directly with pilot clients. The role requires deep production experience in LLM evaluation, RAG, agent systems, observability, async Python, and backend platform engineering.
Summary Generated by Built In
Company Description

SmartDev is an AI-powered software development company headquartered in Vietnam, part of the Verysell Group. We help global businesses deliver faster and build smarter — combining AI-driven development practices with deep expertise across fintech, healthcare, retail, and enterprise technology. Our team of engineers, architects, and AI specialists works across the full stack: from custom software and cloud solutions to generative AI, MLOps, and intelligent automation. At SmartDev, we believe that great architecture and AI-first thinking are a competitive advantage — and we build our teams accordingly. 

The company is at a pivotal point in our journey: transitioning from a pure IT outsourcing provider to an AI-enabled solutions partner

We have over 200 talented employees working both on-site and remote/hybrid at locations:

  • Danang: 81 Quang Trung, Hai Chau 1 ward, Hai Chau District, Danang
  • Hanoi: Zodiac building, 19 Duy Tan Street, Dich Vong Hau ward, Cau Giay District, Hanoi

Job Description

You'll be the founding technical lead for Veris EvalOps, building the platform that answers the two questions every AI-deploying business needs answered: is this system safe to launch, and is it still working correctly a month later. You'll take it from first line of code to first paying clients in 6–7 months.

Role Summary

This is a zero-to-one build, not a maintenance role. You'll architect and ship two commercial modules — a pre-production Release Gate that turns “looks good” into a reproducible readiness score, and a Knowledge Health Monitor that continuously audits the knowledge base an AI draws from — while hiring and leading the engineers who build them alongside you. There's no principal architect above you to escalate to: you make the calls and live with them, with the first pilot client live by month 3.

Key Responsibilities

• Own the evaluation engine. LLM-as-judge scoring, rule-based checks, groundedness verification, hallucination detection, and regression comparison — every readiness score comes from here.

• Build the tracing and observability layer. Distributed tracing across LLM calls, RAG retrievals, and agent workflows, built on OpenTelemetry, capturing every token, tool call, cost, and latency metric.

• Ship the knowledge health pipeline. Ingestion and continuous analysis of enterprise knowledge sources — stale-content detection, contradiction analysis, and coverage-gap mapping.

• Own platform core and integrations. Multi-tenant architecture, RBAC, API connectors, dashboards, and the CI/CD hooks that let the Release Gate plug into client engineering workflows.

• Build and lead the team. Hire and run 5–7 engineers across three streams — Platform Core, Release Gate, Knowledge Health — and own every architecture decision end to end.

Qualifications

Must-have

• 5+ years in engineering. Including 2+ years leading a team of 3–8 through a complete build cycle — architecture to shipping to paying users. Not a first-time lead role.

• LLM evaluation methodology. LLM-as-judge design, RAGAS/DeepEval-style metrics, golden dataset construction, regression testing for AI systems, and hallucination detection — not just the library calls, the mechanics behind them.

• LLM observability and tracing. OpenTelemetry-based tracing across LLM calls, RAG retrievals, and multi-step agent trajectories; cost/latency attribution; drift and anomaly detection.

• RAG system architecture. Production experience across the full pipeline — chunking, embeddings, a vector store (Pinecone, Weaviate, Qdrant, or pgvector), retrieval, re-ranking.

• AI agent systems. Production experience with agent patterns (ReAct, Plan-and-Execute, supervisor/sub-agent), tool-call evaluation, and guardrails.

• Backend platform engineering. Production-grade async Python (FastAPI, Celery), multi-tenant SaaS architecture, PostgreSQL/Redis, and CI/CD integration.

• Build-vs-integrate judgment and client-facing comfort. Can weigh integrating Langfuse/Braintrust vs. building from scratch, and work directly with pilot clients during onboarding and results review.

Nice-to-have

• LLM APIs and model ecosystem. Multi-provider experience (OpenAI, Anthropic, Azure OpenAI, Bedrock) and routing/prompt-management at scale.

• MLOps and experiment tracking. Background with MLflow, Weights & Biases, or equivalent experiment-tracking tooling.

• Security, compliance, and AI governance. EU AI Act and NIST AI RMF awareness, PII handling in AI pipelines, and red-teaming basics — increasingly a qualification question in enterprise security reviews.

Skills Required

  • 5+ years of engineering experience
  • 2+ years leading a team of 3–8 through a complete build cycle from architecture to shipping to paying users
  • Experience designing LLM-as-judge systems, RAGAS/DeepEval-style metrics, golden datasets, AI regression testing, and hallucination detection mechanics
  • Production experience with OpenTelemetry-based tracing across LLM calls, RAG retrievals, and multi-step agent trajectories
  • Experience with cost and latency attribution, drift detection, and anomaly detection
  • Production RAG architecture experience covering chunking, embeddings, vector stores, retrieval, and re-ranking
  • Production experience with AI agent patterns, tool-call evaluation, and guardrails
  • Production-grade asynchronous Python, FastAPI, Celery, multi-tenant SaaS architecture, PostgreSQL, Redis, and CI/CD integration
  • Ability to make build-versus-integrate decisions and work directly with pilot clients
  • Experience with multiple LLM providers including OpenAI, Anthropic, Azure OpenAI, or Bedrock
  • Experience with MLflow, Weights & Biases, or equivalent experiment-tracking tools
  • Awareness of EU AI Act, NIST AI RMF, PII handling in AI pipelines, and red-teaming basics
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Nyon
245 Employees

What We Do

Founded in 1990, Verysell Group is a renowned global software development company, headquartered in Switzerland. We offer various services from staff augmentation to building bespoke software solutions for industries such as Fintech, Payment tech & InsurTech. Central to our business is the offshore development center (ODC) SmartDev LLC in Vietnam where we develop software for external customers and our in-house products such as VeryPay - a mobile money payment platform. Our offices are located in 8 cities on 4 continents. This allows us to cater to a diverse clientele, including start-ups, scale-ups, and large global enterprises. Our family of brands SmartDev, VeryPay, Smart81 & VeryPlay Studio helps us focus on specific market segments and client types. SmartDev is a broad base outsourcing partner for enterprise clients worldwide. Recognized as the Winner of SME100® Fast Moving Companies, SmartDev specializes in building FinTech apps, insurance software, and more. Smart81 is a subsidiary of SmartDev and is instrumental in accelerating growth for start-ups and scale-ups by advising, collaborating, and crafting enterprise-level technology through our offices in California, Singapore, and London – the world's leading centers of startup economy. VeryPay is a mobile payments technology provider offering technology solutions for African mobile network operators, while VeryPlay Studio serves as a full-cycle game developer, merging Switzerland's quality standards with Asia's production expertise. The Applied AI Lab, another innovative extension of the Verysell Group, showcases our commitment to AI revolution by providing high-level consulting services, aiming to integrate AI functionalities in business, thereby enhancing productivity and expanding product offerings. AI Lab is also an in-house center of excellence in this field with a mandate for AI transformation.

Similar Jobs

BJAK Logo BJAK

Technical Lead

Artificial Intelligence • Fintech • Software • Financial Services
Remote
Poland
253 Employees

N-iX Logo N-iX

Technical Lead

Information Technology • Consulting
Remote
Poland
2135 Employees

Masabi Logo Masabi

Technical Lead

Transportation • Travel
Remote
Poland
239 Employees

TRG Solutions Logo TRG Solutions

Technical Lead

Information Technology • Consulting
Remote
Poland
71 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account