Technical Challenge
Responsibilities
- Own and grow the agentic function end-to-end: set architecture across internal (data-engineering & data-science agents) and external (customer/partner-facing agents, MCP servers) surfaces, make build-vs-buy calls, and grow a team under you as we scale.
- Design, build, and deploy production LLM agents that take consequential action (write-back, execute changes) with human-in-the-loop controls — including tool interfaces (MCP, function calling), tiered tool access, approval gates, and rollback.
- Build and maintain eval harnesses for agentic systems — offline and online evaluation, regression testing for non-determinism, guardrails, and observability/tracing.
- Own context and retrieval engineering: go beyond naive chunking to build high-quality retrieval, context assembly, and grounding in proprietary domain data to reduce hallucination.
- Implement graph-based knowledge and retrieval systems (GraphRAG, property graphs, ontology/semantic layers) to ground agents in a complex, structured domain.
- Partner with applied science and the broader data/ML function on pipelines, embeddings, and vector stores where relevant, without owning ML model training.
- Communicate agent architecture, trade-offs, and roadmap to execs and investors; act as a player-coach who writes production code today while owning the function's direction and hiring as the team grows.
Required Skills
- 5+ years shipping production software; strong full-stack/backend engineering (Python core, TS/Node or Go a plus), including production APIs, data models, testing, and CI.
- Proven track record building and shipping LLM agents to production with real users — multi-step, tool-calling, stateful, with orchestration (LangGraph or equivalent) — and the ability to explain the control loop, not just the framework used.
- Experience building eval harnesses for agentic systems (e.g. MLflow, LangSmith, or custom) with fluency in determinism, drift, and guardrails.
- Experience owning a retrieval system in production, including chunking vs. structured retrieval trade-offs and evaluating retrieval quality.
- Experience shipping agents that take consequential, real-world action, with a clear point of view on approval/guardrail/rollback architecture (bonus: an incident where the agent did the wrong thing, and how it was handled).
- Track record leading an agentic initiative or team end-to-end — from architecture to production — including communicating agent systems to both execs and investors.
Preferred Skills
- Experience shipping graph-backed retrieval and explaining why graphs outperform flat vectors for structured domains (GraphRAG, knowledge graphs, ontologies, property graphs, triple stores).
- Experience fine-tuning or distilling an open-source model (LoRA/QLoRA) with measured gains, or strong context/prompt optimization as a substitute.
- Experience serving/operating open-source models (Llama, Qwen, Mistral) via Databricks, vLLM, or similar, with a point of view on self-host vs. hosted-API cost/latency trade-offs.
UP.Labs Summary
Location: Remote
Skills Required
- 5+ years shipping production software with strong full-stack/backend engineering
- Proficiency with Python (core); TypeScript/Node or Go a plus
- Proven track record building and shipping LLM agents to production (multi-step, tool-calling, stateful, orchestration such as LangGraph or equivalent)
- Experience building eval harnesses for agentic systems (MLflow, LangSmith, or custom) and handling determinism, drift, and guardrails
- Experience owning a production retrieval system and evaluating retrieval quality (chunking vs structured retrieval trade-offs)
- Experience shipping agents that perform consequential actions with approval gates, rollback, and human-in-the-loop controls
- Track record leading an agentic initiative or team end-to-end, including communicating architecture and roadmap to execs/investors
- Experience with graph-backed retrieval / knowledge graphs / GraphRAG (preferred)
- Experience fine-tuning or distilling open-source models (LoRA/QLoRA) or strong prompt/context optimization (preferred)
- Experience serving/operating open-source models (Llama, Qwen, Mistral) via Databricks, vLLM, or similar (preferred)
What We Do
We work with global corporate partners to identify the most pressing challenges that they, and broader society, face. Inspired by these complex problems, we launch startups built by proven entrepreneurs, product leaders and technologists that use their agility and talent to develop transformative solutions. After these companies have matured and proven market fit, our corporate partners are able to acquire them, reaping strategic value while enriching their culture and core business. We believe this to be the shortest road to a faster, cleaner, safer, and more accessible future.
Why Work With Us
We launch and innovate 6-8 portfolio organizations a year where no day is the same.
Gallery








