Hybrid — Denver, CO
Full-time
We’re looking for a Software / Applied AI Engineer to build production AI capabilities that make a private AI and data platform scalable, reliable, and repeatable.
You’ll own core components of the AI platform, including agentic and multi-agent systems, reusable AI architectures, evaluation and reliability mechanisms, and platform capabilities such as automated fine-tuning and runtime optimization of infrastructure and models.
This work operates within clearly defined production boundaries. AI behavior must remain within established limits, and systems are not deployed into production without evaluation, clear ownership, operational controls, and reliable rollback mechanisms.
You’ll partner with data and infrastructure teams to translate requirements and feedback into platform capabilities that scale across deployments. This is a hands-on opportunity for someone who thrives in a high-ownership environment and wants to build the infrastructure that enables real-world AI applications.
What You’ll DoBuild and operate production AI capabilities, including agentic and multi-agent workflows, tool calling, orchestration, and reusable patterns that scale.
Design and implement evaluation, monitoring, and quality systems that make AI behavior measurable, reliable, continuously improving, and safe in production.
Build private AI platform capabilities, including automated fine-tuning workflows, model and runtime optimization, and inference performance improvements under real-world constraints.
Implement safety and operational controls that keep AI behavior bounded and production-ready, including policy constraints, approval workflows, auditability, and rollback mechanisms.
Develop practical interfaces and APIs that make AI capabilities easy to integrate across platform services and customer environments.
Improve developer velocity through automation and tooling, using AI tools to accelerate implementation, testing, documentation, and iteration while applying sound engineering judgment.
Partner with data and infrastructure teams to ensure the right context reaches inference and agent workflows with predictable latency, reliability, and cost.
For Senior-level roles, mentor engineers, review designs, and help raise the technical bar across the organization.
In your first 3 months, you will have:
Shipped at least one production AI capability, such as agents, evaluation, fine-tuning, or runtime optimization, that improves platform reliability, performance, or usability.
Established a strong evaluation and rollback model for at least one AI workflow operating within defined production boundaries.
Earned trust through autonomy and execution, becoming a go-to owner for production AI platform capabilities.
In your first year, you will be:
Owning major components of the private AI platform end-to-end, with clear accountability for reliability, performance, and platform adoption.
Shipping reusable AI structures that shorten adoption cycles and scale across deployments, including evaluation, guardrails, orchestration, optimization, and operational playbooks.
Driving platform evolution through enhancements grounded in real-world constraints and measurable outcomes, with safe rollout and rollback as standard practices.
6+ years of experience building and operating production software systems.
Experience shipping AI-enabled platforms or agentic systems is strongly preferred.
Strong fundamentals in distributed systems, performance, and reliability.
Comfortable owning production services end-to-end, including Docker/Kubernetes deployments, REST/gRPC APIs, and disciplined rollout and rollback practices.
Experience building evaluation frameworks, monitoring, and safety or guardrail systems that enable controlled AI behavior in production.
Familiarity with automated evaluation harnesses, drift and quality monitoring, tracing, and structured telemetry.
Strong engineering craft, including clean implementations, thoughtful designs, operational clarity, and thorough documentation.
Experience with technologies such as Python and/or TypeScript/Go, FastAPI-style services, and effective testing practices.
Comfortable working in ambiguous environments and making sound trade-offs involving latency, cost, GPU utilization, and reliability.
Clear communicator and strong collaborator across engineering and commercial teams.
Ownership mindset focused on outcomes rather than tasks.
Production experience building agentic and multi-agent systems, orchestration layers, and evaluation frameworks with clear reliability goals.
Experience with tool calling, workflow orchestration, and measurable, repeatable evaluation loops.
Experience designing reusable AI structures, including tool calling, memory and state patterns, policy constraints, and safety or guardrail systems.
Experience deploying these capabilities through stable APIs across multiple applications or environments.
Experience building fine-tuning workflows and runtime optimization systems for private AI deployments.
Familiarity with inference optimization, batching, caching, GPU efficiency, and vLLM-style serving environments.
Experience building monitoring and quality systems for AI behavior that enable measurable improvement and safe rollback.
Familiarity with offline and online evaluation, tracing, structured logs, metrics, and incident-driven iteration.
Strong systems instincts across data, infrastructure, and security constraints that affect AI in production.
Experience with supporting systems such as Kafka-style event systems, Postgres, time-series databases, and secure deployment patterns.
Work in a high-ownership, real-world startup environment where you can move quickly, build new systems, and see your impact directly.
Use modern AI tools throughout development, testing, documentation, and troubleshooting workflows to accelerate execution.
Take on challenging technical problems at the intersection of infrastructure, cloud, IoT, hardware/software systems, networking, data, and AI.
Collaborate with exceptional teammates and industry leaders across software, AI, and infrastructure.
This role may be filled at either the Senior or Staff level.
Base salary range:
Senior: $130,000–$155,000
Staff: $160,000–$185,000
Eligibility for meaningful equity through stock options in an early-stage, high-growth company.
Eligibility to participate in company benefit plans, which may include health, dental, and vision coverage, a 401(k) with company match, flexible PTO, paid parental leave, commuter benefits, and relocation and visa support for eligible roles.
Skills Required
- 6+ years of experience building and operating production software systems
- Strong fundamentals in distributed systems, performance, and reliability
- Experience with Docker and Kubernetes deployments
- Experience building and owning REST or gRPC APIs
- Experience with disciplined production rollout and rollback practices
- Experience building evaluation frameworks, monitoring, and safety or guardrail systems for controlled AI behavior
- Familiarity with automated evaluation harnesses, drift and quality monitoring, tracing, and structured telemetry
- Strong engineering practices, including clean implementations, thoughtful designs, operational clarity, and thorough documentation
- Experience with Python and/or TypeScript or Go
- Experience with FastAPI-style services and effective testing practices
- Ability to make trade-offs involving latency, cost, GPU utilization, and reliability
- Clear communication and collaboration across engineering and commercial teams
- Production experience building AI-enabled platforms or agentic systems
- Production experience with agentic and multi-agent systems, orchestration layers, and evaluation frameworks
- Experience with tool calling, workflow orchestration, and repeatable evaluation loops
- Experience designing reusable AI structures, memory and state patterns, policy constraints, and guardrail systems
- Experience building fine-tuning workflows and runtime optimization systems
- Familiarity with inference optimization, batching, caching, GPU efficiency, and vLLM-style serving environments
- Experience with AI behavior monitoring, offline and online evaluation, tracing, structured logs, metrics, and safe rollback
- Experience with Kafka-style event systems, Postgres, time-series databases, and secure deployment patterns
What We Do
SourceDirect Talent is a talent advisory and recruiting firm serving seed and early-stage startups. It provides AI-powered recruiting solutions to help customers build go-to-market and engineering teams, alongside global people consulting. Its services cover end-to-end recruitment, immigration, HR, and advisory support, using AI talent agents, sourcing frameworks, and data-driven processes to help growing companies scale hiring and improve recruitment capacity.









