Senior Software Engineer - AI Innovation

Posted Yesterday
4 Locations
In-Office or Remote
Senior level
Artificial Intelligence • Fintech • Software • Financial Services
Worth is the underwriting & onboarding platform that helps financial institutions say yes to small businesses, faster.
The Role
Designs and ships production multi-agent AI systems for KYB, underwriting, case review, and risk monitoring. Owns agent architecture, retrieval, tool integrations, evaluations, observability, MLOps, reliability, and compliance guardrails. Requires strong Python engineering, agent framework expertise, distributed systems experience, Kubernetes operations, and production LLM deployment skills. Mentors engineers and partners with product, security, compliance, and applied science teams. Remote employees travel to Orlando at least twice annually.
Summary Generated by Built In

Worth AI, a leader in AI onboarding and underwriting, is looking for a talented and experienced Senior Software Engineer - AI Innovation to join our team. At Worth AI, we are on a mission to revolutionize decision-making with the power of artificial intelligence helping fintechs, lenders, payment processors, and financial institutions onboard small businesses faster, smarter, and more confidently. We’re building the infrastructure that powers real-time KYB, KYC/IDV, underwriting, and continuous risk monitoring at enterprise scale, and the Worth Score™ our unified credit score derived from 1,200+ data points across 700M+ SMBs.

As a Senior Software Engineer - AI Innovation, you won’t be wiring up demos you’ll be designing and shipping production agent systems that make consequential decisions on regulated financial data. Worth’s platform consolidates onboarding and underwriting into a single AI-powered system, and our agents read, reason over, and act on the messy, high-stakes signals that come with that domain. You’ll own the end-to-end lifecycle: architecting agent graphs, building the retrieval and tool layers they rely on, instrumenting them with evals and observability, and getting them deployed against SOC 2 / GDPR / CCPA guardrails. You’ll partner closely with our Chief AI Officer, applied scientists, product, and platform teams to turn agentic patterns into customer outcomes.

Responsibilities
  • Design and ship multi-step agentic systems (planner/executor, tool-using, multi-agent, human-in-the-loop) that automate KYB, underwriting, case review, and risk monitoring workflows.
  • Architect agent graphs in LangGraph (or comparable frameworks CrewAI, AutoGen, Claude Agent SDK) with explicit state, durable execution, retries, and safe fallbacks.
  • Build and harden the retrieval layer powering our agents chunking strategies, hybrid search, reranking, and grounded citation across SoS filings, IRS records, bank data, and Worth’s 700M+ SMB graph.
  • Own the eval stack: golden sets, offline regression suites, LLM-as-judge, online A/B and shadow evals, and red-teaming for jailbreaks, prompt injection, and PII leakage.
  • Wire agents into Worth’s production systems via well-typed tools, MCP servers, and existing services (decisioning engine, case management, crosswalking). Treat tool surface area as a product.
  • Drive production MLOps for agents: deployment, versioning, traffic shaping, cost/latency budgets, observability (traces, token spend, tool call success), and on-call playbooks for agent incidents.
  • Partner with security, compliance, and legal to keep agents inside Worth’s SOC 2, GDPR, CCPA, and fair-lending posture — building from day one, not bolted on.
  • Translate ambiguous product bets (“what if the underwriter had an AI co-pilot for this?”) into concrete agent designs, prototypes, and shipped features.
  • Mentor engineers across the org on agent patterns, prompt engineering hygiene, eval discipline, and the failure modes of LLM systems.
  • Stay ahead of the frontier new models, frameworks, and patterns — and bring back what actually works in production.

Technology Stack

  • Languages & Runtimes: Python, Node.js, TypeScript
  • Agent / LLM frameworks: LangGraph, LangChain, Claude Agent SDK, MCP, OpenAI SDK
  • Models: Anthropic Claude, OpenAI, open-weight (Llama, Mistral) where appropriate
  • Retrieval & Data: PostgreSQL, pgvector / vector DBs, OpenSearch, Kafka, Redshift, Redis
  • Infra & Orchestration: AWS, Kubernetes (EKS), ArgoCD, Terraform
  • Evals & Observability: LangSmith / Langfuse / Braintrust-style tooling, DataDog, custom eval harnesses

Requirements
  • 8+ years of professional software engineering experience, with at least 2 years building production LLM or agentic systems (not just notebooks or demos).
  • Solid software engineering experience - front-end, APIs, async patterns, queues, databases, and the failure modes of distributed systems.
  • Demonstrated ownership of major features or subsystems in production.
  • Demonstrated experience mentoring junior engineers and raising team quality standards.
  • Demonstrated experience with event-driven systems: enrichment, retries, dead-lettering, backpressure.
  • Experience managing containerized applications in Kubernetes, EKS, ArgoCD, operators, Kustomize.
  • Deep, hands-on experience with at least one modern agent framework (LangGraph strongly preferred) and a track record of shipping agents that actually run, fail gracefully, and recover.
  • Real experience with evals you’ve built golden sets, run offline and online evaluations, and used them to make ship/no-ship calls.
  • Production MLOps fluency: you’ve deployed LLM workloads under real latency, cost, and reliability constraints, and you instrument what you ship.
  • Strong proficiency in Python; comfortable in TypeScript / Node.js for integrating with Worth’s services.
  • Clear, calibrated communicator - able to explain agent trade-offs to product, security, and customers without hand-waving.
  • Operates with extreme ownership in ambiguous, fast-moving environments. Excited to work alongside a team that values “One Team”, “Extreme Ownership”, and “Create Raving Fans.”

Success Metrics

  • Agent Quality: Measurable improvements in task success rate, grounding accuracy, and hallucination rate on Worth’s eval suites, tied to customer-visible outcomes.
  • Production Reliability: Agents you own meet defined SLOs for latency (P90/P99), tool-call success rate, and cost per task.
  • Velocity: New agent capabilities go from prototype to production in weeks, not quarters, without skipping evals or guardrails.
  • Risk Posture: Zero material incidents tied to prompt injection, PII leakage, or unsafe tool use on agents you own.
  • Force Multiplier: Patterns, tools, and eval scaffolding you build are adopted by other engineers across Worth.

Bonus Points (nice to haves, not requirements)

  • Prior experience in fintech, lending, payments, KYB/KYC, fraud, or AML — or any other regulated, high-stakes data domain.
  • Experience building MCP servers or other structured tool interfaces for LLMs.
  • Background in classical ML (ranking, scoring, calibration) you can bring to bear alongside LLM systems.
  • Experience designing explainable / auditable AI workflows for regulated environments (SOC 2, model risk management, fair lending).
  • Open-source contributions to agent frameworks, eval tooling, or retrieval libraries.
  • Hands-on AWS depth (EKS, MSK, RDS, S3, Lambda) and IaC with Terraform.

**All Remote Hires — will be required to travel to Orlando, Florida at least twice per year for Town Halls and team collaboration, in addition to orientation in Orlando, Florida.


Benefits
  • Health Care Plan (Medical, Dental & Vision)
  • Retirement Plan (401k, IRA)
  • Life Insurance
  • Flexible Paid Time Off
  • 9 paid Holidays
  • Family Leave
  • Remote
  • Hybrid work (for Orlando Associates)
  • Free Food & Snacks (Orlando)
  • Wellness Resources

Skills Required

  • 8+ years of professional software engineering experience
  • At least 2 years building production LLM or agentic systems
  • Software engineering experience with front-end, APIs, asynchronous patterns, queues, databases, and distributed systems
  • Ownership of major production features or subsystems
  • Experience mentoring junior engineers and raising team quality standards
  • Experience with event-driven systems, including enrichment, retries, dead-lettering, and backpressure
  • Experience managing containerized applications in Kubernetes, EKS, ArgoCD, operators, and Kustomize
  • Hands-on experience with a modern agent framework, preferably LangGraph
  • Experience shipping agents that fail gracefully and recover
  • Experience building golden sets and running offline and online evaluations for production AI systems
  • Production MLOps experience deploying LLM workloads under latency, cost, and reliability constraints
  • Strong proficiency in Python
  • Comfort with TypeScript and Node.js
  • Ability to communicate agent trade-offs to product, security, and customers
  • Extreme ownership in ambiguous, fast-moving environments
  • Fintech, lending, payments, KYB/KYC, fraud, AML, or other regulated-domain experience
  • Experience building MCP servers or structured LLM tool interfaces
  • Classical machine learning experience in ranking, scoring, or calibration
  • Experience designing explainable or auditable AI workflows for regulated environments
  • Open-source contributions to agent frameworks, evaluation tooling, or retrieval libraries
  • Advanced AWS experience with EKS, MSK, RDS, S3, or Lambda
  • Infrastructure-as-code experience with Terraform

Worth Compensation & Benefits Highlights

  • Healthcare Strength — Medical, dental, and vision coverage are offered alongside HSA/FSA options and life insurance, forming a broad core package. Employer-paid healthcare for individuals is described as solid.
  • Leave & Time Off Breadth — Unlimited/flexible PTO, paid holidays, and bereavement leave are included, with time off characterized as generous in places. These elements signal above-baseline leave breadth for a startup.
  • Parental & Family Support — Parental/family leave is part of the package and is described as generous. This adds depth to the core leave offering.

Worth Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Winter Park, Florida
84 Employees
Year Founded: 2023

What We Do

Worth is the AI-powered platform that consolidates onboarding, underwriting, and risk monitoring for fintechs, lenders, payment processors, and financial institutions. Founded in 2023, we built Worth to replace slow, manual underwriting with a single system that verifies, scores, and monitors small and medium-sized businesses (SMBs) in real time. At the center of the platform is Crosswalking Technology. Our proprietary AI/ML models intelligently match businesses across disparate data sources, ensuring the highest level of accuracy and reliability in SMB entity resolution. By integrating multiple first- and third-party authoritative data sources into our crosswalk-matching logic, Worth ensures that businesses are correctly identified, even in cases of duplicate addresses, name variations, or incomplete records. This data moat spans 186 integrations and 25 global and local partners across 200+ countries and territories, resolving fragmented SMB signals into a database of 350M+ SMBs with a 98% data match rate. Our product suite — Worth Pre-Fill, Custom Onboarding, Case Management, Decisioning Engine, Perpetual Risk Monitoring, and Worth Wallet — is available via API, SDK, or fully white-labeled, enabling financial institutions to consolidate their entire onboarding and underwriting stack into one platform. Customers using Worth have increased approval rates by 37%+, reduced application abandonment by 43%+, cut vendor costs by 25%, and reduced time to revenue by 55%+. We're SOC 2 Type II certified and GDPR and CCPA compliant, and have raised $55M in funding to date. Today, 50+ customers rely on Worth to onboard and underwrite their SMB customers faster and more accurately.

Why Work With Us

We're solving a genuinely hard problem: turning fragmented SMB data into one durable, explainable identity that banks and lenders can trust. Backed by $55M in funding and already live with 50+ customers, we're a tight-knit team with real traction, where your work would help shape the roadmap.

Gallery

Gallery
Gallery
Gallery

Worth Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: Not Specified
HQOrlando

Similar Jobs

Worth Logo Worth

Senior Software Engineer

Artificial Intelligence • Fintech • Software • Financial Services
In-Office or Remote
4 Locations
84 Employees

Worth Logo Worth

Solutions Engineer

Artificial Intelligence • Fintech • Software • Financial Services
In-Office or Remote
4 Locations
84 Employees

Worth Logo Worth

Senior Devops Engineer

Artificial Intelligence • Fintech • Software • Financial Services
In-Office or Remote
4 Locations
84 Employees

Worth Logo Worth

Senior Software Engineer

Artificial Intelligence • Fintech • Software • Financial Services
In-Office or Remote
5 Locations
84 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account