Senior Software Engineer - AI Innovation

Worth AI

  • Orlando, FL
  • 4 days ago
  • Remote

    Highlights

    We’re building the infrastructure that powers real-time KYB, KYC/IDV, underwriting, and continuous risk monitoring at enterprise scale, and the Worth Score our unified credit score derived from 1,200+ data points across 700M+ SMBs. Build and harden the retrieval layer powering our agents chunking strategies, hybrid search, reranking, and grounded citation across SoS filings, IRS records, bank data, and Worth’s 700M+ SMB graph.

    Numbers & Facts

    LocationOrlando, FL (
    Remote
    )

    Description

    Worth AI, a leader in AI onboarding and underwriting, is looking for a talented and experienced Senior Software Engineer - AI Innovation to join our team. At Worth AI, we are on a mission to revolutionize decision-making with the power of artificial intelligence helping fintechs, lenders, payment processors, and financial institutions onboard small businesses faster, smarter, and more confidently. We’re building the infrastructure that powers real-time KYB, KYC/IDV, underwriting, and continuous risk monitoring at enterprise scale, and the Worth Score our unified credit score derived from 1,200+ data points across 700M+ SMBs.

    As a Senior Software Engineer - AI Innovation, you won’t be wiring up demos you’ll be designing and shipping production agent systems that make consequential decisions on regulated financial data. Worth’s platform consolidates onboarding and underwriting into a single AI-powered system, and our agents read, reason over, and act on the messy, high-stakes signals that come with that domain. You’ll own the end-to-end lifecycle: architecting agent graphs, building the retrieval and tool layers they rely on, instrumenting them with evals and observability, and getting them deployed against SOC 2 / GDPR / CCPA guardrails. You’ll partner closely with our Chief AI Officer, applied scientists, product, and platform teams to turn agentic patterns into customer outcomes.

    Responsibilities

    • Design and ship multi-step agentic systems (planner/executor, tool-using, multi-agent, human-in-the-loop) that automate KYB, underwriting, case review, and risk monitoring workflows.
    • Architect agent graphs in LangGraph (or comparable frameworks CrewAI, AutoGen, Claude Agent SDK) with explicit state, durable execution, retries, and safe fallbacks.
    • Build and harden the retrieval layer powering our agents chunking strategies, hybrid search, reranking, and grounded citation across SoS filings, IRS records, bank data, and Worth’s 700M+ SMB graph.
    • Own the eval stack: golden sets, offline regression suites, LLM-as-judge, online A/B and shadow evals, and red-teaming for jailbreaks, prompt injection, and PII leakage.
    • Wire agents into Worth’s production systems via well-typed tools, MCP servers, and existing services (decisioning engine, case management, crosswalking). Treat tool surface area as a product.
    • Drive production MLOps for agents: deployment, versioning, traffic shaping, cost/latency budgets, observability (traces, token spend, tool call success), and on-call playbooks for agent incidents.
    • Partner with security, compliance, and legal to keep agents inside Worth’s SOC 2, GDPR, CCPA, and fair-lending posture — building from day one, not bolted on.
    • Translate ambiguous product bets (“what if the underwriter had an AI co-pilot for this?”) into concrete agent designs, prototypes, and shipped features.
    • Mentor engineers across the org on agent patterns, prompt engineering hygiene, eval discipline, and the failure modes of LLM systems.
    • Stay ahead of the frontier new models, frameworks, and patterns — and bring back what actually works in production.

    Technology Stack

    • Languages & Runtimes: Python, Node.js, TypeScript
    • Agent / LLM frameworks: LangGraph, LangChain, Claude Agent SDK, MCP, OpenAI SDK
    • Models: Anthropic Claude, OpenAI, open-weight (Llama, Mistral) where appropriate
    • Retrieval & Data: PostgreSQL, pgvector / vector DBs, OpenSearch, Kafka, Redshift, Redis
    • Infra & Orchestration: AWS, Kubernetes (EKS), ArgoCD, Terraform
    • Evals & Observability: LangSmith / Langfuse / Braintrust-style tooling, DataDog, custom eval harnesses

    Requirements

    • 8+ years of professional software engineering experience, with at least 2 years building production LLM or agentic systems (not just notebooks or demos).
    • Solid software engineering experience - front-end, APIs, async patterns, queues, databases, and the failure modes of distributed systems.
    • Demonstrated ownership of major features or subsystems in production.
    • Demonstrated experience mentoring junior engineers and raising team quality standards.
    • Demonstrated experience with event-driven systems: enrichment, retries, dead-lettering, backpressure.
    • Experience managing containerized applications in Kubernetes, EKS, ArgoCD, operators, Kustomize.
    • Deep, hands-on experience with at least one modern agent framework (LangGraph strongly preferred) and a track record of shipping agents that actually run, fail gracefully, and recover.
    • Real experience with evals you’ve built golden sets, run offline and online evaluations, and used them to make ship/no-ship calls.
    • Production MLOps fluency: you’ve deployed LLM workloads under real latency, cost, and reliability constraints, and you instrument what you ship.
    • Strong proficiency in Python; comfortable in TypeScript / Node.js for integrating with Worth’s services.
    • Clear, calibrated communicator - able to explain agent trade-offs to product, security, and customers without hand-waving.
    • Operates with extreme ownership in ambiguous, fast-moving environments. Excited to work alongside a team that values “One Team”, “Extreme Ownership”, and “Create Raving Fans.”

    Success Metrics

    • Agent Quality: Measurable improvements in task success rate, grounding accuracy, and hallucination rate on Worth’s eval suites, tied to customer-visible outcomes.
    • Production Reliability: Agents you own meet defined SLOs for latency (P90/P99), tool-call success rate, and cost per task.
    • Velocity: New agent capabilities go from prototype to production in weeks, not quarters, without skipping evals or guardrails.
    • Risk Posture: Zero material incidents tied to prompt injection, PII leakage, or unsafe tool use on agents you own.
    • Force Multiplier: Patterns, tools, and eval scaffolding you build are adopted by other engineers across Worth.

    Bonus Points (nice to haves, not requirements)

    • Prior experience in fintech, lending, payments, KYB/KYC, fraud, or AML — or any other regulated, high-stakes data domain.
    • Experience building MCP servers or other structured tool interfaces for LLMs.
    • Background in classical ML (ranking, scoring, calibration) you can bring to bear alongside LLM systems.
    • Experience designing explainable / auditable AI workflows for regulated environments (SOC 2, model risk management, fair lending).
    • Open-source contributions to agent frameworks, eval tooling, or retrieval libraries.
    • Hands-on AWS depth (EKS, MSK, RDS, S3, Lambda) and IaC with Terraform.

    **All Remote Hires — will be required to travel to Orlando, Florida at least twice per year for Town Halls and team collaboration, in addition to orientation in Orlando, Florida.

    Benefits

    • Health Care Plan (Medical, Dental & Vision)
    • Retirement Plan (401k, IRA)
    • Life Insurance
    • Flexible Paid Time Off
    • 9 paid Holidays
    • Family Leave
    • Remote
    • Hybrid work (for Orlando Associates)
    • Free Food & Snacks (Orlando)
    • Wellness Resources

    Similar Jobs

    See more jobs