Build platform capabilities for agentic workflows, including tool/function calling, orchestration, state/context management, and human-in-the-loop approvals. Create self-service tooling and paved paths that allow engineering teams to consume AI platform capabilities independently.
Numbers & Facts
Location
Fort Collins, CO (Remote)
Description
Duration: Permanent/Direct Hire Compensation: $150-175k Base + Bonus + Benefits Location: Fort Collins, CO (on-site) Remote Eligibility: CA, DC, FL, GA, IL, IN, MN, MS, NC, NV, NY, OH, OR, PA, SC, TN, TX, VA, and WA
Responsibilities:
Build reusable platform services and frameworks for production AI applications.
Establish common patterns for LLMs, RAG, AI agents, and machine learning services.
Develop shared APIs, SDKs, libraries, templates, and internal tooling.
Build platform capabilities for agentic workflows, including tool/function calling, orchestration, state/context management, and human-in-the-loop approvals.
Develop reusable RAG capabilities including ingestion, chunking, embeddings, retrieval, ranking, and grounding.
Build AI evaluation, guardrails, monitoring, and observability.
Support model serving, inference services, vector/search infrastructure, and AI data pipelines.
Build CI/CD and deployment patterns for AI applications.
Establish LLMOps/MLOps practices for versioning, testing, deployment, monitoring, and rollback.
Create self-service tooling and paved paths that allow engineering teams to consume AI platform capabilities independently.
Design and operate secure, scalable AWS infrastructure for AI workloads.
Partner with AI Engineers, Software Engineers, Data Engineers, Data Scientists, Security, SRE, and Product teams.
Help establish architecture and engineering standards for production AI.
Requirements:
5+ years of software engineering, platform engineering, SRE, DevOps, or cloud infrastructure experience.
Strong hands-on AWS experience.
Strong Python development experience.
Hands-on experience building or supporting production Generative AI / LLM applications.
Experience designing and implementing RAG solutions.
Experience with embeddings, vector search/vector databases, retrieval, and grounding.
Experience with AI agents / agentic workflows, including tool or function calling.
Experience with LLMOps/MLOps practices.
Experience implementing AI evaluation, guardrails, and observability.
Hands-on experience with Kubernetes, containers, and Terraform/IaC.
Experience building APIs, services, shared frameworks, or platform capabilities used by multiple engineering teams.
Strong understanding of distributed systems, reliability, monitoring, and production operations.
Experience working with sensitive or regulated data.
Nice to Have:
Experience with multiple LLM providers and model-routing/model-gateway architectures.
Experience with AI orchestration or agent frameworks.
Experience with vector databases and enterprise search platforms.
Experience with event-driven or real-time architectures.
Experience building internal developer platforms or self-service engineering tooling.
Experience with fraud, risk, reconciliation, or financial workflow use cases.
Fintech, banking, payments, or other regulated-industry experience.
Experience with model serving and inference infrastructure.
Experience optimizing AI systems for latency, scalability, and cost.