Senior AI Platform & Production Engineer

LABUR LLC

  • San Francisco, CA
  • 22 days ago
  • Remote
  • $80–$85 Per Hour

Highlights

LABUR is partnering with a client to hire a Senior AI Platform & Production Engineer for a hands-on engineering role focused on building, operating, and continuously improving shared Enterprise AI platform capabilities and the production AI systems that depend on them. Experience with observability tooling (tracing, monitoring, alerting, debugging) and incident response, with a demonstrated history of evolving production systems post-launch.

Numbers & Facts

LocationSan Francisco, CA (
Remote
)
Salary$80–$85 Per Hour

Description

We respectfully request that 3rd parties refrain from contacting us regarding this posting.

Overview

LABUR is partnering with a client to hire a Senior AI Platform & Production Engineer for a hands-on engineering role focused on building, operating, and continuously improving shared Enterprise AI platform capabilities and the production AI systems that depend on them. The ideal candidate combines strong production software engineering fundamentals with practical experience operating LLM and AI applications at scale, and is comfortable diagnosing complex issues across platform infrastructure and application layers. This role is fully remote.

Responsibilities

  • Build, operate, and evolve the shared Enterprise AI platform alongside the production AI applications it supports
  • Develop and enhance the LLM API Gateway, developer tooling, onboarding workflows, and reusable platform services
  • Design and maintain secure MCP and tool integration infrastructure for agentic AI systems
  • Implement and harden authentication, authorization, policy enforcement, and audit controls across the platform
  • Establish and maintain production-grade observability, including tracing, monitoring, alerting, debugging, and operational dashboards
  • Lead production reliability efforts such as KTLO support, incident response, root-cause analysis, and remediation
  • Optimize AI systems for quality, reliability, latency, scalability, and cost while driving incremental enhancements to deployed applications
  • Collaborate with application engineers, platform teams, and security stakeholders to define standards and enable reliable AI development and operations

Qualifications

  • Strong production software engineering skills in Python and/or TypeScript/Node.js with experience designing backend APIs and distributed systems
  • Hands-on experience with AWS, Kubernetes/EKS, and CI/CD pipelines
  • Proven track record building, operating, and optimizing LLM or AI applications in production for reliability, latency, and cost
  • Deep understanding of authentication, authorization, security, policy enforcement, and audit controls
  • Experience with observability tooling (tracing, monitoring, alerting, debugging) and incident response, with a demonstrated history of evolving production systems post-launch
  • Familiarity with MCP, tool integrations, agentic-system architectures, or shared AI platform and gateway development (preferred)
  • Experience with React or full-stack application development is a plus; strong collaboration, documentation, and communication skills expected

Compensation

$80-$85/hr - Dependent on fit and experience

Similar Jobs

See more jobs