Job Title: Overwatch - Observability and Evaluation Engineer
Location: Charlotte, North Carolina
duration 12+ months
40 hours a week
100% onsite
VISA TYPE (any as long as eligible to work)
(NOTE: local candidates will be top priority if you can find anyone; whoever is willing to relocate and can go onsite, make sure they are strong)
We are seeking an Overwatch Observability and Evaluation Engineer to implement telemetry, tracing, dashboards, evaluation suites, service objectives, runbooks, and release-readiness evidence for high-priority AI agent releases. This role is focused on ensuring agentic AI systems are measurable, reliable, production-ready, and continuously monitored.
Key Responsibilities
" Design and implement observability frameworks for LLM applications and AI agent workflows.
" Build telemetry, tracing, logging, metrics, and dashboarding capabilities for production AI systems.
" Develop evaluation suites to measure prompt quality, model behaviour, task success, groundedness, latency, reliability, and safety.
" Define and track SLOs, SLIs, service readiness criteria, and operational health indicators.
" Create automated test and evaluation workflows for model, prompt, and agent performance analysis.
" Build runbooks and operational readiness evidence for production releases.
" Partner with AI engineering, platform, DevOps, product, and risk teams to support priority agent launches.
" Analyse production signals, identify performance issues, and recommend improvements for reliability and user outcomes.
Must-Have Skills
LLM Evaluation, AI Agent Evaluation, Observability, Tracing, Telemetry, Metrics, Dashboards, SLOs, Runbooks, Test Automation, Prompt Performance Analysis, Model Performance Analysis, Python, Production Operations
Preferred Skills
LangSmith, OpenTelemetry, Grafana, Prometheus, Datadog, CloudWatch, Azure Monitor, AI Reliability Engineering, LLMOps, MLOps, Incident Response, Release Readiness
Qualifications
" Bachelor's degree in Computer Science, Data Science, AI/ML, Engineering, or a related field.
" Experience implementing observability, evaluation, or monitoring for AI/ML, LLM, or agent-based systems.
" Strong Python scripting and automation skills.
" Ability to support production-grade AI releases with clear documentation, metrics, and operational readiness discipline.