Experience with at least one of: agent frameworks (LangGraph, DSPy, custom), eval frameworks (Inspect, Braintrust, OpenAI evals, custom), tracing tools (Langfuse, OpenTelemetry), or RLHF / preference data pipelines. Observability and tracing: span-level visibility into agent steps, prompt rendering, tool calls, and intermediate reasoning - built so a radiologist or a model owner can debug a failure mode in minutes, not hours.