Preferred Qualifications:** + 6+ years of experience in applied science, machine learning, evaluation systems, or related technical fields + Strong experience designing evaluation methodologies, experiments, or measurement systems for complex intelligent or distributed systems + Experience analyzing large-scale production or experimental data to derive actionable insights and drive product or system improvements + Strong coding and prototyping skills in Python or similar languages, with the ability to work closely with engineering teams on production-facing systems + Demonstrated ability to lead cross-team technical direction through scientific depth, influence, and strong problem framing + Advanced degree in Computer Science, Machine Learning, Statistics, Applied Mathematics, or related field + Experience building or evaluating LLM- or agent-based systems in production + Familiarity with agent frameworks such as LangChain, LangGraph, OpenAI SDK, or equivalent orchestration frameworks + Experience with evaluation frameworks for AI systems, including benchmarking, regression analysis, and human-in-the-loop assessment + Experience with observability systems, telemetry analysis, or distributed tracing data in large-scale environments + Background in AI safety, guardrails, and responsible AI measurement + Experience with experimentation platforms, causal inference, or statistical methods for product and model evaluation + Experience working with cloud-scale monitoring platforms such as Azure Monitor / Application Insights or equivalent Applied Sciences IC6 - The typical base pay range for this role across the U.S. is USD $165,600 - $296,400 per year. **Technical Focus Areas** + Evaluation science for agent and multi-agent systems: offline, online, and continuous evals; benchmark design; synthetic data; task success measurement + Agent and multi-agent architectures: planners, tool use, memory, orchestration, and coordination patterns + Applied machine learning and statistical methods for behavioral analysis, anomaly detection, experimentation, and regression detection + Observability data for AI systems: traces, logs, metrics, evaluations, and cost/performance signals + Safety and responsible AI signals: policy compliance, risk detection, auditability, and safe logging + Benchmarking and experimentation for agent systems, including A/B tests, canaries, and staged rollouts + Explainability and diagnosis for complex agent workflows and model-driven decision paths **Qualifications** **Required Qualifications:** + Bachelor's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python + OR equivalent experience.