Practical experience with AI/LLM evaluation frameworks (e.g., Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, TruLens) - building eval suites, scoring rubrics, and golden datasets. Working knowledge of eval metrics for generative AI: hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, BLEU/ROUGE/semantic similarity where applicable.
Numbers & Facts
Location
Bellevue, WA
Description
Design and execute test plans for AI/ML-driven features, including model outputs, prompts, and integrated application behavior
Build and maintain automated test suites covering functional, regression, integration, and API testing
Evaluate model outputs for accuracy, consistency, bias, hallucination, and edge-case failures
Develop evaluation frameworks and golden datasets/test cases to benchmark model performance over time
Test prompt engineering changes, model version upgrades, and fine-tuning outputs for regressions
Perform adversarial and red-team style testing to surface safety, security, and robustness issues
Validate data pipelines feeding into AI models (data quality, schema, drift detection)
Collaborate with data scientists/ML engineers to define acceptance criteria and quality metrics for models
Test latency, scalability, and reliability of AI services under load
Contribute to CI/CD pipelines, integrating automated and model-evaluation tests
Hands-on experience testing LLM-based products (chatbots, copilots, RAG systems, AI agents) - designing test cases for non-deterministic, generative outputs
Practical experience with AI/LLM evaluation frameworks (e.g., Ragas, DeepEval, LangSmith, Promptfoo, OpenAI Evals, TruLens) - building eval suites, scoring rubrics, and golden datasets
Working knowledge of eval metrics for generative AI: hallucination rate, faithfulness/groundedness, relevance, answer correctness, toxicity/bias scoring, BLEU/ROUGE/semantic similarity where applicable
Experience with prompt regression testing - validating prompt changes and model/version upgrades against baseline eval sets
Strong proficiency in Python for writing eval scripts, test harnesses, and data validation logic
Familiarity with SQL and data validation techniques
The environment is primarily Microsoft Azure-based, with Snowflake also playing a key role. Their GenAI initiatives largely leverage OpenAI and Claude models.