Design and build diverse positive-example evaluation datasets - develop principled labeling taxonomies and sampling strategies that maximize coverage of real-world content diversity (across languages, formats, content types, and adversarial patterns), leveraging LLMs as tools to surface gaps and expand coverage systematically. Define and own the evaluation methodology for AI-powered safety models - establish frameworks for measuring continuous recall capability across risk severity tiers; define launch criteria, regression thresholds, and ongoing monitoring requirements so that no model ships or degrades without clear, evidence-based quality signals.