LLM - AI Quality Analyst (Personalization) - English
Location: Remote (USA Approved States Only)
Eligible States: Approved U.S. States only (Excluding Texas and Illinois)
Contract: Short-Term (4 Months)
Openings: 350
Start Date: Immediate
Rate: $30/hour USD
Experience: 1+ Year## About Turing
Turing, based in San Francisco, California, is a leading research accelerator for frontier AI labs and a trusted partner for enterprises deploying advanced AI systems.
Turing accelerates AI research through high-quality data, advanced training pipelines, and top AI researchers specializing in coding, reasoning, STEM, multilinguality, multimodality, and agents. Turing also helps enterprises transform AI from proof of concept into reliable systems that deliver measurable business impact.
Role Overview
As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini. You will assess how effectively the model uses information from past Gemini conversations, Gmail, Google Search, and YouTube activity to generate relevant and helpful responses.
This role combines creativity and analytical thinking. You will create prompts based on your personal experiences and evaluate responses across:
- Grounding
- Integration
- Helpfulness
Key Qualifications
- Ability to read and write English with a high degree of comprehension.
- Willingness to use your primary personal Google account (not a testing account) and enable personal data sources.
- Full-time availability in your local time zone.
- Ability to work within a global 24-hour operations team.
- Exceptional analytical thinking and evaluation skills.
- Experience designing creative multi-turn prompts based on personal context.
- Understanding of personalization concepts, including incorrect personalization, poor inferences, and forced connections.
- Strong attention to detail when reviewing Side-by-Side (SxS) model responses.
- Excellent written communication skills with the ability to write clear, structured rationales.
- Ability to provide detailed annotations and constructive feedback.
- Strong communication and collaboration skills.
- Self-motivated and able to work independently in a remote environment.
- Desktop or laptop with a reliable internet connection.
Responsibilities
- Design and execute multi-turn conversational prompts (typically 1-5 turns) requiring AI to use personal information and experiences.
- Evaluate personalized model responses against the original prompt intent.
- Review responses for Grounding issues, flawed inferences, and hallucinations.
- Assess Integration quality to ensure natural use of personal information without overnarrating.
- Compare and stack-rank Side-by-Side (SxS) model responses.
- Write clear, defensible rationales referencing specific conversation turns.
- Extract and verify Debug Info to confirm the proper use of chat summaries and data sources.
- Maintain strict data hygiene by deleting evaluation conversations to prevent impact on future chat history.
Education & Experience
Required
- BS/BA degree or equivalent experience in:- Policy
- Law
- Ethics
- Linguistics
- Journalism
- Computer Science
- Related analytical fields
Preferred
- Experience in:- Data Annotation
- AI Quality Evaluation
- Content Moderation
- Related roles
Availability
- Minimum 4 hours per day
- Minimum 30 hours per week
- 4 hours overlap with PST
Commitment Options:
- 30 Hours/Week
- 40 Hours/Week
Vetting Process
All 3 steps are required:
- Screener
- 1 of 3 Assessments (Mandatory)
- Language Vetting
Evaluation Process
- Shortlisted candidates receive a Job Interest Form.
- After profile review, an assessment will be shared and must be completed within 24 hours.
- Successful candidates will be contacted regarding pre-onboarding requirements.
You should be proficient in:
- AI Content Evaluation
- AI Evaluator
- AI Quality Analyst
- AI Quality Evaluation
- Analytical Thinking
- Annotations
- Content Moderation
- Data Annotation
- English Proficiency
- Gemini
- Generative AI
- Grounding
- Integration
- Large Language Models
- LLM Analyst
- LLM Evaluation
- Personalization
- Prompt Design
- Prompt Engineering
- Quality Assurance
- Remote Contractor
- Response Ranking
- Side-by-Side Evaluation
- SxS Review
- Google Account