AI Evaluation and Rater Jobs: Entry Points and Pay
Every figure on this page is read from 29 live listings on this board, 1 of which publish a rate. Updated 05 September 2026. See the full pay dataset.
Evaluation roles are the core of AI training work: judge output, rank it, explain why.
What these roles pay
| Measure | Figure |
|---|---|
| Median advertised rate | $55/hr |
| Range across listings | $45 to $65/hr |
| Listings with a published rate | 1 of 29 |
Who is hiring
- Turing: 26 live roles
- Mercor: 1 live role
- micro1: 1 live role
- Terac: 1 live role
Where you can work from
These listings are open to applicants in 58 countries. The most represented are United States, Bangladesh, Hong Kong, India, Indonesia, Japan. Check your own country with the eligibility checker.
What the work actually involves
Taken from the live listings themselves:
- About Turing: Based in San Francisco, California, Turing is the world’s leading research accelerator for frontier AI labs and a trusted partner for global enterprises deploying advanced AI systems.
- Role Overview: As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini.
- Role Overview: As an AI Quality Analyst, you will evaluate a new personalization feature for Gemini.
Typical responsibilities
Pulled from the live listings in this category:
- Guide research and program teams to close the gap between ambiguous program requirements and rater-ready instructions across a range of subject-matter domains.
- Design clear, non-contradictory rating guidelines and rubrics that raters can apply consistently, including under edge cases.
- Evaluate draft guidelines for ambiguity, internal contradiction, and coverage gaps, and revise until raters can apply them without escalation.
- Translate program specifications from domains such as finance, retail, insurance, legal, and sports into precise, discipline-specific rater instructions.
- Collaborate with other subject matter experts and program leads to ensure consistency and accuracy across guideline sets.
- Design and implement advanced methodologies for evaluating AI system safety, focusing on ethical jailbreaks, LLM red teaming, prompt injection, and tool-use abuse scenarios.
What employers ask for
- Demonstrated ability to work across multiple subject-matter domains (e.g., finance, retail, insurance, legal, sports) and translate domain-specific nuance into clear, unambiguous instructions.
- Strong track record of resolving ambiguity and contradiction in written specifications, able to point to concrete before/after examples.
- Demonstrable career progression.
- Ability to engage reliably for at least 35 hours/week during weekdays.
- Strong written communication skills and the ability to explain complex or nuanced guidance clearly and precisely.
- 2+ years of expertise in adversarial machine learning, LLM red teaming, AI safety evaluation, or a closely related security domain
Which platform to start with
Turing currently lists the most roles in this category (26 of 29). Platforms differ mainly in eligibility and how they screen, not in the kind of work. See which platforms hire in your country before you spend time on applications.
How to apply
Applications are completed on the hiring platform, not on this site. Most of these platforms screen with a short qualification step before routing you to paid work, so a complete profile matters more than a long cover letter. Applying is free and none of these platforms charge a fee to start.
Live ai evaluation and rater jobs roles
- AI Quality Analyst - English via Turing
- AI Quality Analyst (Gemini) - Chinese via Turing
- AI Quality Analyst (Personalization) - Arabic via Turing
- AI Quality Analyst (Personalization) - Dutch via Turing
- AI Quality Analyst (Personalization) - German via Turing
- AI Quality Analyst (Personalization) - Hindi via Turing
- AI Quality Analyst (Personalization) - Indonesian via Turing
- AI Quality Analyst (Personalization) -Japanese via Turing
- AI Quality Analyst (Personalization) - Korean via Turing
- AI Quality Analyst (Personalization) - Polish via Turing
- AI Quality Analyst (Personalization) - Portuguese via Turing
- AI Quality Analyst (Personalization) - Russian via Turing
- AI Quality Analyst (Personalization) - Spanish via Turing
- AI Quality Analyst (Personalization) - Thai via Turing
- AI Quality Analyst (Personalization) - Turkish via Turing
- AI Quality Analyst (Personalization) - Vietnamese via Turing
- AI Quality Analyst - Portuguese (Portugal) via Turing
- AI Rater Guidelines Writer (Linguist / Instructional Designer) via Mercor · $45-$65/hr
- Business Account Rater -Bilingual (KO, DE, FR) via Turing
- Business Account Rater - English US via Turing
- Business Account Rater - Italian via Turing
- Business Account Rater - Portuguese (BR) via Turing
- Business Account Rater - Spanish via Turing
- AI Jailbreak & Prompt-Injection Security Expert via micro1
- Personal Account Rater -English via Turing