TL;DR: An AI data annotator labels raw data that has one right answer, such as boxes on images or tags on text; an AI evaluator judges what a model produces and writes down why. On Mercor's job board in October 2026, Bengali document annotation paid $10 to $11 an hour while a clinical AI conversation evaluation paid $100. An RLHF specialist sits between the two, turning side-by-side judgments into training data.
In late 2021, Sama's team in Nairobi read and labeled 150 to 250 passages of toxic text per nine-hour shift for OpenAI, for a take-home wage of about $1.32 to $2 an hour.
In 2025, Google's contractor GlobalLogic started its raters, most of them in the US, at $16 to $21 an hour to judge Gemini's answers. Both jobs put human judgment into a model.
I read 315 open contractor listings on October 5, 2026 to see where one job ends and the other starts.
- 1Meta's Segment Anything model drew 99.1% of the 1.1 billion masks in its own training set, with no human in the loop.
- 2xAI cut 500 generalist AI tutors in September 2025 and said it would grow its specialist tutor team tenfold.
- 3Of Mercor's hourly listings, 6 titles named annotation or transcription and 49 named evaluation, review, audit, benchmark or AI safety work.
- 4OpenAI's InstructGPT labelers agreed with each other only 72.6% of the time on which answer was better.
What is the difference between an AI data annotator and an AI evaluator?
An AI data annotator labels raw data before a model trains on it: a box around a car, a tag on a sentence, a transcript of a scanned page.
An AI evaluator judges output the model has already produced, scoring or ranking it against a rubric and explaining the score.
The annotator works where an item has one right answer. The evaluator works where two good answers can differ.
Our AI training data annotator and AI evaluator and trainer role guides draw the same line. The third role, the RLHF specialist, takes the evaluator's side-by-side judgment and produces it at volume as training data.

The labels on job boards blur this. Mercor's RN Annotators listing asks nurses for "annotation, clinical review, model evaluation, and user-feedback investigation" in one job, at $55 to $65 an hour. Read the duties, not the title.
- Input: raw images, text, audio, documents
- Output: a label, box, mask or transcript
- Checked against a gold answer
- Speed is measured per item
- Input: a model's answers to a prompt
- Output: a score or ranking plus a written reason
- Checked by agreement with other raters
- Judgment matters more than speed
What each role does day to day
The annotator: volume against a guideline
Annotation runs on items per hour. Building Microsoft's COCO image dataset took over 70,000 worker hours for 2.5 million labeled objects in 328,000 images. Tracing object outlines alone needed more than 22 worker hours per 1,000 shapes.
The text side runs on the same clock. The Sama labelers in TIME's investigation read passages of around 100 to over 1,000 words each and sorted them into categories of harm.
Current listings look similar. Mercor's Bengali PDF annotation project asks native speakers to map every region of a scanned page, set its reading order and transcribe it in the original script.
The listing stresses that delivered work is "human-authored throughout," with no parsing model filling in the map.
The evaluator: two answers and a reason
An evaluator's task starts with two outputs from the same prompt. Mercor's Multimodal Image Expert listing pays $30 an hour to compare pairs of AI-generated images with the model names hidden.
Raters weigh prompt fit, anatomy, rendered text and artifacts, then "write a short rationale explaining your choice."
The rationale is the part annotation does not have. Meta asked its Llama 2 raters to mark each choice as significantly better, better, slightly better, or negligibly better, instead of picking a side. Evaluation also runs past training.
Scale AI's Subject Matter Expert posting covers audits, scoring rubrics and gold standards, with no contributor management at all.
Where the RLHF specialist fits
The RLHF specialist does an evaluator's comparison, but as training data rather than a verdict on a release. Each pick between two answers becomes one row of preference data, which trains a reward model or feeds a method such as DPO.
The role covers two stages of the same pipeline:

The volumes split the roles. Llama 2 used 27,540 hand-written ideal answers, after its authors set aside millions of third-party examples for fewer, better ones. Its preference stage used over 1 million comparisons.
The InstructGPT data behind its headline result, labelers preferring a 1.3 billion parameter model over the 175 billion parameter GPT-3, came from about 40 contractors, hired on Upwork and through Scale AI.
So the RLHF specialist sits closer to the evaluator than to the annotator. Its output sets it apart from evaluation: rows of training data at a steady quality bar, week after week, not a report on how a model performed.
Our guide to annotation for LLM fine-tuning covers the data formats.
How quality is measured in each role
Teams check annotation against an answer key or against other workers. For category labels, COCO sent each image to 8 Mechanical Turk workers at once. The paper found their combined labels had better recall than any single expert.
Metrics like accuracy against gold items and overlap between boxes fit annotation because a right answer exists. Our post on annotation quality metrics covers them.
Evaluation has no answer key, so teams measure agreement instead. The published rates are lower than most buyers expect:

Anthropic's helpful and harmless assistant paper reported about 63% agreement between its researchers and its crowdworkers.
It also found it easier to get high-quality work from contractors on Upwork paid by the hour than from Mechanical Turk workers paid by the task.
The top two bars carry the bigger warning for evaluators. In the MT-bench study, GPT-4 agreed with human experts 85% of the time, more than the experts agreed with each other.
A model can now do part of the evaluator's job, provided humans keep checking it.
Skills and background
An annotator needs speed, a steady hand with the tool and strict guideline discipline. The work moves fast, so small errors repeat across thousands of items.
Language and domain raise the bar: the Bengali PDF project requires native command of the script, and medical labeling goes to clinicians.
An evaluator needs to write. Evaluation listings ask for a reason behind each judgment, like the image listing's rationale, and the specialist ones ask for credentials.
Scale AI's Human Frontier Collective posting says over 90% of its members hold doctorates across 70+ fields. Most work 10 to 25 hours a week on evaluation and research projects.
The coding end is narrower still. Mercor's SWE-Bench Task Auditor wants 3+ years of professional engineering and open-source maintainer history. The auditor has to "detect answer leakage / reward hacking" in benchmark tasks.
That is the coding AI trainer skill set applied to evaluation.
Pay: what the postings say
Pay rises with the amount of judgment in the task. Language decides the floor. Here are the posted hourly ranges, from labeling at the bottom to expert evaluation at the top:

The Bengali rows show the gap inside one language. Mercor's English and Bengali AI safety listing pays $16 to $22 an hour to red-team chatbots and annotate their failures, about double the document annotation rate.
The same safety project pays $17 to $25 for Vietnamese, Indonesian and Malay speakers.
| Platform and role | Work | Posted rate |
|---|---|---|
| Outlier (run by Scale AI), generalist | Evaluate and rank model reasoning, write rubrics | Up to $15/hr |
| GlobalLogic (for Google), generalist rater | Rate Gemini answers | From $16/hr |
| GlobalLogic, super rater | Rate Gemini answers, in specialist pods | From $21/hr |
| DataAnnotation, generalist | AI training tasks | $25 to $50/hr |
| DataAnnotation, software engineer | Coding tasks | $40 to $150+/hr |
| Mercor, clinical mental health expert | Pick the more realistic of two simulated sessions | $100/hr |
The Guardian's reporting adds that GlobalLogic's raters earn more than "their data-labeling counterparts in Africa and South America."
For more on how location sets labeling rates, see our annotation costs by country comparison and our review of how DataAnnotation pays.
Demand: which role is growing
Evaluation and specialist work are growing. Generalist labeling is shrinking. The last three years of company moves line up behind that:
The Segment Anything paper shows how far model-assisted labeling has gone. Hand-drawn masks took 34 seconds each at the start of the project and 14 seconds once the model helped. Then the model made the rest.
Scale AI's 2025 cuts hit its data-labeling business, a month after Meta's $14.3 billion investment.

xAI was blunt in its September 2025 notice: it no longer needed "most generalist AI tutor positions" and would hire specialists in STEM, finance, medicine and safety.
TechCrunch reported Mercor's annualized run rate crossing $2 billion in July 2026. Its listings lean the same way: 49 evaluation-type titles against 6 annotation titles on my count.
Career paths: annotator to evaluator
Annotation is the usual way in, and evaluation is the usual way up. Our RLHF specialist guide lists data annotation as a common step up, trading volume for judgment and higher rates.
From there the ladder runs rater, senior specialist, quality lead and guideline owner.
Mercor posts the top rung as a job of its own. Its AI Rater Guidelines Writer listing, a W-2 role at $45 to $65 an hour, turns vague program requirements into rubrics "raters can apply without escalation."
An annotator who already knows a field, such as nursing or law, can skip ahead into the domain expert AI trainer track.
Neither role is safe from automation, but evaluation is safer for now. Google's RLAIF study found AI feedback came within 2 points of human feedback on summarization, 71% against 73% in win rate. Human evaluators judged both results.
Our post on synthetic data vs human annotation covers where machine labels hold up.
Which role you need
Hire annotators when the data has a right answer and the volume is large: images, sensor data, transcripts, document structure. Hire evaluators when the question is whether a model's answer is good, and someone has to say why.
Hire an RLHF specialist when those judgments need to become training data on a schedule.

Most teams need more than one. A team fine-tuning a model on support tickets might label the tickets first, then rank the model's replies, then evaluate each release. Our guide to building annotation teams covers the staffing ratios.
Hire annotators and AI evaluators from Asia
Second Talent matches you with vetted data annotation specialists, AI evaluation specialists and RLHF specialists in 24 hours, at 50-70% below US cost.
Every hire carries a 90-day, one-time replacement guarantee, and we employ them through our own EOR in 9 Asian markets.
The full AI training specialists range includes coding and multilingual trainers. Tell us what you are training and we will send a shortlist.
Frequently Asked Questions
Is an AI trainer the same as an AI evaluator?
In most listings, yes. Platforms use "AI trainer" or "AI tutor" for the whole family of tasks. Evaluator tends to mean judging finished output, while trainer also covers writing ideal answers and prompts.
Do AI evaluators need to code?
Only for coding projects. General evaluation asks for clear writing and a domain. Code evaluation, such as auditing SWE-Bench tasks, asks for years of professional engineering.
Can one contractor do both jobs?
Yes, and many listings combine them. The skills differ, though: annotation rewards speed under a fixed guideline, while evaluation rewards written reasoning. Test for each skill on its own.





