Skip to content

AI Data Annotator vs AI Evaluator: Roles, Skills, Pay and Where RLHF Fits

Matt Li By Matt Li Co-Founder and Director 10 min read
TL;DR: An AI data annotator labels raw data that has one right answer, such as boxes on images or tags on text; an AI evaluator judges what a model produces and writes down why. On Mercor's job board in October 2026, Bengali document annotation paid $10 to $11 an hour while a clinical AI conversation evaluation paid $100. An RLHF specialist sits between the two, turning side-by-side judgments into training data.

In late 2021, Sama's team in Nairobi read and labeled 150 to 250 passages of toxic text per nine-hour shift for OpenAI, for a take-home wage of about $1.32 to $2 an hour.

In 2025, Google's contractor GlobalLogic started its raters, most of them in the US, at $16 to $21 an hour to judge Gemini's answers. Both jobs put human judgment into a model.

I read 315 open contractor listings on October 5, 2026 to see where one job ends and the other starts.

Key takeaways
  1. 1Meta's Segment Anything model drew 99.1% of the 1.1 billion masks in its own training set, with no human in the loop.
  2. 2xAI cut 500 generalist AI tutors in September 2025 and said it would grow its specialist tutor team tenfold.
  3. 3Of Mercor's hourly listings, 6 titles named annotation or transcription and 49 named evaluation, review, audit, benchmark or AI safety work.
  4. 4OpenAI's InstructGPT labelers agreed with each other only 72.6% of the time on which answer was better.

What is the difference between an AI data annotator and an AI evaluator?

An AI data annotator labels raw data before a model trains on it: a box around a car, a tag on a sentence, a transcript of a scanned page.

An AI evaluator judges output the model has already produced, scoring or ranking it against a rubric and explaining the score.

The annotator works where an item has one right answer. The evaluator works where two good answers can differ.

Our AI training data annotator and AI evaluator and trainer role guides draw the same line. The third role, the RLHF specialist, takes the evaluator's side-by-side judgment and produces it at volume as training data.

Two by two map of AI training roles. Raw data with a label: AI data annotator, drawing boxes, masks, tags and transcripts with one right answer per item. Raw data with written judgment: domain annotator, doing clinical, legal or code labeling and building the label taxonomy. Model output with a label: RLHF specialist, picking the better of two answers and writing the reason and an ideal reply. Model output with written judgment: AI evaluator, scoring output against a rubric and building eval sets.

The labels on job boards blur this. Mercor's RN Annotators listing asks nurses for "annotation, clinical review, model evaluation, and user-feedback investigation" in one job, at $55 to $65 an hour. Read the duties, not the title.

AI data annotator
  • Input: raw images, text, audio, documents
  • Output: a label, box, mask or transcript
  • Checked against a gold answer
  • Speed is measured per item
AI evaluator
  • Input: a model's answers to a prompt
  • Output: a score or ranking plus a written reason
  • Checked by agreement with other raters
  • Judgment matters more than speed

What each role does day to day

The annotator: volume against a guideline

Annotation runs on items per hour. Building Microsoft's COCO image dataset took over 70,000 worker hours for 2.5 million labeled objects in 328,000 images. Tracing object outlines alone needed more than 22 worker hours per 1,000 shapes.

The text side runs on the same clock. The Sama labelers in TIME's investigation read passages of around 100 to over 1,000 words each and sorted them into categories of harm.

Current listings look similar. Mercor's Bengali PDF annotation project asks native speakers to map every region of a scanned page, set its reading order and transcribe it in the original script.

The listing stresses that delivered work is "human-authored throughout," with no parsing model filling in the map.

The evaluator: two answers and a reason

An evaluator's task starts with two outputs from the same prompt. Mercor's Multimodal Image Expert listing pays $30 an hour to compare pairs of AI-generated images with the model names hidden.

Raters weigh prompt fit, anatomy, rendered text and artifacts, then "write a short rationale explaining your choice."

The rationale is the part annotation does not have. Meta asked its Llama 2 raters to mark each choice as significantly better, better, slightly better, or negligibly better, instead of picking a side. Evaluation also runs past training.

Scale AI's Subject Matter Expert posting covers audits, scoring rubrics and gold standards, with no contributor management at all.

Where the RLHF specialist fits

The RLHF specialist does an evaluator's comparison, but as training data rather than a verdict on a release. Each pick between two answers becomes one row of preference data, which trains a reward model or feeds a method such as DPO.

The role covers two stages of the same pipeline:

Four steps where each role touches a model: 1, training data, the annotator labels images, text, audio and documents; 2, fine-tuning, the RLHF specialist writes ideal answers, 27,540 for Llama 2; 3, preference data, the RLHF specialist ranks pairs of answers, over 1 million comparisons for Llama 2; 4, evaluation, the AI evaluator scores the trained model against rubrics and test sets.

The volumes split the roles. Llama 2 used 27,540 hand-written ideal answers, after its authors set aside millions of third-party examples for fewer, better ones. Its preference stage used over 1 million comparisons.

The InstructGPT data behind its headline result, labelers preferring a 1.3 billion parameter model over the 175 billion parameter GPT-3, came from about 40 contractors, hired on Upwork and through Scale AI.

So the RLHF specialist sits closer to the evaluator than to the annotator. Its output sets it apart from evaluation: rows of training data at a steady quality bar, week after week, not a report on how a model performed.

Our guide to annotation for LLM fine-tuning covers the data formats.

How quality is measured in each role

Teams check annotation against an answer key or against other workers. For category labels, COCO sent each image to 8 Mechanical Turk workers at once. The paper found their combined labels had better recall than any single expert.

Metrics like accuracy against gold items and overlap between boxes fit annotation because a right answer exists. Our post on annotation quality metrics covers them.

Evaluation has no answer key, so teams measure agreement instead. The published rates are lower than most buyers expect:

Bar chart of published agreement rates on pairwise preference judgments: GPT-4 judge vs human experts 85% and human expert vs human expert 81% (MT-bench, 2023, ties excluded), InstructGPT held-out labelers 77.3% and training labelers 72.6% (OpenAI, 2022), Anthropic researchers vs crowdworkers 63% (2022).

Anthropic's helpful and harmless assistant paper reported about 63% agreement between its researchers and its crowdworkers.

It also found it easier to get high-quality work from contractors on Upwork paid by the hour than from Mechanical Turk workers paid by the task.

The top two bars carry the bigger warning for evaluators. In the MT-bench study, GPT-4 agreed with human experts 85% of the time, more than the experts agreed with each other.

A model can now do part of the evaluator's job, provided humans keep checking it.

How the listings were counted. I read the 315 listings in the public data of Mercor's explore page on October 5, 2026, and grouped the hourly ones by job title. A title is not the whole job, so treat the split as a snapshot.

Skills and background

An annotator needs speed, a steady hand with the tool and strict guideline discipline. The work moves fast, so small errors repeat across thousands of items.

Language and domain raise the bar: the Bengali PDF project requires native command of the script, and medical labeling goes to clinicians.

An evaluator needs to write. Evaluation listings ask for a reason behind each judgment, like the image listing's rationale, and the specialist ones ask for credentials.

Scale AI's Human Frontier Collective posting says over 90% of its members hold doctorates across 70+ fields. Most work 10 to 25 hours a week on evaluation and research projects.

The coding end is narrower still. Mercor's SWE-Bench Task Auditor wants 3+ years of professional engineering and open-source maintainer history. The auditor has to "detect answer leakage / reward hacking" in benchmark tasks.

That is the coding AI trainer skill set applied to evaluation.

Pay: what the postings say

Pay rises with the amount of judgment in the task. Language decides the floor. Here are the posted hourly ranges, from labeling at the bottom to expert evaluation at the top:

Range chart of posted hourly pay: toxic-text labeling in Kenya in 2022 $1.32 to $2, Bengali document annotation $10 to $11, Bengali AI safety review $16 to $22, generalist AI trainer $25 to $50, rater guidelines writer $45 to $65, US nurse annotation $55 to $65, SWE-Bench task auditor $70 to $90, coding AI trainer $40 to $150.

The Bengali rows show the gap inside one language. Mercor's English and Bengali AI safety listing pays $16 to $22 an hour to red-team chatbots and annotate their failures, about double the document annotation rate.

The same safety project pays $17 to $25 for Vietnamese, Indonesian and Malay speakers.

Platform and roleWorkPosted rate
Outlier (run by Scale AI), generalistEvaluate and rank model reasoning, write rubricsUp to $15/hr
GlobalLogic (for Google), generalist raterRate Gemini answersFrom $16/hr
GlobalLogic, super raterRate Gemini answers, in specialist podsFrom $21/hr
DataAnnotation, generalistAI training tasks$25 to $50/hr
DataAnnotation, software engineerCoding tasks$40 to $150+/hr
Mercor, clinical mental health expertPick the more realistic of two simulated sessions$100/hr
Sources: Outlier, The Guardian (Sept 11, 2025), DataAnnotation, Mercor. Read October 5, 2026.

The Guardian's reporting adds that GlobalLogic's raters earn more than "their data-labeling counterparts in Africa and South America."

For more on how location sets labeling rates, see our annotation costs by country comparison and our review of how DataAnnotation pays.

Demand: which role is growing

Evaluation and specialist work are growing. Generalist labeling is shrinking. The last three years of company moves line up behind that:

Apr 2023
Meta releases SA-1B: a model draws 99.1% of 1.1 billion masks itself
Jun 2025
Meta puts $14.3 billion into Scale AI for a 49% non-voting stake
Jul 2025
Scale AI cuts 14% of staff and 500 contractors, largely in data labeling
Sep 2025
xAI drops 500 generalist tutors, about a third of its annotation team
Jul 2026
Mercor's run rate passes $2 billion, per TechCrunch
Sources: Kirillov et al. (2023), CNBC (June 12, 2025), TechCrunch (July 16, 2025; Sept 13, 2025; July 9, 2026).

The Segment Anything paper shows how far model-assisted labeling has gone. Hand-drawn masks took 34 seconds each at the start of the project and 14 seconds once the model helped. Then the model made the rest.

Scale AI's 2025 cuts hit its data-labeling business, a month after Meta's $14.3 billion investment.

Three rings: 99.1% of SA-1B's 1.1 billion masks were drawn by the model itself; xAI cut about 1 in 3 of its annotation team in 2025; over 90% of Scale AI's Human Frontier Collective experts hold doctorates.

xAI was blunt in its September 2025 notice: it no longer needed "most generalist AI tutor positions" and would hire specialists in STEM, finance, medicine and safety.

TechCrunch reported Mercor's annualized run rate crossing $2 billion in July 2026. Its listings lean the same way: 49 evaluation-type titles against 6 annotation titles on my count.

Career paths: annotator to evaluator

Annotation is the usual way in, and evaluation is the usual way up. Our RLHF specialist guide lists data annotation as a common step up, trading volume for judgment and higher rates.

From there the ladder runs rater, senior specialist, quality lead and guideline owner.

Mercor posts the top rung as a job of its own. Its AI Rater Guidelines Writer listing, a W-2 role at $45 to $65 an hour, turns vague program requirements into rubrics "raters can apply without escalation."

An annotator who already knows a field, such as nursing or law, can skip ahead into the domain expert AI trainer track.

Neither role is safe from automation, but evaluation is safer for now. Google's RLAIF study found AI feedback came within 2 points of human feedback on summarization, 71% against 73% in win rate. Human evaluators judged both results.

Our post on synthetic data vs human annotation covers where machine labels hold up.

Which role you need

Hire annotators when the data has a right answer and the volume is large: images, sensor data, transcripts, document structure. Hire evaluators when the question is whether a model's answer is good, and someone has to say why.

Hire an RLHF specialist when those judgments need to become training data on a schedule.

Heatmap of which role fits which task. Boxes, masks and transcripts: annotator best, RLHF specialist limited. Clinical or legal labeling: annotator best, others limited. Ranking two chatbot answers: RLHF specialist best, evaluator good, annotator limited. Writing ideal answers: RLHF specialist best, evaluator good. Scoring a release on a rubric and writing rubrics and guidelines: evaluator best, RLHF specialist good.

Most teams need more than one. A team fine-tuning a model on support tickets might label the tickets first, then rank the model's replies, then evaluate each release. Our guide to building annotation teams covers the staffing ratios.

Hire annotators and AI evaluators from Asia

Second Talent matches you with vetted data annotation specialists, AI evaluation specialists and RLHF specialists in 24 hours, at 50-70% below US cost.

Every hire carries a 90-day, one-time replacement guarantee, and we employ them through our own EOR in 9 Asian markets.

The full AI training specialists range includes coding and multilingual trainers. Tell us what you are training and we will send a shortlist.

Frequently Asked Questions

Is an AI trainer the same as an AI evaluator?

In most listings, yes. Platforms use "AI trainer" or "AI tutor" for the whole family of tasks. Evaluator tends to mean judging finished output, while trainer also covers writing ideal answers and prompts.

Do AI evaluators need to code?

Only for coding projects. General evaluation asks for clear writing and a domain. Code evaluation, such as auditing SWE-Bench tasks, asks for years of professional engineering.

Can one contractor do both jobs?

Yes, and many listings combine them. The skills differ, though: annotation rewards speed under a fixed guideline, while evaluation rewards written reasoning. Test for each skill on its own.

Hiring developers in Southeast Asia?

Get Cost Guide
Matt Li

Written by

Matt Li is a tech-driven entrepreneur with deep expertise in global talent strategy, digital experience optimization, e-commerce, and Web3 innovation. He is the Co-Founder of Second Talent, a US-based company that connects businesses with top-tier tech professionals worldwide. Since launching the company in 2024, Matt has led its growth by leveraging technology to streamline remote hiring and scale distributed teams. With a background spanning product, operations, and innovation, Matt brings a cross-disciplinary perspective to the evolving digital economy. His work sits at the intersection of global talent, emerging technology, and scalable digital transformation.

More posts by Matt Li →

How would you like to talk?

WhatsApp us Prefer texting at your own pace? Just hit us up on WhatsApp. We promise no spam and a hassle-free experience.

Loading available times…