Skip to main content

Volga Partners

AI Evaluation

How Do You Know a Human Graded It?

A buyer’s checklist for the trust and rigor behind AI evaluation work.

In short

Human evaluation is only worth paying for if a human actually did it, and that’s hard to verify after the fact. Here’s what a trustworthy partner does to prove it, and the questions to ask before you trust their data.

When a company buys human evaluation for its AI models, it’s buying judgment: a person reading a model’s output and deciding whether it’s accurate, useful, safe, or well made. Judgment is also the hardest thing to verify after the fact. A finished dataset looks the same whether a trained specialist produced it with care or someone pasted the task into a chatbot and forwarded the reply.

Until recently, that second possibility was a fringe concern. Now it’s one of the first things careful buyers raise, and with good reason. If you’re paying for human judgment and a model quietly supplies it, the evaluation no longer measures anything, and there is no point. In short, you’ve asked one model to grade another and called it human review.

No single safeguard rules this out. What protects you is a set of overlapping quality assurance practices that make cut corners visible, whether the shortcut is a chatbot or plain carelessness. They’re worth walking through, because they double as a checklist for any human evaluation provider you’re weighing.

How do you stop raters from using AI tools to do the work?

It starts with where the work happens. A large share of our work runs on the client’s own platform, where the workspace is visible and access is controlled. Much of the rest is recorded live, so a reviewer can easily retrace how someone reached an answer rather than seeing only what they submitted. Automated work tends to give itself away in the process: its pace and the absence of hesitations and second-guessing that mark a person thinking. Our team leads watch for those signs as part of normal oversight, and completed work also undergoes AI-detection checks.

None of these settles the question alone. A recording is useless if no one reviews it, and detectors miss some cases while flagging innocent ones. All of this, layered together, makes machine-written answers hard to slip through unnoticed, and they raise the cost of trying high enough that most people won’t.

How do you measure and report quality?

Confirming that a person did the work is only the first bar. The next is whether the work holds up. We check it by having more than one person evaluate the same item independently, often three or four for anything contested or high stakes. How closely they agree (commonly known as inter-rater agreement) is a useful signal. Disagreement is even more useful: when reviewers split, the item goes back for another look, is discussed, and is settled on a single answer before it’s marked. Those splits tend to cluster around the exact places where the guidelines are vague, so they also serve as an early warning about what needs clarifying.

Where the checking happens depends on the setup. On a client’s platform, we use that platform’s QA tooling. For our own tools, our in-house quality assurance team runs it. When the work turns on a point of language that a general reviewer might miss, a senior linguist for that language reviews and rectifies it. Two things always stay constant: quality is checked while the dataset is being built, not only at the end, and the whole set is reviewed once more before it ships.

How do you handle ambiguous cases and unclear guidelines?

How an operation handles the obvious items tells you very little. What matters is the case the guidelines never anticipated: the ambiguous one that sits between two categories, or the format nobody expected. Anyone can label clear examples. The edges are where quality is decided.

The tempting move is to guess and keep going. It’s also the worst choice, because a private guess gets repeated inconsistently by everyone who later encounters the same case.

Our rule is to surface it. The rater sets the item aside, finishes the cases the guidelines already cover, and brings the outlier back to the team. We decide how that kind of case should be handled, then write the decision into the annotation guidelines so it never gets argued twice. The next person who meets it, and someone usually does, finds an answer waiting. Over the life of a project, the guideline stops being a static brief and becomes a record of every hard call the team has already reasoned through.

How do you keep judgments consistent over time and across languages?

All of this rests on people working from the same instructions, which is harder than it sounds because the instructions change and the people don’t stay the same. Raters improve, drift, arrive, and leave. Every language you add to a multilingual evaluation is another place where one team’s reading of a rule can drift from another’s.

We handle it by keeping a single set of instructions that everyone works from. The guidelines are written and maintained in English, and everyone we bring on, regardless of language, must be able to work from them in English. A rater delivering in Czech, Hindi, or Romanian is calibrated against the same definitions as everyone else, not a loose translation. Working from English definitions doesn’t flatten nuance, either, because the senior linguist review described above catches points a general reviewer would miss. Different trainers run the sessions, but they teach the same requirements to achieve the same result.

Consistency also has to be maintained while the work is live. The same parallel review that measures quality is what catches drift early, correcting a reviewer who’s starting to slide before it becomes a batch you have to redo. A final pass before delivery confirms that the finished set still matches what the client asked for.

How do you train raters on your quality framework?

The guidelines are not a document we hand over at kickoff and forget. They are revised as the work raises new questions, sometimes weekly, and each meaningful change comes with a training session for the people on the project. The retraining is what makes the change count; an update that stays in the document only widens the gap between the written rule and the work.

None of this is exotic. It’s a stack of unglamorous practices (recorded work, independent review, resolved disagreements, a living guideline, retraining when it changes) that only holds up if a partner runs all of them consistently, even when no one is watching.

That’s what to probe when you evaluate one. When you buy human evaluation, you’re buying a person’s judgment, and you have every right to see how it’s safeguarded. A partner who has done the work can answer in specifics and will probably be glad you asked. When the specifics aren’t there, that tells you something too.

Ask us these five questions.

We can walk you through how each safeguard runs on a live program — the platforms, the review layers, and the guidelines that grew out of real edge cases.

Talk to our team