AI evaluation

Human evaluation for the AI you ship

Verified domain professionals judge your model's outputs against the criteria that matter in their field. Structured comparison, not anonymous star ratings.

On a 30-minute call we scope what good looks like in your domain, then score 20 of your model's outputs with verified experts and send back the results. No cost, no commitment.

Source: "Take one tablet twice daily with food. Do not exceed 4 tablets in 24 hours."

Output A

Tome un comprimido dos veces al día con comida. No exceda 4 comprimidos en 24 horas.

92
Verified expert review Licensed pharmacist, native speaker
Output B

Tome un comprimido dos veces al día. Máximo 4 comprimidos por día.

41

What the expert caught

✓ Fluent translation ✗ Dropped "with food" ✗ "Per day" is not "per 24 hours"
B reads naturally, so a linguist would pass it. Losing the food instruction changes absorption, and "per day" invites a second dose before midnight. Neither is a language error.

Your model is only as good as the people evaluating it

Workforce size does not predict evaluation quality. Domain expertise, verified identity and representative coverage do. A rater who cannot judge the field will pass outputs that a professional would reject on sight, and the scores will look clean the whole way through.

The risk of evaluating with the wrong humans

  • Raters who cannot judge the domain
  • Flat scores that hide real gaps
  • Failures your users find first

What experts catch that raters miss

Four domains where fluent, plausible output is routinely wrong in ways only a practitioner sees.

Reviewed by a verified design and build professional

Model output: generated floor plan revision

Corridor width reduced to 900mm to accommodate the additional storage bay. All room dimensions preserved.

✓ Dimensions internally consistent ✓ Format followed ✗ Fails code compliance
Egress width is under minimum. It reads fine unless you have actually built to this code.

Reviewed by verified product designers

Model output: two generated interface variants

Variant A and variant B of a checkout summary screen, same content, different visual hierarchy.

✓ B preferred, 7 of 10 experts ✓ Position randomised
B's hierarchy survives on a small screen. A collapses into one grey block the moment it is on a phone.

Reviewed by a verified finance professional

Model output: summarised financial analysis

Revenue grew 34% year on year, driven by strong subscription performance in the fourth quarter.

✓ Arithmetic correct ✓ Format followed ✗ Misreads the metric
It is treating deferred revenue as income. The numbers add up and the conclusion is wrong.

Reviewed by a verified native professional

Model output: localised product copy

A customer-facing onboarding message, translated from English into the target market language.

✓ Grammatically correct ✓ Terminology consistent ✗ Wrong register
Technically fine, and far too formal. Nobody speaks to a customer this way here.

Illustrative examples of the failure modes we test for.

Where most evaluation panels are thinnest

Model quality is not evenly distributed across languages, and neither is the supply of people qualified to judge it. Our depth is in the markets the large vendors cover last.

Southeast Asia

Indonesia, Malaysia, Vietnam and the Philippines. Verified professionals evaluating in Bahasa Indonesia, Bahasa Melayu and English, in the domains they actually work in.

Wider Asia

Recruited to brief where the panel does not already reach, using the same verification process rather than a different standard for harder markets.

Latin America

Spanish and Portuguese evaluation with native professionals who hold the domain expertise your content needs, not generalist linguists.

US, UK and Europe

Established coverage across our four domains, useful as the baseline you compare emerging-market performance against.

Language fluency and domain expertise are recruited together. One without the other is why evaluation misses things.

A method you can trust

Structured comparison, not vague ratings. The design choices below are what make the results hold up when two models are close.

Forced-choice pairwise

Experts agree far more on whether A beats B than on what a 4 out of 5 means, so inter-rater reliability is much higher without heavy rubric training.

Per-criterion checks

Factually correct? Format followed? Safe? Binary checks give you an absolute quality bar for regression testing, which pairwise alone cannot.

Position randomised

Left and right order flips for every rater. Position bias is large, and left unhandled it manufactures a winner that does not exist.

Ties allowed

Raters can mark a tie or flag both outputs as unusable. Forcing a pick between two bad answers invents signal you will later act on.

What an evaluator sees

Fintech assistant eval Task 4 of 14
EnglishBahasa IndonesiaKorean+ more
Prompt: "I want to withdraw from my e-wallet but there's an admin fee. Is that normal?"
Response A

Admin fees on withdrawals are standard and commonly applied across providers...

Response B

Totally normal! Most e-wallets charge a fee for that, it's nothing to worry about...

Prefer APrefer BSameBoth bad
Accuracy●●●●○
Naturalness●●●○○
Cultural fit●●○○○
Why? Type your reasoning, or record a voice note Flag this task

Model names never reach the evaluator. A and B are flipped per rater and the flip is recorded, so position bias cannot manufacture a winner.

Not all failures are equal

Every criterion check is graded by severity, so a style complaint never outranks a clinical or compliance error. You see what blocks a release, and what is worth knowing but not worth holding it back for.

Critical Wrong in a way that causes harm, breaks a rule, or fails compliance. Blocks release.
Major Changes the meaning or the decision a user would make. Needs a fix before it ships.
Minor Register, tone or style. Worth knowing, rarely worth blocking a release for.

In the example above, dropping "with food" is critical. It would pass any review that only counted language errors.

How it works

1

Scope and rubric

We define what good looks like in your domain, with you.

2

Recruit experts

Verified professionals matched to the domain and language.

3

Run evaluation

Pairwise comparison plus per-criterion checks.

4

Deliver scored data

Win rates, criterion pass rates, and the reasons behind them.

5

Re-run each release

The same harness, so you can see movement between versions.

What you get back

Not a slide deck of impressions. A dataset you can act on and re-run.

Win rates per comparison

Which output won, by how much, and whether the margin is wide enough to act on.

Criterion pass rates

How often each output cleared each check, split by critical, major and minor severity.

Expert rationales

The reason behind every judgment, in the evaluator's own words. This is where the fixes come from.

A reusable harness

The rubric, criteria and evaluator profile, documented, so the next release is measured the same way.

Start with a pilot, not a contract

Evaluation only works when the criteria are right, so we scope it with you on a call rather than guessing. Thirty minutes to agree what good looks like in your domain, then we put 20 of your outputs in front of verified experts and send back scored results with rationales.

Free, and yours to keep whether or not you continue. Full engagements are scoped per study. Most start in the low four figures and run one to three weeks, depending on domain, language and how many evaluators the comparison needs.

  • 30 minutes, no contract, no card
  • Criteria agreed with you, not assumed
  • 20 outputs scored by verified experts
  • Results and rationales sent back
  • You keep the rubric either way

Built on verification, not volume

Identity-verified expertsReal name, role, employer and work history, checked at sign-up
You own the dataScores, rationales and datasets are yours, exportable any time
Transparent compositionYou know who evaluated, by domain, seniority and market
Quarterly recalibrationProfiles re-checked so expertise stays current

Who evaluates with us

AI product teams

You are shipping AI features into a domain your team does not practise in, and you need judgment from people who do.

AI startups

You need an evaluation function before you can afford to build one, and you need it to survive customer scrutiny.

Localisation managers

Your linguists catch language errors. They are not licensed to catch the clinical, legal or technical ones sitting underneath fluent copy.

Agencies and consultancies

Verified expert evaluation and fieldwork for your client engagements, with recruitment and quality controls handled.

Frequently asked questions

Every evaluator is verified against their public professional profile at sign-up: real name, current role, employer and work history. They are then matched to your brief by domain, pass a custom screener with disqualifiers and consistency checks, face in-study quality controls, and are re-verified quarterly. A self-declared job title is never enough.
Evaluators agree far more on whether A is better than B than on what a 4 out of 5 means, so inter-rater reliability is much higher without heavy rubric training. Rating scales also suffer central tendency, where everyone clusters at 3 and 4, which kills sensitivity when two models are close. We pair forced choice with per-criterion binary checks so you also get an absolute quality bar.
CAD and AEC, design, business including finance and operations, and translation and localisation. For domains beyond these we run fresh targeted recruitment against your brief using the same verification process.
It depends on how close the outputs are and how much separation you need. Small expert panels give reliable direction on clear differences, while tighter comparisons need more raters per pair. We scope the number honestly at the start rather than quoting a default.
Yes. Evaluation runs in the language your product ships in, with native professionals who also hold the relevant domain expertise. Language fluency alone is not enough for specialist content, which is why we match on both.
You do. Scores, rationales, transcripts and the underlying dataset are yours, exportable at any time. We do not reuse your evaluation data or share it with anyone.

See it on your own outputs

A 30-minute call to scope the criteria, then 20 of your outputs scored by verified experts in your domain. No cost, and the rubric is yours to keep.