Skip to content

AI talent

AI Prompt QA and Evaluation Specialist

Builds the test sets and judging methods that say whether the system is good enough to ship, and keeps saying so after each change.

The role

What the job actually is.

Evaluation is the discipline that makes a probabilistic system releasable. This role designs the datasets, decides how outputs are judged and by whom, and turns that into something that runs on every change rather than once before launch. The judging question is the hard one: model judges are cheap and biased, human judges are expensive and inconsistent, and the specialist has to know when each is appropriate and what to do when they disagree. Their real output is a number the programme trusts enough to make a release decision on.

What we screen for

The questions that separate the field.

Asked by someone who has built the thing, and designed to catch this role's specific failure rather than to confirm a general impression.

  1. Whether they can design a dataset for a task they have not seen before. Give them one in the interview.
  2. How they handle disagreement between a human judge and a model judge.
  3. Whether they measure regression against a baseline, not only accuracy against a target.
  4. What they do about the cases that are rare, expensive and the reason the system would be pulled.

The common mis-hire

The one you have probably already made.

A manual tester applying scripted test cases to a probabilistic system. Every run produces a different result, the scripts are marked as failures, and within a month the suite is being ignored because it cries wolf.

In the estate

Where this role works, and what we screen it against.

The layers this family works at, lit. These are the tools we screen against. Naming one says we can test for it, not that we have delivered on it.

Evaluation and observability

Spans every layer. Without it a system is shipped on impressions.

  • LangSmith
  • Weights and Biases

Also screened against

  • Langfuse
  • Ragas
  • MLflow
  • DeepEval
Role families we place here

Experience and delivery

Copilots and agents inside a business process, and the interaction design that makes an uncertain system usable.

Orchestration and agents

Where an agent's steps, tools and state are defined, and where its failures are caught before a user meets them.

Models

The models themselves, and the platforms an enterprise hosts them through.

  • Anthropic (Claude)
  • OpenAI
Role families we place here

Data and grounding

What the model is grounded in, and the integration work that gets enterprise data to where it can reach it.

Systems you already run

The seven platform desks Yallo staffs. Almost no AI work is greenfield; it lands here.

Role families we place here

Governance, risk and safety

Spans every layer. Named as what governance roles are screened against; what any of them obliges is your counsel's call.

  • EU AI Act
  • ISO/IEC 42001
  • ISO/IEC 23894
  • NIST AI Risk Management Framework
  • OWASP Top 10 for LLM Applications
Role families we place here
Naming a technology here says we screen against it, not that we have delivered on it. The role families on each layer are the ones we place there.

Seniority

What changes between mid, senior and lead.

The grade is a description of what the person owns, not a band. Rates come with the shortlist.

Mid
Builds and runs test sets against an established evaluation method.
Senior
Designs the evaluation method itself, including the judging approach, and owns the release recommendation.
Lead
Sets the evaluation standard across the portfolio and defines what evidence a system must produce before it is allowed near users.

In a programme

When this role is needed, and what blocks it.

Needed before the first release and consistently under-scoped until the week of it. The dependency is subject-matter access: the test set is only as good as the people who can say what a correct answer looks like, and their time has to be booked in advance. Retained, because the evaluation has to run again on every model change.

Ask

Send the brief, get a screened AI Evaluation Specialist shortlist.

Tell us the programme, the stack and the timeline.