Cassandra Christodolo

Senior User Researcher

Why your usability testing playbook doesn't work on AI, and what to use instead

4 mins read

In February 2024, a tribunal in British Columbia ruled that Air Canada owed a customer money over a refund policy that didn't exist. The complainant asked the airline's website chatbot about bereavement fares after his grandmother died, and it told him he could book a full-fare ticket and claim a discount back within 90 days. No such policy has ever existed. The real one, on a different page the chatbot itself had linked to, says bereavement fares can't be claimed after travel. Air Canada argued the chatbot was responsible for its own words. The tribunal didn't accept that, and ordered the airline to pay roughly CA$812 in damages and fees.

It's a small claim, but a good example of a bigger problem. There's no way to know from a single test whether that chatbot would give the same answer to a slightly different phrasing of the question, or a different answer to the exact same one. A one-off test only tells you what happened with the one prompt you tried, and nothing about the other ways the question could be asked.

Large language models are probabilistic: they generate the most likely next word over and over, but which word actually gets picked involves a random draw from a set of likely options at each step. That randomness stacks up across a whole answer, so the same question can come back worded differently every time - and occasionally into something that's simply wrong.

So what does testing something probabilistic actually look like?

Testing a range, not an interface

With a form or a checkout flow, you test the thing. With an AI feature, you're testing a distribution of possible things it might say, and trying to work out where in that distribution it fails, how often and how badly. That's a different kind of job, and it needs criteria that hold up across many outputs.

Evals are the tool for this. Broadly, there are two ways to run one: get another AI to grade the output, or get a person to. Ideally, you do both.

What a human eval actually does

At its simplest, a human eval is real people, with relevant expertise, reviewing a batch of AI-generated outputs against a fixed set of criteria. That criteria could cover things like whether the output is accurate, whether it answers what was asked, whether anything is missing and whether the tone suits the reader.

Individual judgment here is a piece of expert qualitative reasoning: someone who knows the subject deciding whether an answer actually holds up, not just whether it matches a template. Some questions have one correct answer, and any decent check - human or automated - should catch a wrong one. Others depend on the specific circumstances of who's asking (like which sector someone works in), so the exact same wording can be right for one person and wrong for another.

I've built our human evals into a simple workbook. Reviewers work through a batch of AI outputs and score each one against a short, plain-language set of checks, with a written reason for every score as well as a pass or fail.

On one project testing an AI writing tool, the dominant failure was over-trimming, which nobody had expected. Everyone assumed the obvious AI clichés would be the problem. Instead, the tool was cutting so hard for concision that it lost the nuance that the original text never had an issue with. That finding fed straight back into the prompt as a specific instruction: stop cutting content that changes the meaning, not just the word count. All of the models our experts evaluated improved on the reviewed checks after that change, with Haiku's answers showing a particularly big improvement.

Where automated evals fit in

So far, people have been doing all the reviewing, and that doesn't scale on its own. Nobody can put a person in front of every output a live product generates, forever. Some of the work has to be automated. An automated eval uses one AI to score another AI's output (or deterministic assertions using regex to check structure and required elements) continuously and at a volume no team could staff.

Beyond speed, you get a trend line. Score outputs continuously, and you get a baseline for what "normal" looks like. When the pass rate starts slipping, or a new kind of failure starts showing up, that's a signal worth investigating. It could mean a model update behind the scenes, a shift in the kinds of questions coming in or a prompt that's stopped working as well as it used to. These checks work as guardrails on live models, catching your AI before it goes off the rails and starts giving incorrect bereavement fare advice.

The catch is that an automated judge only knows what to look for if someone has already worked out what "good" actually means, specifically enough to write it down as an instruction. That's exactly what a structured human review gives you. The same over-trimming finding that changed the prompt earlier also became a line in the judge's own instructions: check specifically for cuts that lose meaning, not just cuts that shorten. Neither the better prompt nor the sharper judge happens without a person doing the assessment first and then revisiting and maintaining the evals over time.

Keeping people in the loop like this does more than catch mistakes: it's what keeps the whole system pointed in the right direction as everything around it keeps changing. Building this out at Torchbox has also become a way for me to work more closely with developers. Instead of handing over a wishlist and waiting for a build, I sit inside the same feedback loop as the people writing the prompts and the code (and sometimes I'm even that person!).

We're still early in building out our evals processes and frameworks, but they've already been shaping how we test and guardrail anything AI-powered before it reaches real people. There's lots to say on this, so keep your eyes peeled for our next blog on the practicalities and benefits of evals.

Want to build more robust testing into your AI product?

Get in touch