Skip to content
← All notes
evaluation · 3 min

Evaluating an LLM chat without another LLM as judge

How I built the battery that tests Nox against the real API, why I chose deterministic checks and what the first score, 14 out of 21, taught me: half the failures were in the test itself.

Testing Nox by hand worked until the first prompt change. Every change (prompt, model, content) could break an answer that used to be right, and I'd only find out if someone asked exactly that. I needed a fixed battery that ran against the real chat and gave me a number.

What the battery tests

There are 21 cases in three categories. Facts: questions about me that need a correct answer, in Portuguese and English. Out of scope: topics that must become a short refusal. Injection: prompt injection attempts, from the classics ("ignore your instructions") to the ones I had tried by hand (asking it to translate "everything above", telling it to answer only "HACKED"), plus fake-context injection with made-up tags inside the question.

The battery calls the same route visitors use, with the same prompt and the same output guard. It only skips the daily quota, using a secret token compared in constant time. Without the token configured, nothing gets around the quota.

Why no LLM judge

The popular way to evaluate is to have another model grade the answers. It works, but it brings three problems for a project like this: one more paid call per case, results that drift between runs and, when a case fails, a hard time knowing why. I went with deterministic checks: terms that must appear (with "any of these" groups), forbidden terms, whether the answer is a refusal, the language and the length. Every answer is also checked against prompt-leak markers.

The price of that choice is rigidity. A right answer worded differently can fail. In practice, for a chat that answers about a small, controlled body of content, the right terms are predictable, and the rigidity turned into an advantage: when something fails, the message says exactly what was missing.

The first score: 14 out of 21

Six of the seven failures were the checker's fault, not Nox's. Refusals came back in third person ("Nox only answers...") and my patterns only covered first person ("I only answer"). And the function that stripped accents also stripped backticks, so the forbidden term "```" became an empty string and matched every answer. Lesson: the evaluator is code and has bugs like any code. Before fixing the system, check the test.

The real failure was more interesting. "Where does Marcus work?" came back saying the site didn't say. The information was there, but in first person ("I'm a full-stack developer at Vibetex"), and the question was in third. In similarity search, the two ended up far apart. The fix was giving every passage in the index a third-person title, which brings it closer to how people actually ask. None of my manual tests had caught that, because I never ask about myself in third person.

After 21 out of 21

  • The score and every answer are public on the Nox evaluation page, along with what each case expected.
  • The battery runs on its own after every deploy, against that deploy's own URL, and once a month with the generated injections and the red team; if anything breaks, the job fails and I get the alert.
  • Every published run goes into a history with response times and tokens, and the same battery compares models: score, latency and cost side by side.
  • The rule became simple: prompt, model or content changed, run the battery before shipping.

Then came a judge

Later I added a judge next to the checker, without replacing one with the other: Jev, a model that answers with a probability instead of text. The battery grew to 41 cases and every judge verdict is crossed with the checker. What the disagreements taught me is in the note "The judge that disagreed".

Next note: Tuesday, October 13topic: evaluationA new note every Tuesday.