Measured, not asserted
AI Quality Report
Every figure on this page comes from a specific, dated, automated test run. The run links are live — click them.
Most AI products ask you to trust a claim. We publish the test results instead, including the ones that do not flatter us.
What we measure
Every night we replay a 160-case decision-logic suite against the live model, across all twelve languages we test. It scores whether the AI makes the right decision — does it notice an answer is vague, catch a respondent contradicting themselves, choose the right input control, follow a signal the respondent never made explicit. It does not score writing style.
Latest result
140–142/160 (88–89%)
Cases passed, last four nightly runs
That is the spread on an identical set of tests. The most recent run, 29 September 2026, passed 140 of 160.
We publish a range rather than a single percentage because a single percentage would be wrong by the next morning. Nothing about the tests changed between these runs — the movement is the model giving a different answer to the same question on a different night. Any AI quality figure quoted to one decimal place is hiding this.
By category — including our worst number
| What it tests | Last four nights |
|---|---|
| Choosing the right input control | 32/32100% |
| Catching a respondent contradicting themselves | 46/4896% |
| Detecting a vague answer worth following up | 45/4894% |
| Branching implicitly on an unprompted signal — our weakest category | 17–19/3253–59% |
That last row is the honest one: branching implicitly on an unprompted signal is our weakest area, and on every night above it accounts for at least 72% of everything the suite got wrong. It scores a behavior our shipped engine does not yet fully implement. We report it at the same size as the good numbers, because a quality page that hides its worst row is marketing, not measurement.
One thing to know before you click a run link: those four nights each scored 192 cases, not 160, and their summaries say so. The extra 32 tested a “recall” question type — offering someone back an answer they gave in an earlier conversation — which we have removed from the product. It was built without the identity model it needed, so we took it out rather than ship something we would have to warn you about.
Those 32 cases passed on every night above, so what you see here is those same four runs with a category that is no longer in the suite subtracted — arithmetic on the linked runs, not a re-measurement, and it moves every figure on this page against us rather than for us. We would rather print a smaller number than a present-tense claim about tests we no longer run.
The reason the two disagree is that every night above replays one fixed build of our engine, cut on 17 September 2026 — that is what makes four nights comparable to each other at all. We removed the question type on 23 September 2026, after that build. So those runs still score the older suite, and will until the pinned build moves past the removal. When it does, the run links and the totals here agree again and this paragraph goes.
Languages
The suite runs in all twelve languages we test: English, Spanish, French, German, Brazilian Portuguese, Italian, Dutch, Polish, Japanese, Simplified Chinese, Korean, and Turkish.
English results gate our builds — a regression fails the build and the change does not ship. The other eleven languages run every night and report results, but do not yet block a release. They are advisory, and we label them that way deliberately: the cases were written directly in each language rather than machine-translated, but they have not been reviewed by native speakers.
That review has not started. We have not yet settled how to source native reviewers for eleven languages, and until a language is reviewed it stays advisory. We would rather say that than let “runs nightly” be read as “verified”.
We deliberately do not publish a per-language league table. 10 cases per language means one case is 10.0 percentage points. Across four consecutive nights on an identical set of tests, four of the eleven advisory languages moved by a full case purely from model variability. Ranking languages off that would be presenting noise as a quality difference. When a language’s test set is large enough for its number to mean something, it will appear here.
Native-speaker review will not change that. Review fixes whether a case is a fair test of the language; it does not make the sample bigger, and a perfectly reviewed 10-case set still swings 10.0 points per case.
What this does not tell you
- It measures one decision point at a time, not a whole conversation end to end.
- It is not a guarantee. We do not offer a contractual accuracy commitment, and we would rather say so than imply one.
- One night is one measurement. That is why this page reports a range over four nights, and why we republish when the numbers move rather than leaving a stale figure up.
Cadence
The suite runs nightly. This page is refreshed from those runs — on real change, not on a fixed publishing calendar. The figures above were last transcribed from the 29 September 2026 run.
Want the rest of how we handle your data?
The home page has the respondent-facing side of this — what people answering your conversations can see and change.
How your data is handled