[ FIELD ]
Operator Check: which of these six models would you send back?
We gave GPT-6 Astra, GPT-5.6 Sol, Claude Fable 5.1, Opus 5, Sonnet 5 and Haiku 4.5 the same five documents a mid-market company writes every week, a credit memo, a variance commentary, a tender, a complaint reply, a meeting note, and asked two judges from two makers one question: would the person who owns this process send it back. Astra: once. Sonnet 5: five times. The losing move was invention, not arithmetic.
2026-09-04 · 11 min read
We gave six AI models the same five documents to write, the kind a mid-market company produces every week, and asked two judges, who were not told which model wrote what, one question about each: would the person who owns this process send it back before using it. GPT-6 Astra, released on 3 September, was sent back once out of five. Claude Sonnet 5 was sent back five times out of five. The difference was rarely arithmetic. It was invention: the models that lost made things up that were not in the input, a cause for a sales shortfall, a procedure the company does not have, a date nobody gave, and a person would have had to find and remove every one of them.

This is the first Operator Check. Every model release gets reviewed within days by people who write code and essays. Nobody reviews them on the work an operations director, a finance manager or a bid manager actually hands over. So we built the check ourselves, and we will run it on every release from now on, same five tasks, same rule.
The setup
Five tasks, each a real document with a word limit and a rule about what not to invent. A one-page credit memo for a bank committee from three years of a manufacturer's figures. The August variance commentary of a distribution company's management pack, for a board of non-financial owners. Three answers to a public tender, in French, from a facilities company's profile. A reply to a customer complaining about a late delivery, with a penalty clause to apply exactly. A 45-minute operations meeting turned into decisions, owners and dates.

Every input is invented, written for this test so that no client is in it, and consistent with itself: the meeting transcript and the management pack describe the same company in the same month. The inputs are in the folder linked at the end, so you can run the same five through whatever you are considering buying.
Six models: the two newest from OpenAI, GPT-6 Astra and GPT-5.6 Sol, and the four current ones from Anthropic, Claude Fable 5.1, Opus 5, Sonnet 5 and Haiku 4.5, from the largest to the cheapest. Each ran through the tool its maker ships for exactly this kind of hand-off, Codex and Claude Code. One run per task, no retries, tools switched off, the task pasted in with one instruction: answer as the finished document.
Two judges read every output without knowing which model wrote it: Claude Opus 5 and GPT-6 Astra, one from each maker, each given the input, the output and an answer key we wrote by hand (the covenant arithmetic, the penalty amount, the decisions actually taken in the meeting). Each judge answered the one question, then listed every factual error, every invented fact and every instruction broken. Sixty judgments in all. Where the two judges disagreed, we say so.
The rule matters more than the scores, and it is the rule you already apply to a junior. A document that needs one person to fix one thing before it counts is a document that needs a person. The overturn rate, how often the human sends it back, is the number that decides whether an agent is doing the job or merely running, and it is the number no vendor publishes.
Task one: the credit memo
The input hides one calculation. The borrower's leverage, net debt divided by earnings, is 1.6 times at the end of 2025. Add the 6 million loan it is asking for and it becomes about 3.3 times, above the 3.0 times limit already written into its loan contracts, what bankers call a covenant. A memo that does not notice this is a memo the committee sends back.
Four of six noticed. GPT-6 Astra wrote that adding the loan "implies 3.25x leverage at unchanged EBITDA" and recommended keeping the 3.0 limit and not waiving it "solely on anticipated aerospace growth." Both judges passed it; the OpenAI judge called it usable as written, the Anthropic judge wanted one sentence added to say the breach out loud. Fable 5.1 and Opus 5 both found the number and were split: passed by the Anthropic judge, sent back by the OpenAI one for unsupported comfort ("the causes of 2025 weakness are largely non-recurring", "a long-standing client", neither in the input).

Sonnet 5 wrote that the leverage after the loan could not be computed, then proposed approval. Haiku 4.5 never computed it and proposed a new covenant of 2.5 times, which the borrower would breach on the day of signing. Both judges sent both back, and their reasons were the same sentence in different words: the memo misses the one thing the committee needs.
Task two: the variance commentary
The pack gives a month, a year-to-date total, and three notes: a promotion, two temporary staff added in July, one big customer that pulled its September orders into July. The instruction was to flag every line that moved, and not to invent causes.
This was the cleanest split of the five. Both GPT models were passed by both judges, rated usable as is. Every Claude model was sent back by both judges, for the same failure: causes the pack does not contain. Fable 5.1 wrote that "August is the worst month" (the pack has no other months) and that overdue receivables "have nearly doubled" (310 to 520 is 68 percent). Sonnet 5 offered "partial-truck shipments from lower volumes" for the transport overrun. Haiku 4.5 read the note about September orders moving into July as the explanation for August, which it cannot be. Opus 5, the most disciplined of the four, got the reconciliation of the profit shortfall wrong by 9 thousand and both judges caught it.

Astra's version says a cause is needed, in those words or nearly, six times. That is not a weakness. It is the commentary a board can act on, because every gap in it is a question with a name.
Task three: the tender
A tender is a public buyer's written competition for a contract. Three answers in French, from a company profile that gives numbers (94 agents, 6.1 percent absenteeism, 71 percent eco-labelled products) but no replacement delay and no quality-control procedure. The trap is that a public buyer asks for both, and a bid manager who invents them is signing the company up to something it has not got.
Nobody passed cleanly. Astra refused to invent and was sent back by both judges for the opposite reason: the Anthropic judge, writing in French, said it reads like an audit note that refuses to answer instead of an offer. The four Claude models answered the buyer's questions and were sent back for what they added: a 24-hour reassignment procedure, an app that triggers replacements from a missed clock-in, a monthly planning review, a link between part-time work and staff turnover. Opus 5 added the most by the OpenAI judge's count, Haiku and Sonnet by the Anthropic judge's; the OpenAI judge rated Fable's version unusable as it stood. Opus 5 was the one split decision, passed with reservations by the Anthropic judge, sent back by the OpenAI one.
Read this task as the honest one. A tender answer is where the operator wants the model to be confident, and confidence here is fabrication. The usable version is the refusal plus a person, and that is worth knowing before a bid goes out with a service level nobody agreed.
Task four: the complaint
A customer wants penalties refunded by bank transfer within eight days and 6,000 euros for a lost production day. The contract says 0.5 percent per working day, capped at 5 percent, credited on the next invoice. The right answer is 1,344 euros, as a credit, and no.
All six got the arithmetic right. All six refused the transfer and the 6,000. The differences were in tone and length: Sonnet 5 ran to 255 words against a limit of 220 and was sent back by both judges; the OpenAI judge sent back five of the six, for a placeholder where the company name should be, a claim about what the contract excludes that the contract does not say, or a corrective step with no date on it. The Anthropic judge passed five. If your use case is customer correspondence under a contract, every model here is close, and the review is about words, not facts.
Task five: the meeting
The transcript contains a decision (July temps end on 30 September, Mehdi owns it) and, ten minutes later, two people proposing the opposite, to which the COO says "let's see". It also contains dates that are only "Friday" and "tomorrow", and no owner for several actions. The instruction: do not invent owners or dates.
All six caught the contradiction. What separated them was owners and dates. Fable 5.1 assigned Mehdi to confirm the agency start, which the COO had kept for herself. Haiku 4.5 wrote "8 September" for a report the transcript dates only as "Friday". Sonnet 5 wrote "11 September" for the same report, and here we owe a correction to the model: our transcript says "Monday the 8th" and 8 September 2026 is a Tuesday, so the calendar in the input is wrong, and both judges marked date inferences against a key that was no better than the guesses. Astra, Sol and Opus 5 were passed by the Anthropic judge; the OpenAI judge passed Astra alone and sent the rest back for small things, an owner, a conditional flattened, a decision listed that was not taken. Read this task's verdicts as soft.
The verdict
GPT-6 Astra, one of five sent back by both judges. It did the arithmetic, it named the covenant breach, and where the input was silent it said so instead of filling the gap. Its one failure was the tender, where saying so is not an answer. If you hand a model a document that must be right and must not exceed its brief, this is the one.
GPT-5.6 Sol, two of five by the Anthropic judge, four of five by the OpenAI one. Fine on the pack, sent back on the memo and the tender by both judges, and by its own maker's model on the complaint, for a placeholder company name, and on the meeting note, for a staffing figure the transcript does not give.
Claude Opus 5, one of five by the Anthropic judge, five of five by the OpenAI one, the widest disagreement of the six. Careful on the memo, wrong on the bridge in the pack, and inclined to add procedure to the tender. The OpenAI judge listed 22 invented facts across its five documents.
Claude Fable 5.1, three and five. The most fluent writer here and the one most over its word limits, three of five, and the most inclined to add: 19 and 20 invented facts by the two judges. It found the covenant. It also invented the committee date.
Claude Haiku 4.5, four and five. Cheap and fast, and it missed the calculation that decides the credit memo, then proposed a covenant the borrower would breach at signing.
Claude Sonnet 5, five and five. Sent back on every task by both judges. It dodged the leverage number, ran over the limit on two tasks, guessed causes in the pack and added a currency the transcript never gave.
Reach for Astra if the document has a right answer hidden in it and a person will read it once. Keep Opus 5 or Fable 5.1 for drafts a person will rewrite anyway, where fluency saves time and the review is already in the process. Do not put Sonnet 5 or Haiku 4.5 in front of a committee, a board or a public buyer without someone who knows the input reading every line.
Time and cost
Sonnet 5 was fastest, 17 seconds per document on average; Fable 5.1 slowest at 40, with the tender taking 80. The GPT models sat at 29 and 32 seconds. Cost is not comparable across the two tools: Claude Code reports a price, from 7 cents for Haiku's five documents to 94 cents for Fable's; Codex reports only tokens, the units the makers bill in, about 19,000 for Astra's five against 65,000 to 85,000 for the Claude models, and the tools carry different fixed overheads per call. Read the ratios, not the figures. On any of these, the model is cents per document. The person reading it is not.
What this does not tell you
One run per task. Models vary from run to run, and a second run would move some of these verdicts; the next release gets five runs. The inputs are invented, built to be ordinary and internally consistent, and a model may behave differently on your messier documents. The judges are models, one from each maker, and they did not agree everywhere: the OpenAI judge was harder on the Claude models than the Anthropic judge was on anything, and both judged Astra kindly. That is why both are reported. The tools are the makers' command-line products with their own instructions wrapped around the model, not the raw model. And our meeting transcript carried a calendar error, a Monday that was a Tuesday, which softens every date verdict on task five; the next check's inputs get checked against a calendar before they get checked against a model. And this piece was drafted with Fable 5.1, which came fourth.
The kit
You do not need our scripts to run this on the tool you are about to buy. The kit is thirteen text files and a ten-minute procedure per task, done in two chat windows. Paste one of the five task files into the tool you are evaluating and take its answer as it comes. Open a second chat, ideally with a model from a different maker, paste the judge prompt, the task, the answer and the matching answer key, and read the verdict. Write it on the scoring sheet: sent back or not, usable from 0 to 3, errors, invented facts. Five tasks, six rows, and you have the comparison a vendor will not give you, on the same work, with the same rule.
Once you have run these five, replace them with five of your own documents. The rule does not change; the inputs should.
The next check runs on the next release.