The Disclaimer Says Check The Output. No Product Gives You Tools.
A widely-shared post argues every major AI tool admits it makes mistakes, then ships zero interface for catching them. The fixes it proposes are mostly UI.
Three disclaimers, one missing feature
Gemini says "AI can make mistakes, so double-check responses". Claude says "Claude is AI and can make mistakes. Please double-check responses." ChatGPT says "ChatGPT can make mistakes. Check important info." All three in small grey type, at the bottom of the box.
In a post published on 27 September 2026, the developer writing as Glyph takes those sentences at face value and asks the obvious follow-up: if checking the output is a mandatory part of the workflow, where is the interface for doing it? His answer is that there isn't one, and that the disclaimers function as legal liability transfer rather than product design. His conclusion is blunt — a tool that tells you to verify its work while giving you no way to track verification is, in his framing, closer to a grift than a product. He extends the same criticism to Ollama, the local-model runner, which he argues needs these features more than the frontier labs do because the models it runs are weaker.
What the proposed interface actually looks like
The specifics matter more than the polemic, because most of them are user-interface changes rather than model changes.
For chat, he proposes a two-column worksheet: model output on the left, a human notes field on the right recording what work went into checking each claim, and a checkbox you tick only once you believe you've checked it. For coding assistants, he argues the absence of this pushes verification into code review, which lets the nominal author ship a pull request they never read.
For research, the reversal is sharper. Today, citations appear as tiny inline annotations showing a domain name and a sub-16-pixel icon. He wants the citation to be the primary object: publication date, author name where available, and the literal unmodified quotation — extracted by an ordinary program, not by the model — displayed larger than anything the AI wrote. The generated summary gets demoted to the small grey text currently occupied by the disclaimer.
Worth unpacking the jargon here. RAG, or retrieval-augmented generation, is the standard trick where a system searches documents first and feeds the results into the model as context. Grounding is the vendor term for tying output to those retrieved sources. Glyph concedes those little citation links do point at real structures in the pipeline rather than hallucinated tokens. His point is that a model can still garble retrieved text the same way it garbles anything else, so the presentation must never look authoritative.
Related: provenance. When a chatbot shows you a data table, some of it came mechanically from an API or tool call and some is model output. Both are rendered identically. He wants those visually separated, plus a pass where numbers are treated as a spreadsheet and verified with ordinary arithmetic, with the working shown.
Genuinely new versus already known
Little of this requires a model breakthrough, which is the argument's strength. Citation cards, checkboxes, provenance styling and arithmetic verification are harness work — the software wrapped around the model — and could ship against today's models unchanged.
Two proposals do require lab cooperation. Banning first-person language and apologies from the output is a model-behaviour change; he argues the labs' own benchmark claims prove they can control output tightly. And exposing temperature — the randomness dial that determines how varied an answer is between runs — he raises but immediately qualifies, noting that setting it to zero doesn't simply produce reliable results. The reproducibility section is the least developed part of the post.
The MCP critique cuts against the industry's current direction. Model Context Protocol lets chatbots call external tools and take actions. Glyph's objection: if natural-language input is imprecise enough to need constant clarification, granting that same system destructive powers is backwards. He'd rather see task-specific buttons — a dedicated control for scanning a codebase for OWASP Top 10 vulnerabilities, say — backed by smaller purpose-trained models.
Questions You Should Be Asking
- If our staff are required to verify AI output, where in our tooling is that verification recorded — and can we produce the record if a claim we shipped turns out to be wrong?
- Ask your vendor directly: what percentage of citations shown in your interface link to text a user can read unmodified, versus a model-written paraphrase?
- How many of our merged pull requests were authored by an assistant and reviewed by someone who assumed the author had already checked them?
- When a table appears in a chat response, can anyone on the team tell which cells came from an API call and which came from the model?
- Are we buying agentic tool access because it solves a problem we have, or because it is what vendors are currently selling?
What To Watch Next
The signal is where the disclaimer sits. If any major vendor moves verification from grey legalese into a functional interface element — a checkbox, a citation card, a provenance marker — it means they accept design responsibility for the error rate rather than delegating it to the user. Everything else in this debate follows from that one choice.
- 1Keep a separate verification log — paste each AI claim with its source link and a verified/unverified flag, since no chatbot tracks this for you.
- 2Ask the model to list every factual claim it made as a numbered checklist, then check them off one by one against primary sources.
- 3Treat any AI output you haven't personally sourced as a draft hypothesis, and never forward it to colleagues or clients unmarked.
Ready to implement AI in your business?
Our team builds the AI systems you just read about. Start with a free 30-minute discovery meeting.
