Every claim gets a grade
I stopped trying to make my AI helper more accurate and started asking it to show its work. Four small habits came out of it, and they work just as well in a paper notebook.
The answer that costs me the most from an AI assistant isn’t the wrong one. It’s the wrong one that arrives in exactly the tone of a right one.
Wrong answers that announce themselves are easy. You catch them, you move on. The ones that sting come with the same steady confidence as the checked ones, so there’s nothing on the surface to sort them by, and you act on them. Three weeks later you find out which kind it was.
Librarians have a name for this, and we teach it to eighteen-year-olds: source evaluation. What surprised me is how little of what I knew about evaluating a source I was applying to the tool sitting on my own desk.
My assistant is called samson. At some point I stopped trying to make samson more accurate, because accuracy was never really the lever I could reach. I asked it to be legible instead. Every claim it hands me now comes with a little mark saying how well it’s supported, and the rules that produce that mark are written down where both of us can see them.
Four habits came out of it. None of them need a computer.
One: every claim carries a grade
Three marks:
[P]primary, confirmed. I opened the thing.[C]circumstantial. It fits the evidence, but the evidence doesn’t quite prove it.[X]unverified. I’m telling you this because it’s probably useful, not because I checked.
Every factual sentence gets one. Not just the important ones. All of them.
The marks themselves are almost beside the point. What does the work is the half second before you type one, because you can’t put [P] next to a sentence you haven’t actually checked without noticing that you’re doing it. The grade is a speed bump disguised as a notation. Most of the checking I do now is checking I only started because I had to pick a letter.
The place it slips is the sentence after the fact. Numbers get graded without thinking about it. So do dates, dollar figures, page counts. Then one line later comes the sentence that connects the number to the argument, the one that starts “which means,” and it goes out unmarked because it reads like conversation rather than like a claim.
It is a claim. It’s usually the claim that matters, and it’s the one people quote back at you. So it gets a grade too, and in practice it’s the one that comes back [C] most often.
I use the same three marks in my own notes now. It turns out I was doing the confident-wrong-answer thing to myself long before I had anything to blame it on.
Two: an absence is a claim
“There’s no record of that” doesn’t feel like a claim. It feels like the absence of one, which is why it slides through unchecked.
It is a claim, and a big one. It says something about the entire universe of a source. And it’s one of the easiest things in research to get wrong, because a search that finds nothing looks exactly like a search that ran correctly and found nothing.
So the habit is: before an absence gets written down or acted on, run the search that would have disproved it, and keep a note of what you ran. Then try it a second way, against a different source. One tool reporting nothing is a fact about that tool.
The shape that gets me most often isn’t a missing record. It’s the wrong key:
- Searching a person’s name in a system that indexes on an ID number.
- Searching an organization under the name it uses today, when it filed under the one it used in 2019.
- Searching a field that exists, but is spelled a little differently in this particular export.
Every one of those runs clean, returns fast, and returns nothing. And nothing looks the same no matter why it happened.
Three: run the positive control
That little note is the whole third habit, borrowed straight from the lab bench. Before you trust a negative result, run the same test on a case you know is positive. If the known-positive comes back negative, your instrument is off, and your negative result isn’t a result at all.
Here’s the one that taught me to write it down.
I was searching a stack of financial filings for a particular term. The search came back empty across the whole set. That was interesting, because the term should have been in at least a few of them, and an unexplained absence in a pile of filings is the kind of thing that makes you sit up.
It wasn’t a finding. The search tool was quietly failing on the compressed files and reading none of their contents. It reported zero matches because it had read zero bytes, and it finished without complaint, because from the tool’s point of view nothing had gone wrong.
The part I keep thinking about is that the control was right there. A document I knew contained the term had been in the same run, and it had come back empty too, in the same output, at the same moment. Nobody looked at it. The null was interesting and the control was boring, and the interesting thing got all the attention.
So now: a null whose positive control also comes back empty is a broken lookup, not a finding. Stop there. Don’t write it up. Don’t even mention it in passing as a thing worth looking into, because “worth looking into” is how a broken search becomes a fact six weeks later.
This habit costs almost nothing. It’s one extra search. It’s the highest-yield item on this list, and it’s the one I most often catch myself skipping, always for the same reason: the result was too good to slow down for.
Four: write down what fooled you
Most note systems save what you learned. Mine mostly saves what fooled me.
One trap per file. A short name, a one-line description of the shape of the mistake, and enough detail to recognize it next time it shows up in a different costume. Not “the Wayback Machine has an availability API.” Instead: this API will tell you a capture doesn’t exist when it does, so use the other endpoint, and treat an absence reported by this tool as a fact about the tool.
That’s a different kind of note than a fact, and it earns its keep differently. Facts get looked up. Traps get walked into, and the only defense is having read the description recently enough to feel the shape of it coming.
Two things make it work.
Everything is in an index. A note nobody loads isn’t a note, it’s a file. The index is one line per entry with the hook in it, and it’s the thing that gets read first, every time.
Every rule traces to a specific day. Nothing goes in because it sounds like good practice. If I can’t point at the day it cost me something, it doesn’t go in the file, because a rule with no story behind it is just clutter competing for attention with the rules that earned their place.
And the corollary I keep having to relearn: a rule that needs a second reminder needs a mechanism instead. When I find myself writing the same warning down twice, the right move isn’t a firmer warning. It’s a checklist, a required format, a little script, something that makes the mistake harder to make. Writing it down a third time is how a document gets long enough that nobody reads any of it.
What it costs
It’s slower. Not slightly.
Answers now arrive with their sources attached, and a real share of them say some version of “I haven’t checked that.” A question that used to get a clean paragraph back sometimes gets a table with three rows marked [X] and a note about which lookup would settle it. That’s uncomfortable in a way I didn’t expect. Confident prose feels like progress, and a table of unverified rows feels like being stuck.
But the comparison that matters isn’t “fast answers versus slow answers.” It’s “slow answers versus the same fast answers plus the afternoon in October you spend working out which of them was quietly wrong.” I’ve had that afternoon. It is not faster.
The test I use for whether any of this is working is small and specific: when I ask “is that right?”, does the answer change? For a long time it did, often enough that I stopped trusting the first version of anything. It changes less now. That’s the whole measurement.
None of this is really about AI, in the end. What samson did was make the problem visible, because it produces enough claims per hour that a habit which used to take a year to show up now shows up by Thursday. Every one of these would have made my work better in 2015, on a reference desk, with a legal pad.
Four habits you can start tomorrow
- Mark every factual claim as confirmed, circumstantial, or unverified. Deciding the mark is what forces the check.
- Treat “there’s no record” as a claim: run the search that would disprove it, keep a note of what you ran, and confirm from a second source.
- Before trusting a null, run the same search on something you know is there. If that comes back empty too, it’s your tool, not your topic.
- Keep a file of what has fooled you, indexed, one trap per entry, and only add a rule when you can name the day it cost you something.
