Skip to main content

Judging Evidence in AI Experiments

Use this article to assess the strength of a claim after reading its methods and results. Model cards and benchmarks cover identifying model artifacts; data and evaluation covers designing the experiment. Here the focus is the inference from that evidence to a conclusion.

A note can contain ten links and still be mostly opinion. To judge it, check what each source actually supports.

A specification can establish how an API is defined. It cannot establish that every implementation works. A benchmark can report performance in one setup. It cannot choose the best tool for every repository. A personal test can guide my own workflow, but one run does not become a general result.

Match the evidence to the claim​

If the note claims...It should usually have...That still does not prove...
A protocol or API behaves a certain wayA versioned specification or current official documentationThat implementations are complete, secure, or interoperable
A model or product can do a taskA model card or documentation plus a task-specific testGeneral superiority outside the tested setting
A benchmark resultThe task and harness versions, model, date, configuration, and result recordTransfer to another repository, language, or machine
A personal observationA reproducible run record and retained failuresPopulation-level performance
A recommendationCriteria, alternatives, costs, and a failure thresholdOne best choice for everyone
A forecastAssumptions, scenarios, and uncertaintyA verified future fact

Official documentation is the right source for an interface and for what a vendor says its product does. It is interested evidence about quality. Independent benchmarks reduce that bias, but introduce their own choices of tasks, configurations, and publication thresholds.

Bias is easier to manage when it has a name​

Before trusting a comparison, I check what it leaves out. The omissions are often more important than the last decimal place.

  • Selection bias appears when the test excludes inconvenient models, languages, tools, or failed runs.
  • Vendor bias appears when product positioning is repeated as a measured result.
  • Benchmark bias includes contamination, narrow tasks, metric choice, and confusion between model and harness quality.
  • Temporal bias turns a dated snapshot into a durable ranking.
  • Hardware bias hides assumptions about memory, power, accelerators, or network access.
  • Language and domain bias generalizes English or Python results to settings that were never tested.
  • Automation bias treats fluent output or an agent's completion message as proof that the work is correct.
  • Workflow bias turns one person's preference for terminals, IDEs, or automation into advice for everyone.

This list does not make a note neutral. It makes the remaining perspective visible enough to challenge.

Write recommendations as decisions, not slogans​

A compact claim record makes the missing pieces hard to hide:

claim: gpt-oss-20b can be a practical local agent model on a 16 GB machine
kind: recommendation
as_of: 2026-08-10
facts:
- provider says the MXFP4 model runs within 16 GB of memory
assumptions:
- runtime supports Harmony and the exact quantization
- context and KV cache stay within the remaining budget
missing_evidence:
- accepted-result rate on my coding tasks
- latency and peak memory on my hardware
alternatives:
- smaller dense model
- larger model on 24 GB
- hosted model for escalation
falsifier: repeated tool-call or patch failures above the task threshold

This format separates a provider statement from local assumptions and names the test that would change the recommendation. It is more useful than writing that a model is "powerful" or "promising."

A hash proves less than it first appears​

Saving a content-addressed copy of a source makes later review reproducible. An exact excerpt check can prove that quoted characters occur in the saved file.

It cannot prove that the file came from the claimed publisher, that it is current, that the surrounding context preserves the meaning, or that the quote supports the conclusion. Those questions still need source checks and semantic review.

I keep four labels separate:

  • a fact has an identified source and a passage that supports it;
  • a discovery lead is worth investigating but remains unverified;
  • a hypothesis is a testable explanation or prediction;
  • a preference depends on goals and trade-offs.

Not every note needs an archived copy of every page. Strong custody matters most when a claim changes quickly, is disputed, costs a lot to reproduce, or could cause real harm.

A review that stays small enough to use​

Extract the claims before polishing the prose. Label each one, then ask whether its evidence is strong enough for the decision it supports. Look for a credible counterexample and keep failed runs beside successful ones. Separate durable ideas from dated product snapshots.

Finally, write down what would change the conclusion and set the next review date from that. The facts, tests, and judgments should still be easy to tell apart when the review is done.

Explore connectionsOpen network