The Lab

The instrument the reports are readings from.

Ask a chatbot what people are saying about something and it will tell you, fluently, with no idea where the answer came from. This was built because that is not good enough to publish.


It is called Ouroboros, after the snake that eats its own tail, because what it settles gets written back into the thing it reads from. It has one job: to make a claim about what is happening in a field, and to make that claim traceable back to the rows it was counted from — so that a reader who doubts it can go and check rather than deciding whether to trust the tone.

The distinction it exists for is the difference between a doctor and someone playing a doctor on television. Both are fluent. Only one of them can show you why. This can still be wrong — and has been, on this page, below — but it is wrong in a way you can find.

Public activity on open-source projects
Research papers
Company filings
Public posts
Recorded talks
SENSORSOne storeNOTHING DELETEDRead in stagesCHEAPEST FIRSTThe ledgerCLAIM FIRSTA field reportPUBLISHED HEREwhat was settled goes back in
Collection is the easy half. The two stages on the right are what separate a reading from an opinion: nothing reaches a report without a written claim behind it, and what the ledger settles is written back into the store, so the next question starts from a settled answer rather than from scratch.

What it watches

Five sensors, each collecting a different kind of public evidence, and each running on its own schedule into ClickHouse, a database built to count across hundreds of millions of rows in seconds rather than to serve an application. Nothing is ever deleted from it. That is a deliberate and slightly costly decision: you cannot go back and decide retroactively that something was worth keeping, so everything is kept.

SensorHoldsCovers
What people build in public140,957,350 activity recordsacross 17,087,365 projectsJanuary 2023March 2026
What researchers publish290,240 papersJanuary 2023August 2026
What public companies tell their regulator19,572 filingsacross 530 companiesJanuary 2023August 2026
What people say about those projects22,946 postsacross 2,359 accountsSeptember 2025August 2026
What gets said out loud in recorded talks3,148 segmentsacross 10 channelsJanuary 2024August 2026
Every window is published with the month it stops. One sensor stopped collecting in March and has not been extended — and its window also holds a pause of more than a year in the middle. Both are facts about the instrument, and a reason to distrust any claim it makes about months it did not collect.

Who does the reading

Most questions are answered by counting, which is what the database is for and costs nothing worth mentioning. The ones that are left need something to actually read a passage and judge what it says — and which model does that is a measured question here, not a preference.

Everything goes through one gateway — LiteLLM, which gives every provider the same shape so nothing downstream has to know which one it is talking to, pointed in turn at OpenRouter, which resells hundreds of models from one account. Between them, trying a different model is a line of configuration rather than a rewrite, and that is what makes the measurement practical: for each job, every candidate reads the same sample and is scored against a reference model before the winner does the bulk run. Some jobs draw a field of candidates; one had a single candidate, measured against the reference all the same rather than trusted on reputation. Different jobs end up with different winners, and the winner is rarely the expensive one.

Categorizing passages of company filings

Candidates measured
5
Chosen
openai/gpt-oss-120b
Read by the winner
18,260 passages
Read by the reference
545

Sorting what people said in public posts

Candidates measured
1
Chosen
deepseek/deepseek-v4-flash
Read by the winner
9,570 posts
Read by the reference
672
The reference model is the expensive one, and it never does the bulk run — it reads a sample twice to establish how much it agrees with itself, which is the ceiling every candidate is scored against, and then audits the winner. Paying up the price ladder has more than once bought nothing: on the filings, two models costing tens of times more were further from the reference than the one that won.

Alongside the sensors there is a shelf of 65 documents kept in full rather than counted — papers, posts, and articles worth reading properly rather than tallying. A number that comes out of a database query is only ever as good as the question; the shelf is what gets consulted when the question itself is the thing in doubt.

The claim gets written down first

Everything so far is collection and arithmetic, and neither of those makes anything trustworthy. What does is the ledger: the claim first, then the method for testing it, and only then is the test run. What the numbers say and what they are taken to mean are recorded as separate things, because the gap between those two is where most confident nonsense lives.

There are 22 claims on the record on the one subject it has been pointed at so far, of which 1 came back refuted and 14 came back mixed — the evidence pointing both ways at once. Claims that failed stay on the record. A ledger that quietly loses its wrong answers is a marketing document, and the whole point of writing the claim before the test is that it becomes impossible to pretend afterward that you expected the result you got.

What the first pass found

Companies have been putting real numbers behind their AI claims less and less over the last three years.

Clean, quotable, and consistent with what everyone already suspects. It is also an artifact of the paperwork: the mix of document sections shifted toward the ones that never quantify anything. Inside every individual section, the rate was flat.

What was actually happening

Nothing fell. The talk moved — out of the announcements companies make voluntarily and into the risks they are obliged to disclose.

A better finding than the one it replaced, and it only appeared because the first one had to be written down as a claim, tested inside each section separately, and allowed to fail.

Both statements are true about the same filings. Only the second one is about companies rather than about paperwork.

What gets published, and what does not

Most of what the ledger holds is not on this site, and will not be. A finding that cannot be stated without three qualifications is a finding that has not been understood yet. The reports carry the ones that survived, along with the limits of the data they came from — including, on more than one of them, the fields that were measured and then cut.

That is the honest summary of this whole section. There is a rented model in the middle of it, and around that model a set of notes, rules, and habits that between them can be pointed at a real question and produce an answer that does not have to be taken on faith. The reports are what that looks like when it works.