Retro-futurist terrace overlooking a spired city at sunset

How we keep score

Every number on this site can be traced to a published source, a line on the page, and a rule anyone can read. This page explains the rules in plain language.

Optimistic missionNeutral measurementOpen evidence

The three promises

The foundation of everything we do.

A spired city on an island at sunrise

Optimistic mission

We built this because progress, setbacks, and limits in AI alignment deserve to be measured and shown to everyone, not argued about in the dark. The hope is in why we exist and how the site looks.

A balanced scale before a spired city at sunrise

Neutral measurement

The rules that sort a result do not know which direction it points. A failure and a success get the same form, the same checks, and the same space on the page. The hope never touches the data.

A spired city inside a glass case

Open evidence

Every card links to the source it came from and the exact line we read. The whole dataset is public and free to reuse. If we get something wrong, the correction stays visible.

Three things we never mix up

A clear record means three parts are always separate. All three are shown for every result.

Gauge with its needle pointing straight up

What happened

Did the number move toward more control, away from it, or not at all. We always record this, even for tiny changes.

Gauge face showing four rising bars

How big

Before and after, in the test’s own units, shown as a number out of 100 wherever the measure allows it.

Gauge with a needle across a graded arc

How sure

What the evidence can actually establish. Sometimes a report gives enough to say chance is unlikely. Often it does not, and then we say so instead of guessing.

A company can report a big improvement with no way to check it. We show the improvement and we show the gap. Neither one hides the other.

How a result becomes a card

Our agents read every eligible report and pull out each result one at a time, with the exact sentence or figure it came from. A person checks every one of those against the source before it becomes part of the record. The agents never write to the Ledger. They write to a queue, and a human decides what leaves it.

  1. A tray of published reports

    1. A published report arrives.

  2. A robot reading a glowing report

    2. Agents read it and pull out every result they can find.

  3. A printer issuing a receipt with a verified check

    3. Each result gets a receipt, the exact line in the source.

  4. A person checking a record against its source

    4. A person checks every field against the source and approves or rejects.

  5. An approved record card with a verified seal

    5. The approved record becomes a card in plain language and a full record in the research view.

What we count

A finding

One thing somebody tried, one behavior it was meant to change, one test, one number.

A comparison

An older model and its successor on the same test. It can tell you which did better. It cannot tell you why.

A baseline

One model, one test, one number, no fix involved. The plain measurements that everything else is built on.

An incident

Something went wrong in a real or test setting. We record what it touched, what safeguards were on, and how long it took to become public.

The ecosystem

How openly companies disclose, share materials, and correct themselves. Tracked on its own and never mixed into the score.

How sure are we, in three questions

We look at the evidence through three simple questions.

Who ran the test

The company that makes the model, an outside evaluator with access the company provided, or an outside group on its own.

What kind of test

A controlled experiment where one thing changed and everything else stayed the same, a comparison between models or conditions, or a single case that was observed or replayed.

Has anyone checked

No repeat test documented, the company repeated it, or an outside group repeated it and got the same answer.

We keep a letter grade in the research view for sorting. We do not show it on cards, because a letter looks like a verdict and these three questions are what the letter is made of.

Same rules both ways

Here are two results in the same form. Both have the same sections, the same size, and the same link to the evidence.

Not enough evidence yet

What we do not do

  1. We never say why a company or a critic did something. We record what they reported.

  2. We never take money from the companies that build frontier models. That rule can only change through a published governance note, in advance.

  3. We never turn a test result into a claim that AI is safe, or that alignment is solved, or that it has failed.

  4. We never print a single overall score until there is enough evidence from enough companies to earn one. The conditions are published, and today they are not met.

  5. We never delete a record quietly. Corrections and withdrawals leave a visible trace.

A stack of books and published reports

Where the evidence comes from

We read published research, company safety reports and system cards, government evaluator reports, and reports from independent testers. A news story or a post can point us to a source. It can never be the source.

When something goes wrong, we keep two clocks. One counts from the day it happened to the day it was disclosed. One counts from the day the company found it. We show both and never blend them.

Browse the source ledger

The Progress Report

Every quarter we issue a report card with ten subjects, one for each way we measure control. The marks are in words. Getting better. Getting worse. Mixed directions. Uncertain. Holding steady. Not enough evidence yet. Under every mark is a line that says what it is based on, in the same size type.

The first report card will say Not enough evidence yet on most lines. That is the honest state of the record, and we would rather show an empty card with the rules for filling it than a full one nobody should trust.

Read the current Progress Report
A stack of quarterly report cards

The ten subjects

The ten control domains we track.

  1. Deception and scheming resistance
  2. Corrigibility and shutdown compliance
  3. Instruction hierarchy
  4. Scope discipline and recklessness
  5. Autonomous action limits
  6. Jailbreak and misuse resistance
  7. Monitoring and oversight effectiveness
  8. Interpretability and internal verification
  9. Containment and infrastructure
  10. Training integrity

Mistakes, disputes, and versions

Anyone can dispute any record through a public form. While a dispute is open the card wears a badge, and when it closes the reasoning is published either way. Every card carries a version number, so if we correct something you can see what changed and when. A shared image cannot be recalled, so each one carries its version and points back to the live record.

Open the dispute form

Who runs this

The Works is run by one editor who approves every record and whose name is on each one. The agents that do the reading run under published rules. They earn more autonomy only by a long track record, they lose it on any material error, and the judgment calls stay with a person no matter what.

Read the full methodology