The AI Commentator Observatory

Forty-six of the loudest voices on artificial intelligence, graded on five separate things instead of one. Sort it however you like, hide the columns you do not care about, and pivot it to see the shape of the whole field.

Here's Your Takeaway

  • There is no overall trust score here, and that is on purpose. A person can be superb on frontier capability and terrible on economics. Collapsing that into one number destroys the only information worth having.
  • Ask a better question. Not "is this person trustworthy" but: what should I believe from them, in which domain, on what evidence, and who pays them.
  • Deep technical knowledge and a serious conflict of interest travel together. The people closest to the frontier are usually employed by it. That is not a scandal, it is the structure of the field, and the table lets you see both columns at once.
  • Where a grade and a measured track record disagree, believe the track record. Nine of these people also appear in the prediction scorecard, and two of them look very different once you count what actually came true.

The explorer

It opens on the question most people arrive with: who actually understands this, and who is paid to reach a conclusion. That is the Technical column sorted high to low, with Conflict sitting right beside it. Change any of it.

View
Role
Conflict Measured
Group by
Columns

Grades come from the research package described at the bottom. Measured accuracy is a different kind of evidence entirely: it comes from the prediction database behind the AI Predictions Scorecard, which scores individual dated predictions on direction, timing and magnitude. Where somebody has only one or two scored predictions the number is marked thin, because one lucky call is not a track record. Turn on the Direction called right column for the blunter measure: forget how close the timing was, how often did they even get the direction right?

What the five columns actually mean

ColumnThe question it answersWhat a high grade does not mean
TechnicalDo they understand how the systems work, and do their capability claims survive testing?That their economic or social predictions are any good.
ForecastingHow good are their explicit, dated, falsifiable predictions?That they are right about today. Forecasting and description are different skills.
UpdatingDo they visibly change position when the evidence changes?That they were right the first time. Updating well often means having been wrong.
TransparencyDo they disclose their money and their relationships?That there is no conflict. It means the conflict is visible, which is the next column.
ConflictCould their income or position benefit from the conclusion they are reaching?That they are lying. Extreme conflict plus high transparency is an honest, interested party.

The trap this table is built to avoid. A single trust score would let you skip the thinking. It would also be wrong: it would average a person's superb frontier knowledge together with their weak macroeconomics and hand you a number that describes neither. Every column here stands alone, and none of them add up.

How far is the claim from the evidence?

The most useful idea in the whole research package, and the one you can apply without any table at all. A statement can be completely true and the conclusion drawn from it enormous. Confidence should fall as you go down this ladder.

    Worked example, using a real event from this year. "Astra reached OpenAI's critical cyber threshold" sits at distance 1. "AI can now hack anything" is distance 3. "Cybersecurity as a profession is finished" is distance 4. "AI will destroy civilisation" is distance 5. The first one is checkable. The last one is a feeling wearing a fact's clothing.

    What AI can actually do right now

    Prediction arguments get much shorter when both sides agree on the baseline. This is the current state, separated into what is established, what has been demonstrated in controlled conditions, and what remains unresolved.

    The events that moved the baseline

    A capability-update event is a moment when something previously argued about became something observed. These are the ones the research flags, with the evidence and the reading kept separate.

    Predictions on the record

    Falsifiable, dated, and therefore worth something. A forecast with no date and no threshold cannot be wrong, which is why it is also worthless.

    Who pays for the explaining

    Funding is not one thing, and treating a VPN advertisement as equivalent to an advocacy campaign is how this analysis usually goes wrong. These are the distinct mechanisms.

    Documented cases

    Channels examined

    Methodology, and what this is not

    These grades are provisional research judgments, not measurements. They come from a research pass completed in September 2026, they carry no confidence intervals, and reasonable people would grade several of these differently. Treat them as a starting point for your own reading, never as a verdict.

    Sources