Evaluate Status: Released v0.1.0

Signal Tester

Check whether your evaluation signal measures the thing you care about, or something adjacent to it.

Runs in your browser. Nothing you enter is sent to a server, and nothing is stored unless you save it to this device.

Split automatically into fragments at sentence boundaries. Each fragment gets its own classification and support check.

Cited evidence, up to four sources

Leave a source blank if you have fewer than four. Nothing here is sent anywhere.

Source 1

Leave blank if undated. An undated source is its own freshness risk, not a neutral one.

Source 2

Leave blank if undated. An undated source is its own freshness risk, not a neutral one.

Source 3

Leave blank if undated. An undated source is its own freshness risk, not a neutral one.

Source 4

Leave blank if undated. An undated source is its own freshness risk, not a neutral one.

Support map

Each claim fragment, its classification, and the strongest evidence overlap found for it.

Idle

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Source gaps

Problems with the evidence itself: unstated source types, or a factual claim resting on weak evidence.

Idle

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Freshness risk

Recency of each source. Undated is its own risk category, not a neutral middle ground.

Idle

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Confidence

A plain ratio, supported fragments over total, with the ledger of everything that moved it.

Idle

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Rewritten evidence calibrated claim

The same claim, restated to assert only what the evidence supports. Every change is explained.

Idle

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

Rater agreement, secondary panel

Optional: where two people independently judged whether each fragment is supported, check whether they agree beyond chance.

Idle

Stale Inputs changed after this result was produced. Run it again to see numbers that match what is on screen.

One pair per line: rater A value, rater B value. For categorical data use the same two labels on both sides.

Warning

This is not a fact checker

This tool evaluates the evidence you supplied. It does not search the web, does not verify a fact against the world, and a fragment marked supported only means it matches what you pasted, not that it is true.
Export
Note Why does evidence narrower than the claim matter?

A source that says "in one trial" and a claim that says "in general" can both be true on their own and still not support each other. The evidence describes one case; the claim describes every case. That gap is where an AI generated summary most often quietly overreaches, because the words sound similar even though the scope changed.

The rewritten claim below narrows every fragment like that back down to what the evidence actually showed, and explains the narrowing so the change can be checked, not just trusted.