WORKING REPLAY DEMO · NOT A MODEL LEADERBOARD

When evidence changes,
show why the answer changed.

A real GitHub observation. An exact commit boundary. A reproducible defect in the rules for revising a conclusion. Move through time and inspect the evidence.

1real GitHub observation
32 / 32adapter development tests
7 → 12 / 12published synthetic audit
0LLM comparison runs

AVAILABLE EVIDENCE → APPLICABILITY → CONCLUSION

QUERY CONTEXT / EXACT VERSION

RECOMPUTED LOCALLY IN THIS BROWSER

Same evidence. Explicit rules.

Original policy
before dependency fix
Candidate policy
with dependency checks
A supported claim is not permission to merge, deploy, pay, or take an external action.
Inspect the recorded results of all 12 published synthetic cases

These are open, author-designed development cases, not a held-out sample. Four variants expose one missing dependency check. The twelfth case adds explicit cycle rejection. The 7/12 → 12/12 result is not an LLM accuracy gain.

CaseExpectedOriginalCandidate
Inspect the real-source replay and what it does not prove

Five queries share one saved observation of one GitHub Actions check. The endpoint returned eight checks; this demonstration deliberately projects one named run. The source is the project owner's repository, not an independent customer deployment.

The check result is historical evidence for the specified run and commit. We do not infer the branch's required checks, trust policy, current release readiness, artifact authenticity, or source completeness. Event time is not the time this local journal received the observation.

The synthetic withdrawal scenario is not a GitHub incident. No third-party bug, revenue, independent review, emotional ability, or superiority of an LLM protocol has been measured.

QueryExpectedObserved
10,000-record stress probe, reproducibility, and next experiment

Timings cover a complete deterministic query, including evidence filtering, dependency checks, and snapshot fingerprinting. They are seven local repetitions per size, not model latency, independent trials, or a production service-level promise. Construction and validation costs are separately recorded.

The next falsifiable experiment compares R4 and R5 on identical information in flat and graph representations, with fixed whole-process budgets and new held-out episodes. A second team must first review the policy boundary and reproduce the pinned case.

Reproduce the demo ↗Review or challenge the policy ↗