WORKING REPLAY DEMO · NOT A MODEL LEADERBOARD
When evidence changes,
show why the answer changed.
A real GitHub observation. An exact commit boundary. A reproducible defect in the rules for revising a conclusion. Move through time and inspect the evidence.
AVAILABLE EVIDENCE → APPLICABILITY → CONCLUSION
RECOMPUTED LOCALLY IN THIS BROWSER
Same evidence. Explicit rules.
before dependency fix
with dependency checks
Inspect the recorded results of all 12 published synthetic cases
These are open, author-designed development cases, not a held-out sample. Four variants expose one missing dependency check. The twelfth case adds explicit cycle rejection. The 7/12 → 12/12 result is not an LLM accuracy gain.
| Case | Expected | Original | Candidate |
|---|
Inspect the real-source replay and what it does not prove
Five queries share one saved observation of one GitHub Actions check. The endpoint returned eight checks; this demonstration deliberately projects one named run. The source is the project owner's repository, not an independent customer deployment.
The check result is historical evidence for the specified run and commit. We do not infer the branch's required checks, trust policy, current release readiness, artifact authenticity, or source completeness. Event time is not the time this local journal received the observation.
The synthetic withdrawal scenario is not a GitHub incident. No third-party bug, revenue, independent review, emotional ability, or superiority of an LLM protocol has been measured.
| Query | Expected | Observed |
|---|
10,000-record stress probe, reproducibility, and next experiment
Timings cover a complete deterministic query, including evidence filtering, dependency checks, and snapshot fingerprinting. They are seven local repetitions per size, not model latency, independent trials, or a production service-level promise. Construction and validation costs are separately recorded.
The next falsifiable experiment compares R4 and R5 on identical information in flat and graph representations, with fixed whole-process budgets and new held-out episodes. A second team must first review the policy boundary and reproduce the pinned case.