Evidence ledger · 2026-09-06 · revision 1

What the tests show.
What they do not.

This page records named local campaigns and their limits. A passing campaign proves the behavior it exercised on that version. It is not an external audit, a production certification, or proof of every claim in the broader research archive.

Measured campaigns

Counts are exact for the recorded run; scope and version travel with the count.

01

Reproduced locally

Canonical languages and compiler core

9,216
transitive cross-view checks
256
quine walks
1,536
byte/word checks
2,934
CA → RV32IM differential cases

Shows: exact round trips and matching finite compiler outputs across the exercised transforms.

Does not show: general semantic understanding, completed control flow, or a finished 18-language toolchain. Ten Rosetta runners remain unfinished.

Inspect the public toolchain
02

Passed locally

GeoSeal CLI and Agent Bus harness

728 / 728
Agent Bus tests
209
CLI tests passed
38 / 38
shell protocol cases

Shows: local command correlation, capability propagation, structured denial, bounded output, and descendant timeout behavior in the tested harness.

Does not show: universal CLI superiority, production penetration-test coverage, broad hosted-provider quality, or multi-machine coordination.

Inspect the Agent Bus surface
03

Exact control

Prime/exponent coordinate

5,000 / 5,000
product composition checks
5,000 / 5,000
decoded L1 checks
5,000 / 5,000
prime-permutation controls

Shows: exact, invertible, compositional arithmetic under the tested prime-power encoding.

Does not show: a task-level advantage over a direct exponent vector. The representation is classical unique-factorization machinery used as a coordinate sidecar.

04

Experimental

Topology on learned states

Underpowered
current statistical conclusion
3+
seeds required per claim
2 controls
size-matched and no-intervention

Shows: a concrete experimental program for testing topology-guided routing and representation ideas.

Does not show: that geometry caused a model gain. Current results need stronger held-out evaluation, controls, and independent replication.

Read an underpowered result

The rule for promoting a claim

Multiple seeds, a verified evaluator, dispersion large enough to judge the effect, and wins against both a size-matched control and the no-intervention baseline.

  1. 1 Run at least three independent seeds.
  2. 2 Verify the evaluator against known-positive and known-negative cases.
  3. 3 Report dispersion; below twice pooled standard deviation stays underpowered.
  4. 4 Preserve failed and refuted branches so later work cannot rewrite history.

Trace a claim back to its boundary