Methods
The open harness, the two evidence modes, and the pre-registration discipline behind every result.
Every claim is tested on an open, dependency-free reference harness, under a discipline that separates what the machinery proves from what only real agents can prove.
The harness
Snapshots, weighted checks, the mint and settlement rules, a hash-chained ledger written by the environment, mock agents, and a runner for a production coding agent. Specs are pre-registered: hypothesis, method, metric, and falsifier stated before results. The results ledger is content-addressed, so a skeptic re-runs and compares hashes rather than trusting prose.
Two evidence modes
- Scripted agents
- Validates the machinery
- Can falsify the harness
- Cannot speak for real agents
- Production coding agent
- Identical protocol
- Measured tokens
- Its numbers are measurements
The two modes are never conflated: every published result is labeled with its mode, and the claim ledger tiers follow the same discipline.
With the method fixed, the results: validation, hypothesis by hypothesis.