Methods

The open harness, the two evidence modes, and the pre-registration discipline behind every result.

Every claim is tested on an open, dependency-free reference harness, under a discipline that separates what the machinery proves from what only real agents can prove.

The harness

Snapshots, weighted checks, the mint and settlement rules, a hash-chained ledger written by the environment, mock agents, and a runner for a production coding agent. Specs are pre-registered: hypothesis, method, metric, and falsifier stated before results. The results ledger is content-addressed, so a skeptic re-runs and compares hashes rather than trusting prose.

Two evidence modes

Mock mode
  • Scripted agents
  • Validates the machinery
  • Can falsify the harness
  • Cannot speak for real agents
Real mode
  • Production coding agent
  • Identical protocol
  • Measured tokens
  • Its numbers are measurements

The two modes are never conflated: every published result is labeled with its mode, and the claim ledger tiers follow the same discipline.

With the method fixed, the results: validation, hypothesis by hypothesis.