Real local computation. Inspectable predictions. No
account or API key.
A capability is a hypothesis about future usefulness.
FROM THE MANUSCRIPT TO A FALSIFIABLE EXPERIMENT
What survives one mission must
earn its place in the next.
Training and future tasks are separate. The policy
makes the predictions. Failure remains part of the
record.
01 / DESIGN THE TRIAL
Make progress measurable.
Choose a future. The same learning process faces a
different test. Every figure below is computed on your
device.
Use your own series, inspect assumptions, or import a
native run
Provide 20–64 integer training observations and 2–8
distinct future tasks. Each task has a calibration
prefix and a held-out suffix. The learner only receives
training data. Public fixtures are disclosed, not a
secret benchmark.
Preparing your experiment…
02 / CAPABILITY PASSPORT
α↗
A policy, waiting to be learned.
Forty observations from Mandate A. Ten candidate
policies. No future-task answers.
Training commitment
Not frozen
Capability commitment
Not frozen
Prior learning cost
Charged in full to treatment
The frozen policy actually drives predictions. No
expected-answer lookup.
03 / FUTURE TASKS
Does the advantage travel?
AWAITING EXECUTION
Gain before human review—Modeled units, after learning, validation
and coordination
Tasks improved—Against the current-stack comparator
Actual
Current stack
With memory
Freeze a policy, then test the next mission.
Comparators · lower prediction error is
better
Arm
Total absolute error
Forecast calls
Results will appear after execution.
B0: last value · B3: fixed trend · B5:
calibration-only selection · B6: frozen memory. B1,
B2 and B4 remain unmeasured. These are deterministic
local baselines, not an LLM leaderboard.
Inspect every prediction and the cost ledger
No run yet.
04 / REVIEW THE EVIDENCE
Your judgment. An explicit record.
Inspect both arms. The timers record elapsed review
time, not verified human attention. Review attaches
to this exact run. Imported reviews stay in the
docket; start fresh timers to record a new review.
Acceptance cannot override a loss, excessive
forecast error or a missing archive.
BOUNDED TRANSFER
HOLD
Execute and review the evidence before accepting a
local claim.
Human review cost is pending.
MANUSCRIPT PROMOTION
HOLD Broader evidence remains open
Strongest-agent comparisons, independent validation,
multi-agent scaling, calibrated α-WU and delayed
real-world outcomes are still required.
Local replay establishes reproducibility.
Independent replay requires an independent process
and reviewer.
E0Simulated
E1Probed
E2Executed
E3Independent replay
E4Stressed
E5External validation
05 / TAKE THE EVIDENCE WITH YOU
A complete record. Even when it fails.
Thirteen manuscript sections. Raw tasks and
predictions, baselines, cost and safety ledgers,
capability, reviewer record, replay and calibration
status. Each file has a SHA-256 checksum.
The same inputs and frozen policy replay in
Python and JavaScript. A matching hash is not a
signature or proof of independence.
THE RESEARCH CONTINUES
Intelligence organizations. Evidence before elevation.
This lab implements a bounded part of the 198-page
AGI ALPHA: A Scalable Substrate for Intelligence
Organizations
manuscript. Explore the theory, the implementation map and
the remaining proof obligations.