Skip to content

Alignment with the 198-page AGI ALPHA manuscript

Version 1.8.0 treats AGI ALPHA: A Scalable Substrate for Intelligence Organizations as the latest research specification. It preserves all earlier repository material and distinguishes an executable local result from an independently established research claim.

Implementation and evidence map

Page references are PDF page numbers. A bounded implementation supports only its documented input domain.

Manuscript requirement Implementation and demonstration Evidence / remaining obligation
Intelligence organization: objective → work → validation → memory Native core/runtime/{models,engine,store}.py; Proof Bloom and Insight Atlas Persistent native missions, reviewed signed exports, browser replay and revocation; no claim of autonomous general intelligence
Objective-native compilation and proof obligations Bloom makeSeed, compile, native Mission validation Editable goals and bounded job plans; automatic general objective decomposition remains research
Distributed agents, nodes and markets Existing native CLI/API, sandbox, signed returns and chain adapter Native execution and reviewed handoffs work; no evidence of a production decentralized validator network or open market
Shared substrate, lineage and capability memory (pp. 69–70) Bloom Chronicle and transfer capability passport Hash-bound inputs, policy, learning scores, rollback and replay; general tool-contract composition remains open
Evidence ladder and honesty boundaries (pp. 36–38) This map, release evidence, complete transfer docket Local implementation/replay demonstrated; independent and delayed evidence remains pending
Evidence Contact Index (pp. 58–59) Bloom ECI gate; transfer core.evidence_contact E2 local execution only; no self-upgrade to independent E3 or external E5
RSI baseline ladder B0–B5 and adjusted advantage (p. 60) Compounding Lab B0, B3, B5 and additional treatment B6 Actual predictions, full learning cost, explicit assumed rates and review time; B1/B2/B4 unmeasured
Move-37 novelty, advantage, risk, persistence and dossier (pp. 59–60) Bloom's bounded stress gate plus transfer negative scenarios and docket Local stress is not the complete high-novelty promotion protocol; novelty and external persistence are unestablished
Freeze learning before future-task evaluation core/runtime/transfer.py:freeze and assets/compounding/engine.mjs:freeze Learner has only A as an argument. B predictions actually invoke the frozen policy. Adversarial and cross-language tests enforce this
Useful transfer / local compounding Four new B tasks per default trial; no-archive and regime-shift scenarios Positive/negative effects computed from raw errors. Disclosed synthetic tasks are not a blinded public benchmark
Complete Evidence Docket (pp. 37–38) docket_files, export_docket, verify_docket; browser ZIP export All 13 canonical sections, raw artifacts, checksums, strict replay and tamper rejection
Action-Reason traces (p. 70) Native journal; transfer 05_agialpha_runs/action_reason_trace.json Fixed experiment actions tied to commitments and outcomes; no claim of access to a model's hidden reasoning
Environment, validation, cost, safety and settlement records (p. 160) Transfer docket plus existing native sandbox/ledger/settlement evidence Exact call counts, observed whole-run time, review provenance; no transfer settlement occurs, no invented energy or monetary measurements
Human review and repair Artifact-bound accept/reject/repair; separate review timers Elapsed or operator-reported time is disclosed; identity and active attention are not externally attested
Capability promotion and rollback Transfer separates local acceptance from manuscript HOLD; Bloom transitive revocation All broad gates remain HOLD until missing comparators, scaling and independent evidence exist
Public task portfolios (pp. 158–161) Custom integer-series inputs and existing native mission schemas No claim that SWE-bench, GAIA, OSWorld or other public portfolios were run by this release
Scaling efficiency, coordination and risk Explicit coordination assumption; full cost ledger No measured multi-agent scaling law or S_N > 1 claim from a single-device experiment
α-WU calibration Docket section 11_alpha_wu_calibration/status.json Explicitly uncalibrated, value null; forecast error is not relabeled α-WU
External economy and $AGIALPHA Existing authenticated native invoice/payment verification, Ascension models Existing local EVM checks remain; transfer lab spends no funds and establishes no external economic value
Independent, stressed and delayed outcomes Separate pending fields and manuscript promotion HOLD Requires actual independent processes, reviewers and real outcomes; local keys, hashes or simulated reviewers do not suffice

Two evidence ladders and two baseline profiles

The broad evidence ladder on page 36 (E0 architecture through E10 public benchmark) and the Evidence Contact Index on pages 58–59 (E0 simulated, E1 probed, E2 executed, E3 independently replayed, E4 stressed, E5 externally validated) serve different purposes. The UI's ECI labels refer only to the latter. Executed replay on the same device does not claim E3. A signed result proves provenance under a pinned key, not organizational independence or truth of a source.

The RSI comparator profile on page 60 names B0 null, B1 incumbent, B2 neighboring system, B3 static policy, B4 strongest single agent and B5 current stack. The task-portfolio profile near page 158 uses B0–B3 for single agent, unstructured swarm, fixed roles and routed constellation. These labels are not interchangeable. The transfer protocol explicitly names manuscript-rsi-p60; its extra B6 is the treatment, not a new paper baseline.

What changed in Proof Bloom

Before 1.8.0, the local gate labeled ECI measured executed advantage. ECI now means Evidence Contact Index. The original positive-result requirement remains as a separate ADVANTAGE gate. Promotion still requires all replay, review, benchmark, stress and lineage conditions. Existing v1 Chronicle events remain replayable: their committed payload contains seed, bundles and reviews, not the display label of a derived gate.

Bloom's unchanged-benchmark reuse remains useful as a recovery/probe demonstration. The Compounding Lab adds the separate question that unchanged-benchmark reuse could not answer: does a frozen policy help on new tasks? Neither local path promotes the paper's broader claims.

Pinned upstream inspection

The source publication is pinned to bd920a6c52d820a087116bf59f2a4236d0494ac0 in agialpha-first-real-loop. That revision's treatment_control.py:_treatment_prediction calls _expected(fixture["payload"]) and does not use its capability argument. Its capability_freeze.py selects hard-coded rules by pair ID. Consequently, those fixtures cannot by themselves establish that a learned frozen capability caused the reported future-task advantage. Some other engine files at that revision are placeholders. These observations are specific to the pinned pilot; they do not dispute the manuscript's proposed architecture.

This release therefore implements an original, bounded learning experiment rather than copying those reported wins. Its treatment calls the learned forecasting policy, scores predictions against untouched suffixes and includes a no-archive control and a changed-regime failure case. The source manuscript remains unchanged and fully attributed.

Acceptance contract

The release requires native Python tests, independent Python/JavaScript implementations with exact cross-replay, complete browser journeys, cost and failure gates, stale/tampered import rejection, dossier verification, keyboard/mobile accessibility, offline recovery, byte-identical manuscript publication and all existing release gates. Cross-language replay is engineering verification; authorship within this project does not make it an independent scientific replication. See the field guide for the exact method and limitations.