Skip to content

See docs/DISCLAIMER_SNIPPET.md

Alpha Asi World Model

preview

Launch Demo

Current runnable path โ€” 1.14.0

Mode: Research training. Explores generated grid worlds using a small learner and local API.

Prerequisites: torch, numpy, FastAPI and the demo requirements; CPU training can be slow.

After installation:

python -m alpha_factory_v1.demos check alpha_asi_world_model
python -m alpha_factory_v1.demos run alpha_asi_world_model

Expected result: Local web service on port 7860; stop with Ctrl+C.

Scope: Grid-world research only; ASI and generalization beyond the tested environments are not established.

The catalog explains installation, stopping, backups and recovery. Browser charts for legacy demos are labeled sample replays. Original research narratives and advanced scripts below are preserved; they do not expand the tested scope stated here.

This repository is a conceptual research prototype. References to "AGI" and "superintelligence" describe aspirational goals and do not indicate the presence of a real general intelligence. Use at your own risk. Nothing herein constitutes financial advice. MontrealAI and the maintainers accept no liability for losses incurred from using this software. Each demo package exposes its own __version__ constant. The value marks the revision of that demo only and does not reflect the overall Alphaโ€‘Factory release version.

ฮฑ-ASI World-Model Demo ๐Ÿ‘๏ธโœจ

The open-ended curriculum engine + MuZero learner that powers the Alpha-Factory v1 multi-agent runtime.
Out-Learn ยท Out-Think ยท Out-Design ยท Out-Strategise ยท Out-Execute


0 Table of Contents

  1. Why this demo matters
  2. Quick-start ๐Ÿฅ‘
  3. Offline setup
  4. High-level architecture ๐Ÿ—บ๏ธ
  5. Meet the agents ๐Ÿค– (โ‰ฅ 5)
  6. Runtime controls ๐ŸŽฎ
  7. Deployment recipes ๐Ÿš€
  8. Safety, antifragility & governance ๐Ÿ›ก๏ธ
  9. Extending the demo ๐Ÿงฉ
  10. Troubleshooting ๐Ÿ”ง
  11. Production checklist โœ…
  12. License & citation

1 Why this demo matters

Missionโ€ƒProve that a constellation of agentic micro-services can independently grow their own synthetic worlds (open-ended POET curriculum), learn a general world-model (MuZero-style), automate strategy research, detect live alpha opportunities across industries, and march toward the ฮฑ-ASI referenced by Vincent Boucher, President of MONTREAL.AI and QUEBEC.AI โšก).

Success criteria โœ“

Pillar Concrete demonstration
Open-Endedness Automatic generation & evaluation of ever harder MiniWorld mazes
World-Models MuZero learner predicts reward/value & policy without ground-truth rules
Multi-Agent โ‰ฅ 5 independent Alpha-Factory agents coordinate via A2A bus
Cross-Industry Alpha StrategyAgent spots profitable โ€œalphaโ€ events (simulated market feed)
Antifragility SafetyAgent can freeze learner on NaN/spike; system self-recovers
Local-First No internet or API keys required; LLM helpers activate only if keys provided

2 Quick-start ๐Ÿฅ‘

# โ–‘ Local Python (CPU or GPU)

pip install -r requirements.txt        # torch, fastapi, uvicornโ€ฆ

# All interactive helpers (`run_ui`, `run_headless`) require these packages.

torch is by far the largest dependency. Tests that import it are skipped when the package is missing. For a short smoke test use:

pytest -m 'not e2e'
# new CLI (after `pip install -e .` at repo root)
alpha-asi-demo --demo        # same as `python -m alpha_asi_world_model_demo --demo`
alpha-asi-demo --demo --no-llm   # force-disable the optional LLM planner
python -m webbrowser http://localhost:7860  # dashboard & Swagger

# โ–‘ One-liner Docker
python -m alpha_asi_world_model_demo --emit-docker
docker build -t alpha_asi_world_model .
docker run -p 7860:7860 alpha_asi_world_model

# โ–‘ Helm (K8s)
python -m alpha_asi_world_model_demo --emit-helm
helm install alpha-asi ./helm_chart

# โ–‘ Notebook
python -m alpha_asi_world_model_demo --emit-notebook
jupyter lab alpha_asi_world_model_demo.ipynb
# โ–‘ Colab
Open `alpha_asi_world_model_colab.ipynb` in Google Colab for an end-to-end guided setup.
Nonโ€‘technical users can run it step by step:
1. Visit the notebook on GitHub and click **Open in Colab**.
2. Wait for the environment to start then choose **Runtime โ†’ Run all** (or run each cell manually).
3. The notebook installs requirements and launches the demo. When no API key is provided it automatically sets `NO_LLM=1`.
4. Interact with the dashboard in the new browser tab and run the final **Shut down** cell when done.
# โ–‘ Shell helper
./deploy_alpha_asi_world_model_demo.sh
# โ–‘ OpenAI Agents bridge
# uses ``OPENAI_API_KEY`` if set
python openai_agents_bridge.py
# โ–‘ Google ADK gateway
ALPHA_FACTORY_ENABLE_ADK=true python openai_agents_bridge.py

Set OPENAI_API_KEY to connect the bridge to the OpenAI Agents platform.

Tip ๐Ÿ’ก Set ALPHA_ASI_SEED=<int> or general.seed in config.yaml to reproduce identical curriculum runs. Tip ๐Ÿ’ก Set ALPHA_ASI_SILENT=1 to hide the startup banner.

Offline setup

When working without internet access, first build a local wheelhouse:

mkdir -p /media/wheels
pip wheel -r requirements.txt -w /media/wheels
pip wheel -r ../../../requirements-dev.txt -w /media/wheels

Install and verify using the wheelhouse from the repository root:

WHEELHOUSE=/media/wheels AUTO_INSTALL_MISSING=1 ./codex/setup.sh
WHEELHOUSE=/media/wheels AUTO_INSTALL_MISSING=1 \
  python check_env.py --auto-install --wheelhouse /media/wheels

See docs/OFFLINE_SETUP.md for a short reference.

Set NO_LLM=1 to disable the planning agent when no API key is available. The deploy_alpha_asi_world_model_demo.sh helper exports this variable automatically. Define ALPHA_ASI_LLM_MODEL=gpt-4o-mini to change the planner's model.

Device selection

config.yaml exposes a device field controlling which accelerator PyTorch uses. Accepted values are cpu, cuda and auto. With auto (the default), the demo runs on GPU when torch.cuda.is_available() returns True and falls back to CPU otherwise.


3 High-level architecture ๐Ÿ—บ๏ธ

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€ Alpha-Factory Bus (A2A) โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                                                                                        โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   curriculum   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   telemetry   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”          โ”‚
โ”‚   โ”‚ StrategyAgentโ”‚โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บโ”‚ Orchestr. โ”‚โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บโ”‚   UI / WS  โ”‚          โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                โ”‚  (loop)   โ”‚โ—„โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”‚  Interface โ”‚          โ”‚
โ”‚          โ–ฒ  โ–ฒ                     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜    commands   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜          โ”‚
โ”‚          โ”‚  โ”‚ new_env/reward                     โ–ฒ                                   โ”‚
โ”‚   plans  โ”‚  โ”‚ loss stats                        โ”‚ halt                              โ”‚
โ”‚          โ”‚  โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”            โ”‚                                   โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”   context       โ”‚            โ”‚                                   โ”‚
โ”‚   โ”‚ ResearchAgentโ”‚โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–บ Learner (MuZero) โ—„โ”€ SafetyAgent (loss guard)      โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                โ”‚   โ–ฒ                                             โ”‚
โ”‚              code patches         โ”‚   โ”‚                                             โ”‚
โ”‚   โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”                โ”‚   โ”‚ gradients                                   โ”‚
โ”‚   โ”‚ CodeGenAgent โ”‚โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜   โ”‚                                             โ”‚
โ”‚   โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                    โ”‚                                             โ”‚
โ”‚                                       โ–ผ                                             โ”‚
โ”‚                            POET Generator โ†’ MiniWorlds (env pool)                    โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
  • All messages flow through a single in-proc A2A topic bus (swap for Redis/NATS at scale).
  • MCP is used by ResearchAgent to attach rich โ€œcontext blocksโ€ to learner queries when an LLM key is supplied.
  • Components comply with OpenAI Agents SDK & Google ADK lifecycle (init/step/shutdown), so they can be re-packaged as micro-services at will.

4 Meet the agents ๐Ÿค– (โ‰ฅ 5)

Topicย ๐Ÿ›ฐ Skill How it contributes to End-to-End Alpha
planning_agent Long-horizon curriculum sketching (optionally via GPT-4o) Keeps learner near its โ€œzone of proximal developmentโ€ โ†’ faster capability gain
research_agent Literature & data mining (papers, patents, SEC filingsโ€ฆ) Injects distilled insights; helps learner transfer skills across domains
strategy_agent Real-time alpha detection (mock market feed ๐Ÿ“ˆ) Signals lucrative industry opportunities; triggers env mutations that mimic them
codegen_agent Auto-ML / network surgery Evolves MuZero hyper-params & architecture โ†’ antifragile optimisation
market_agent Streams synthetic or live financial ticks Provides cross-domain stressor; validates Alpha-capture loops
safety_agent Alignment guardrails Halts on NaN/catastrophe; enforces resource quotas & ethical policies

(If a concrete implementation is absent the stub logs every call, guaranteeing bus liveness even on a clean clone.)


5 Runtime controls ๐ŸŽฎ

REST Use case
GET /agents List active agent topics
POST /command {"cmd":"new_env"} Force-spawn a fresh world
POST /command {"cmd":"stop"} Graceful halt โธ

WebSocket (/ws) streams JSON telemetry every ui_tick steps:
{"t":1234,"r":-0.01,"loss":0.872} โ†’ plug into Grafana or a custom React chart.


6 Deployment recipes ๐Ÿš€

Target Guide
๐Ÿณ Docker Auto-generated Dockerfile (<100 MB slim). GPU builds: swap base for nvidia/cuda:runtime-12.4.
โ˜ธ๏ธ Kubernetes Run --emit-helm; edit values (replicaCount, resources.limits). Works on GKE, AKS, EKS, k3d.
๐Ÿ Pure Python No Docker needed; just pip install -r requirements.txt.
๐Ÿ”’ Air-gapped Offline wheels; set NO_LLM=1 to disable the planner or omit API keys.
๐Ÿ”‘ Cloud LLM mode Export OPENAI_API_KEY โ†’ PlanningAgent & ResearchAgent auto-upgrade to LLM assistants.

7 Safety, antifragility & governance ๐Ÿ›ก๏ธ

  • Reward-hacking firewall โ€” StrategyAgent & SafetyAgent cross-check any sudden reward spike; suspicious events quarantine the environment seed for forensic replay.
  • Loss guard โ€” Threshold loss > 1e3 or NaN triggers global stop.
  • Compute budget โ€” Learner train loop obeys torch.set_grad_enabled(False) for evaluation, cuts GPU utilisation to โ‰ค 80ย %.
  • Policy logging โ€” Every 10โ€ฏk steps, MuZero weights hashed (SHAโ€‘256) + signed for traceability.
  • Audit-ready โ€” All IPC messages dumped to ./logs/audit_<ts>.ndjson (regulator-friendly).

8 Extending the demo ๐Ÿงฉ

One-file hackability yet enterprise scalability.

  1. New env type โ†’ subclass MiniWorld (step/reset/obs), register in POETGenerator.propose.
  2. Swap learner โ†’ Implement .act/.remember/.train in a new class; StrategyAgent can trigger hot-swap via {"cmd":"swap_learner"}.
  3. External micro-service โ†’ Reuse BaseAgent; deploy as HTTP worker that bridges to A2A via WebSockets.

9 Troubleshooting ๐Ÿ”ง

Problem Cause / Fix
โ€œUI stallsโ€ Browser blocked WS โ†’ check console; ensure port 7860 reachable.
CUDA OOM export TORCH_FORCE_CPU=1 or downsize net via CodeGenAgent.
Docker build slow Add build-arg TORCH_WHL=<local-wheel> (offline).
K8s CrashLoop kubectl logs; missing GPU driver or env var.
Hide banner Set ALPHA_ASI_SILENT=1 before launching.

Need help? Open an issue โ†’ @MontrealAI/alpha-factory-core.

10 Production checklist โœ…

  • Ensure python3 --version returns 3.11โ€“3.13.
  • Install dependencies: pip install -r requirements.txt.
  • Launch via ./deploy_alpha_asi_world_model_demo.sh and visit http://localhost:7860.
  • The script sets NO_LLM=1 automatically when OPENAI_API_KEY is unset.
  • Provide an OPENAI_API_KEY to unlock planner features.
  • Set NO_LLM=1 to skip the LLM planner even when a key is provided.

11 License & citation

Apacheโ€‘2.0 ยฉ 2025 MONTREAL.AI

Please cite Alpha-Factory v1 ๐Ÿ‘๏ธโœจ โ€” Multi-Agent AGENTIC ฮฑ-AGI:

MONTREAL.AI (2025). Fully-Agentic ฮฑ-AGI: Foundation World Models for ฮฑ-ASI.
GitHub https://github.com/MontrealAI/AGI-Alpha-Agent-v0

View README on GitHub