Learning & research

MuZero-style Planning

Explore model-based planning, search, policy telemetry and simulated job allocation.

A CLOSER LOOK

How can search evaluate an action before taking it?

For planning researchers: inspect model-based search, replay, training and evaluation in a synthetic job-allocation environment.

Follow the search

Trace the Monte Carlo tree-search implementation, including selection and value propagation. Search estimates depend on the learned model and budget.

demo/MuZero-style-v0/muzero_demo/mcts.py ↗

Understand the CLI

The demo, smoke-tests and eval subcommands have different purposes. Smoke exercises the configured model; a saved checkpoint is needed where the evaluation path expects one.

demo/MuZero-style-v0/muzero_demo/cli.py ↗

REAL REPOSITORY MATERIAL

Inspect. Understand. Reproduce.

Reading the exact source at revision 5b4cebb3.
This browser inspection does not execute the demo.

Repository source1,593 bytesDownload source ↓View on GitHub ↗

demo/MuZero-style-v0/config/muzero_demo.yaml

Select a walkthrough step to explore its source.

Full source text
# MuZero-style AGI Jobs demo configuration
owner:
  pause_planning: false
  max_capital_per_action: 2500.0
  governance_contact: "owner@agijobs.example"

experiment:
  seed: 17
  device: "cpu"
  episodes: 48
  evaluation_episodes: 64
  artifact_dir: "demo/MuZero-style-v0/artifacts"

environment:
  episode_length: 6
  job_pool_size: 5
  success_noise: 0.08
  discount: 0.997
  max_budget: 10000.0
  stochastic_fail_penalty: 0.2

network:
  observation_dim: 27
  hidden_dim: 64
  latent_dim: 48
  reward_support: [-4, 4]
  value_support: [-20, 20]
  policy_temperature: 1.0

planner:
  enable_muzero_planning: true
  default_simulations: 96
  max_simulations: 256
  exploration_constant: 1.5
  dirichlet_alpha: 0.3
  dirichlet_epsilon: 0.25
  temperature: 1.0
  visit_temperature_schedule:
    warmup_episode: 24
    min_temperature: 0.05

thermostat:
  enable: true
  low_entropy_threshold: 0.25
  high_entropy_threshold: 0.65
  min_simulations: 32
  max_simulations: 192
  latency_budget_ms: 140
  simulation_cost_ms: 1.1

sentinel:
  enable: true
  value_error_alpha: 0.1
  value_error_threshold: 0.9
  drift_window: 12
  fallback_on_violation: true

training:
  batch_size: 32
  unroll_steps: 5
  td_steps: 5
  learning_rate: 0.0008
  weight_decay: 0.000001
  replay_capacity: 2048
  warmup_steps: 32
  reanalyse_ratio: 0.25
  value_loss_weight: 0.9
  reward_loss_weight: 1.0
  policy_loss_weight: 1.0
  checkpoint_interval: 16

telemetry:
  enable: true
  flush_interval: 5
  prometheus_format: false
  sample_rate: 1.0

baselines:
  greedy_immediacy_bias: 0.05
  policy_temperature: 0.7

SHA-256 6f98c9ebe6ec98bb696d833b90c53880308246a00799c668567ca601cc66778d

FROM READING TO A REPRODUCIBLE RUN

Try the selected path.

Isolated Python environment

Use Python 3.12 in a virtual environment. Install this demo’s tracked requirements file when present, then run python -m pip check. Some variants have additional requirements: follow the selected implementation’s guide, not an unrelated demo’s dependency list.

Copy the environment setup
python3.12 -m venv .venv-muzero-style-v0
. .venv-muzero-style-v0/bin/activate
python -m pip install -r demo/MuZero-style-v0/requirements.txt --extra-index-url https://download.pytorch.org/whl/cpu
python -m pip check
Dependency files for this demo and its variants (1)
Complete environment setup ↗
SELECTED EXECUTION PATH
python demo/MuZero-style-v0/scripts/run_demo.py smoke-tests --config demo/MuZero-style-v0/config/muzero_demo.yaml

What you should observe

The smoke path checks the selected configuration and model behavior. Use the demo subcommand for the longer training journey and inspect its checkpoint before eval.

The source inspector above reads bundled repository material. Local commands run separately on your computer. Recorded examples may contain historical timestamps, placeholders and simulated metrics.

MAKE IT YOUR OWN

One useful experiment.

Compare search budgets while holding environment and seed fixed. Account for both reward and extra computation.

THE SYSTEM, MADE VISIBLE

Architecture & relationships

Architecture diagram · source preserved below
View original Mermaid source
flowchart LR
    Operators((Mission Owners)) --> demo_MuZero_style_v0[[Demo → MuZero style v0]]
    demo_MuZero_style_v0 --> Core[[AGI Jobs v0 (v2) Core Intelligence]]
    Core --> Observability[[Unified CI / CD & Observability]]
    Core --> Governance[[Owner Control Plane]]

WHEN SOMETHING DOESN’T MATCH

Troubleshooting

Import or dependency error

Confirm the active virtual environment and the selected demo’s requirements. Run python -m pip check; do not install unrelated demo requirements over a working environment.

Unexpected result or missing file

Check the selected entry point, configuration and output argument. Keep the seed and implementation fixed before comparing outcomes.

TRACE THE CHECKS

Verification & next steps

8 tracked test source files are available in this directory. Inspect the tests and their environment before choosing a suite; file counts do not establish test results.

Browse the test sources

For live commissioning, consult the production readiness record.

EVERY VARIANT, PRESERVED

Complete document library

REPRODUCE & INSPECT

Registered commands

Run commands from the repository root after following this demo's guide. Network and owner actions require their documented setup.

No root-level launch command is associated with this source path. Follow the guide or source directory for its own entry point.

Full command catalog and troubleshooting ↗

ORIGINAL DIRECTORYdemo/MuZero-style-v0View on GitHub ↗