Learning & research
MuZero-style Planning Explore model-based planning, search, policy telemetry and simulated job allocation.
Code & guide What to expect This directory includes code and documentation. Its guide defines dependencies, execution modes and what the results demonstrate.
Guides & runbooks 1
Registered commands 0 Environment setup ↗ Guided tour Inspect the sources Try it locally Architecture Complete library
A CLOSER LOOK
How can search evaluate an action before taking it? For planning researchers: inspect model-based search, replay, training and evaluation in a synthetic job-allocation environment.
01 Choose a configuration Read training, search and environment settings. This file determines the experiment; do not mix it with config/default.yaml when comparing runs.
demo/MuZero-style-v0/config/muzero_demo.yaml ↗ Inspect this source ↓ 02 Follow the search Trace the Monte Carlo tree-search implementation, including selection and value propagation. Search estimates depend on the learned model and budget.
demo/MuZero-style-v0/muzero_demo/mcts.py ↗ Inspect this source ↓ 03 Understand the CLI The demo, smoke-tests and eval subcommands have different purposes. Smoke exercises the configured model; a saved checkpoint is needed where the evaluation path expects one.
demo/MuZero-style-v0/muzero_demo/cli.py ↗ Inspect this source ↓
REAL REPOSITORY MATERIAL
Inspect. Understand. Reproduce. Reading the exact source at revision 5b4cebb3. This browser inspection does not execute the demo.
demo/MuZero-style-v0/config/muzero_demo.yaml
Select a walkthrough step to explore its source.
Show more fields ↓ Full source text # MuZero-style AGI Jobs demo configuration
owner:
pause_planning: false
max_capital_per_action: 2500.0
governance_contact: "owner@agijobs.example"
experiment:
seed: 17
device: "cpu"
episodes: 48
evaluation_episodes: 64
artifact_dir: "demo/MuZero-style-v0/artifacts"
environment:
episode_length: 6
job_pool_size: 5
success_noise: 0.08
discount: 0.997
max_budget: 10000.0
stochastic_fail_penalty: 0.2
network:
observation_dim: 27
hidden_dim: 64
latent_dim: 48
reward_support: [-4, 4]
value_support: [-20, 20]
policy_temperature: 1.0
planner:
enable_muzero_planning: true
default_simulations: 96
max_simulations: 256
exploration_constant: 1.5
dirichlet_alpha: 0.3
dirichlet_epsilon: 0.25
temperature: 1.0
visit_temperature_schedule:
warmup_episode: 24
min_temperature: 0.05
thermostat:
enable: true
low_entropy_threshold: 0.25
high_entropy_threshold: 0.65
min_simulations: 32
max_simulations: 192
latency_budget_ms: 140
simulation_cost_ms: 1.1
sentinel:
enable: true
value_error_alpha: 0.1
value_error_threshold: 0.9
drift_window: 12
fallback_on_violation: true
training:
batch_size: 32
unroll_steps: 5
td_steps: 5
learning_rate: 0.0008
weight_decay: 0.000001
replay_capacity: 2048
warmup_steps: 32
reanalyse_ratio: 0.25
value_loss_weight: 0.9
reward_loss_weight: 1.0
policy_loss_weight: 1.0
checkpoint_interval: 16
telemetry:
enable: true
flush_interval: 5
prometheus_format: false
sample_rate: 1.0
baselines:
greedy_immediacy_bias: 0.05
policy_temperature: 0.7
SHA-256 6f98c9ebe6ec98bb696d833b90c53880308246a00799c668567ca601cc66778d
The first source is shown in full. All walkthrough source links and the complete document library remain available without JavaScript.
{"revision":"5b4cebb309a83a7a6749d8911d8bf96a1921e042","sources":[{"file":"demo/MuZero-style-v0/config/muzero_demo.yaml","content":"# MuZero-style AGI Jobs demo configuration\nowner:\n pause_planning: false\n max_capital_per_action: 2500.0\n governance_contact: \"owner@agijobs.example\"\n\nexperiment:\n seed: 17\n device: \"cpu\"\n episodes: 48\n evaluation_episodes: 64\n artifact_dir: \"demo/MuZero-style-v0/artifacts\"\n\nenvironment:\n episode_length: 6\n job_pool_size: 5\n success_noise: 0.08\n discount: 0.997\n max_budget: 10000.0\n stochastic_fail_penalty: 0.2\n\nnetwork:\n observation_dim: 27\n hidden_dim: 64\n latent_dim: 48\n reward_support: [-4, 4]\n value_support: [-20, 20]\n policy_temperature: 1.0\n\nplanner:\n enable_muzero_planning: true\n default_simulations: 96\n max_simulations: 256\n exploration_constant: 1.5\n dirichlet_alpha: 0.3\n dirichlet_epsilon: 0.25\n temperature: 1.0\n visit_temperature_schedule:\n warmup_episode: 24\n min_temperature: 0.05\n\nthermostat:\n enable: true\n low_entropy_threshold: 0.25\n high_entropy_threshold: 0.65\n min_simulations: 32\n max_simulations: 192\n latency_budget_ms: 140\n simulation_cost_ms: 1.1\n\nsentinel:\n enable: true\n value_error_alpha: 0.1\n value_error_threshold: 0.9\n drift_window: 12\n fallback_on_violation: true\n\ntraining:\n batch_size: 32\n unroll_steps: 5\n td_steps: 5\n learning_rate: 0.0008\n weight_decay: 0.000001\n replay_capacity: 2048\n warmup_steps: 32\n reanalyse_ratio: 0.25\n value_loss_weight: 0.9\n reward_loss_weight: 1.0\n policy_loss_weight: 1.0\n checkpoint_interval: 16\n\ntelemetry:\n enable: true\n flush_interval: 5\n prometheus_format: false\n sample_rate: 1.0\n\nbaselines:\n greedy_immediacy_bias: 0.05\n policy_temperature: 0.7\n","format":"text","sha256":"6f98c9ebe6ec98bb696d833b90c53880308246a00799c668567ca601cc66778d","bytes":1593,"download":"/AGIJobsv0/examples/6f98c9ebe6ec98bb-muzero_demo.yaml","source":"https://github.com/MontrealAI/AGIJobsv0/blob/5b4cebb309a83a7a6749d8911d8bf96a1921e042/demo/MuZero-style-v0/config/muzero_demo.yaml"},{"file":"demo/MuZero-style-v0/muzero_demo/mcts.py","content":"\"\"\"Monte Carlo tree search utilities for the MuZero-style demo.\"\"\"\nfrom __future__ import annotations\n\nfrom dataclasses import dataclass\nfrom typing import Iterable, List, Sequence, Tuple\n\nimport torch\n\nfrom .network import MuZeroNetwork, NetworkOutput\n\n\n@dataclass\nclass PlannerSettings:\n \"\"\"High level planner controls exposed to configuration.\"\"\"\n\n num_simulations: int = 32\n temperature: float = 1.0\n\n\nclass MuZeroPlanner:\n \"\"\"Lightweight planner that masks illegal actions and normalises policies.\"\"\"\n\n def __init__(self, network: MuZeroNetwork, settings: PlannerSettings | None = None) -> None:\n self.network = network\n self.settings = settings or PlannerSettings()\n\n def run(self, observation: torch.Tensor, legal_actions: Sequence[int]) -> Tuple[torch.Tensor, float, int, int]:\n \"\"\"Compute an action distribution for ``observation``.\"\"\"\n\n if observation.dim() != 1:\n observation = observation.view(-1)\n with torch.no_grad():\n output: NetworkOutput = self.network.initial_inference(observation.unsqueeze(0))\n logits = output.policy_logits.squeeze(0)\n mask = torch.full_like(logits, float(\"-inf\"))\n mask[list(legal_actions)] = 0.0\n temperature = max(self.settings.temperature, 1e-6)\n policy = torch.softmax((logits + mask) / temperature, dim=-1)\n value = float(output.value.squeeze().item())\n action = int(policy.argmax().item())\n return policy, value, action, self.settings.num_simulations\n\n @staticmethod\n def normalise_visit_counts(counts: Iterable[float]) -> List[float]:\n tensor = torch.tensor(list(counts), dtype=torch.float32)\n if tensor.sum() <= 0:\n tensor = torch.ones_like(tensor)\n tensor = tensor / tensor.sum()\n return tensor.tolist()\n\n\nclass _SearchNode:\n \"\"\"Lightweight container mirroring the tree node interface used by the planner.\"\"\"\n\n def __init__(self, value: float) -> None:\n self._value = value\n\n def value(self) -> float:\n return self._value\n\n\nclass MCTS:\n \"\"\"Simplified Monte Carlo Tree Search used by the CLI/demo entrypoints.\n\n The full MuZero search is intentionally compressed here: we estimate visit counts\n directly from the policy head and propagate the scalar value for downstream\n consumers. This keeps the public API compatible with the planner while avoiding\n a heavy dependency graph for the runnable demo script.\n \"\"\"\n\n def __init__(self, network: MuZeroNetwork, config: dict) -> None:\n self.network = network\n env_conf = config.get(\"environment\", {})\n self.action_space = int(env_conf.get(\"max_jobs\", 5)) + 1\n\n def run(self, observation: torch.Tensor, simulations: int) -> Tuple[_SearchNode, List[float]]:\n if observation.dim() != 1:\n observation = observation.view(-1)\n with torch.no_grad():\n output: NetworkOutput = self.network.initial_inference(observation.unsqueeze(0))\n policy_logits = output.policy_logits.squeeze(0)\n probabilities = torch.softmax(policy_logits, dim=-1)\n visit_counts = (probabilities * float(simulations)).tolist()\n root = _SearchNode(float(output.value.squeeze().item()))\n return root, visit_counts\n\n @staticmethod\n def final_policy(visit_counts: Sequence[float], temperature: float) -> List[float]:\n counts = torch.tensor(list(visit_counts), dtype=torch.float32)\n temperature = max(temperature, 1e-6)\n adjusted = counts / temperature\n adjusted = torch.nan_to_num(adjusted, nan=1.0 / max(len(visit_counts), 1), posinf=1.0, neginf=1e-6)\n probabilities = torch.softmax(adjusted, dim=0)\n return probabilities.tolist()\n\n\n__all__ = [\"MuZeroPlanner\", \"PlannerSettings\", \"MCTS\"]\n","format":"text","sha256":"51faffb9fff0fdff4ba3b87d3aef74a44cac8590ac565dfe4d6260249facc63d","bytes":3805,"download":"/AGIJobsv0/examples/51faffb9fff0fdff-mcts.py","source":"https://github.com/MontrealAI/AGIJobsv0/blob/5b4cebb309a83a7a6749d8911d8bf96a1921e042/demo/MuZero-style-v0/muzero_demo/mcts.py"},{"file":"demo/MuZero-style-v0/muzero_demo/cli.py","content":"\"\"\"Command-line interface for MuZero demo.\"\"\"\nfrom __future__ import annotations\n\nimport argparse\nimport json\nimport sys\nfrom pathlib import Path\nfrom typing import Dict\n\nimport torch\nimport yaml\n\nfrom .environment import JobsEnvironment, config_from_dict, vector_size\nfrom .network import MuZeroNetwork, NetworkConfig\nfrom .planner import MuZeroPlanner\nfrom .training import MuZeroTrainer, TrainingConfig\nfrom .evaluation import compare_strategies\nfrom .telemetry import TelemetrySink, summarise_runs\n\n\ndef load_config(path: str) -> Dict:\n config_path = Path(path)\n if not config_path.exists():\n raise FileNotFoundError(f\"Config file {path} not found\")\n with config_path.open(\"r\", encoding=\"utf-8\") as handle:\n config = yaml.safe_load(handle)\n experiment = config.setdefault(\"experiment\", {})\n experiment.setdefault(\"artifact_dir\", str(config_path.parent / \"..\" / \"artifacts\"))\n return config\n\n\ndef prepare_device(config: Dict) -> torch.device:\n device_name = config.get(\"experiment\", {}).get(\"device\", \"cpu\")\n if device_name == \"cuda\" and not torch.cuda.is_available():\n print(\"CUDA requested but unavailable; falling back to CPU\", file=sys.stderr)\n device_name = \"cpu\"\n return torch.device(device_name)\n\n\ndef run_demo(config_path: str) -> None:\n config = load_config(config_path)\n device = prepare_device(config)\n env_config = config_from_dict(config)\n net_conf = config.get(\"network\", {})\n network_config = NetworkConfig(\n observation_dim=int(net_conf.get(\"observation_dim\", vector_size(env_config))),\n action_space_size=env_config.max_jobs + 1,\n latent_dim=int(net_conf.get(\"latent_dim\", NetworkConfig.latent_dim)),\n hidden_dim=int(net_conf.get(\"hidden_dim\", NetworkConfig.hidden_dim)),\n )\n network = MuZeroNetwork(network_config).to(device)\n env = JobsEnvironment(env_config)\n env.seed(config.get(\"experiment\", {}).get(\"seed\", 17))\n planner = MuZeroPlanner(config, network, device)\n train_conf = config.get(\"training\", {})\n trainer = MuZeroTrainer(\n TrainingConfig(\n batch_size=int(train_conf.get(\"batch_size\", TrainingConfig.batch_size)),\n unroll_steps=int(train_conf.get(\"unroll_steps\", TrainingConfig.unroll_steps)),\n td_steps=int(train_conf.get(\"td_steps\", TrainingConfig.td_steps)),\n learning_rate=float(train_conf.get(\"learning_rate\", TrainingConfig.learning_rate)),\n weight_decay=float(train_conf.get(\"weight_decay\", TrainingConfig.weight_decay)),\n replay_capacity=int(train_conf.get(\"replay_capacity\", TrainingConfig.replay_capacity)),\n discount=float(train_conf.get(\"discount\", env_config.discount)),\n reanalyse_ratio=float(train_conf.get(\"reanalyse_ratio\", TrainingConfig.reanalyse_ratio)),\n value_loss_weight=float(train_conf.get(\"value_loss_weight\", TrainingConfig.value_loss_weight)),\n reward_loss_weight=float(train_conf.get(\"reward_loss_weight\", TrainingConfig.reward_loss_weight)),\n policy_loss_weight=float(train_conf.get(\"policy_loss_weight\", TrainingConfig.policy_loss_weight)),\n checkpoint_interval=int(train_conf.get(\"checkpoint_interval\", TrainingConfig.checkpoint_interval)),\n environment=env_config,\n ),\n network,\n device,\n )\n episodes = int(config.get(\"experiment\", {}).get(\"episodes\", 32))\n trainer.self_play(env, planner, episodes)\n for _ in range(max(episodes // 4, 1)):\n trainer.train_step()\n checkpoint_path = Path(config.get(\"experiment\", {}).get(\"artifact_dir\", \"demo/MuZero-style-v0/artifacts\")) / \"muzero_demo.pt\"\n checkpoint_path.parent.mkdir(parents=True, exist_ok=True)\n trainer.save_checkpoint(str(checkpoint_path))\n results = compare_strategies(config, network, device)\n with TelemetrySink(config) as telemetry:\n telemetry.record(\"demo_results\", results)\n print(\"=== MuZero Demo Results ===\")\n for strategy, value in results.items():\n print(f\"{strategy:>10}: {value:8.3f}\")\n print(f\"Checkpoint saved to {checkpoint_path}\")\n\n\ndef run_smoke_tests(config_path: str) -> None:\n config = load_config(config_path)\n device = prepare_device(config)\n env_config = config_from_dict(config)\n net_conf = config.get(\"network\", {})\n network = MuZeroNetwork(\n NetworkConfig(\n observation_dim=int(net_conf.get(\"observation_dim\", vector_size(env_config))),\n action_space_size=env_config.max_jobs + 1,\n latent_dim=int(net_conf.get(\"latent_dim\", NetworkConfig.latent_dim)),\n hidden_dim=int(net_conf.get(\"hidden_dim\", NetworkConfig.hidden_dim)),\n )\n ).to(device)\n env = JobsEnvironment(env_config)\n planner = MuZeroPlanner(config, network, device)\n observation = env.observe()\n action, policy, meta = planner.plan(env, observation, forced_simulations=8)\n assert 0 <= action < env.num_actions\n assert abs(sum(policy) - 1.0) < 1e-6\n trainer = MuZeroTrainer(TrainingConfig(environment=env_config), network, device)\n trainer.self_play(env, planner, episodes=1)\n metrics = trainer.train_step()\n print(json.dumps({\"plan_action\": action, \"policy\": policy, \"training_metrics\": metrics}, indent=2))\n\n\ndef run_eval(config_path: str) -> None:\n config = load_config(config_path)\n device = prepare_device(config)\n env_config = config_from_dict(config)\n net_conf = config.get(\"network\", {})\n network = MuZeroNetwork(\n NetworkConfig(\n observation_dim=int(net_conf.get(\"observation_dim\", vector_size(env_config))),\n action_space_size=env_config.max_jobs + 1,\n latent_dim=int(net_conf.get(\"latent_dim\", NetworkConfig.latent_dim)),\n hidden_dim=int(net_conf.get(\"hidden_dim\", NetworkConfig.hidden_dim)),\n )\n ).to(device)\n results = compare_strategies(config, network, device)\n print(json.dumps(results, indent=2))\n\n\nparser = argparse.ArgumentParser(description=\"MuZero-style AGI Jobs demo\")\nsubparsers = parser.add_subparsers(dest=\"command\")\n\ndemo_parser = subparsers.add_parser(\"demo\", help=\"Train and evaluate the MuZero planner\")\ndemo_parser.add_argument(\"--config\", required=True, help=\"Path to configuration YAML\")\n\nsmoke_parser = subparsers.add_parser(\"smoke-tests\", help=\"Run smoke tests for the demo\")\nsmoke_parser.add_argument(\"--config\", required=True, help=\"Path to configuration YAML\")\n\neval_parser = subparsers.add_parser(\"eval\", help=\"Evaluate strategies only\")\neval_parser.add_argument(\"--config\", required=True, help=\"Path to configuration YAML\")\n\n\ndef app(argv: list[str] | None = None) -> None:\n args = parser.parse_args(argv)\n if args.command == \"demo\":\n run_demo(args.config)\n elif args.command == \"smoke-tests\":\n run_smoke_tests(args.config)\n elif args.command == \"eval\":\n run_eval(args.config)\n else:\n parser.print_help()\n\n\nif __name__ == \"__main__\":\n app()\n","format":"text","sha256":"d70c60bdab26fb90d2496f96444ebf0b4b1ddcf3d6a5a1e65b32db9f06148420","bytes":6966,"download":"/AGIJobsv0/examples/d70c60bdab26fb90-cli.py","source":"https://github.com/MontrealAI/AGIJobsv0/blob/5b4cebb309a83a7a6749d8911d8bf96a1921e042/demo/MuZero-style-v0/muzero_demo/cli.py"}]}
FROM READING TO A REPRODUCIBLE RUN
Try the selected path. Isolated Python environment Use Python 3.12 in a virtual environment. Install this demo’s tracked requirements file when present, then run python -m pip check. Some variants have additional requirements: follow the selected implementation’s guide, not an unrelated demo’s dependency list.
Copy the environment setup python3.12 -m venv .venv-muzero-style-v0
. .venv-muzero-style-v0/bin/activate
python -m pip install -r demo/MuZero-style-v0/requirements.txt --extra-index-url https://download.pytorch.org/whl/cpu
python -m pip checkCopy Dependency files for this demo and its variants (1) Complete environment setup ↗ SELECTED EXECUTION PATH Copy
python demo/MuZero-style-v0/scripts/run_demo.py smoke-tests --config demo/MuZero-style-v0/config/muzero_demo.yamlWhat you should observe The smoke path checks the selected configuration and model behavior. Use the demo subcommand for the longer training journey and inspect its checkpoint before eval.
The source inspector above reads bundled repository material. Local commands run separately on your computer. Recorded examples may contain historical timestamps, placeholders and simulated metrics.
MAKE IT YOUR OWN
One useful experiment. Compare search budgets while holding environment and seed fixed. Account for both reward and extra computation.
THE SYSTEM, MADE VISIBLE
Architecture & relationships Architecture diagram · source preserved below
View original Mermaid source flowchart LR
Operators((Mission Owners)) --> demo_MuZero_style_v0[[Demo → MuZero style v0]]
demo_MuZero_style_v0 --> Core[[AGI Jobs v0 (v2) Core Intelligence]]
Core --> Observability[[Unified CI / CD & Observability]]
Core --> Governance[[Owner Control Plane]]
WHEN SOMETHING DOESN’T MATCH
Troubleshooting Import or dependency error Confirm the active virtual environment and the selected demo’s requirements. Run python -m pip check; do not install unrelated demo requirements over a working environment.
Unexpected result or missing file Check the selected entry point, configuration and output argument. Keep the seed and implementation fixed before comparing outcomes.
TRACE THE CHECKS
Verification & next steps 8 tracked test source files are available in this directory. Inspect the tests and their environment before choosing a suite; file counts do not establish test results.
Browse the test sources For live commissioning, consult the production readiness record .
EVERY VARIANT, PRESERVED
Complete document library REPRODUCE & INSPECT
Registered commands Run commands from the repository root after following this demo's guide. Network and owner actions require their documented setup.
No root-level launch command is associated with this source path. Follow the guide or source directory for its own entry point.
Full command catalog and troubleshooting ↗