Learning & research
Absolute Zero Reasoner Explore self-generated reasoning tasks, verification and learning experiments.
Code & guide What to expect This directory includes code and documentation. Its guide defines dependencies, execution modes and what the results demonstrate.
Guides & runbooks 3
Registered commands 0 Environment setup ↗ Guided tour Inspect the sources Try it locally Architecture Complete library
A CLOSER LOOK
Can a reasoning loop generate its own practice tasks? For researchers: inspect proposal, sandbox verification, solving and feedback in a bounded synthetic reasoning experiment.
01 02 Follow the default entry point The root launcher uses azr_demo. Follow task generation, the verifier, solving and wall-clock stopping. The cap bounds local demo work; it is not a security sandbox certification.
demo/Absolute-Zero-Reasoner-v0/azr_demo/__main__.py ↗ Inspect this source ↓ 03
REAL REPOSITORY MATERIAL
Inspect. Understand. Reproduce. Reading the exact source at revision 5b4cebb3. This browser inspection does not execute the demo.
demo/Absolute-Zero-Reasoner-v0/absolute_zero_reasoner_demo/config/default_config.yaml
Select a walkthrough step to explore its source.
Show more fields ↓ Full source text azr:
iterations: 50
tasks_per_iteration: 5
random_seed: 42
proposer:
base_temperature: 0.7
max_program_loc: 40
difficulty_step: 0.05
min_difficulty: 0.2
max_difficulty: 1.0
solver:
base_temperature: 0.6
accuracy_floor: 0.35
accuracy_ceiling: 0.95
improvement_rate: 0.08
buffers:
max_size_per_type: 40
rewards:
learnability_weight: 1.0
correctness_weight: 1.0
econ_weight: 0.2
format_penalty: 0.5
market:
human_hour_value_usd: 85.0
baseline_completion_minutes: 15
ai_completion_minutes: 0.5
complexity_bonus_weight: 12.0
guardrails:
target_success_rate: 0.55
target_valid_rate: 0.85
diversity_floor: 0.35
max_iterations_per_run: 500
telemetry:
report_path: reports/absolute_zero_reasoner_report.md
json_path: reports/absolute_zero_reasoner_metrics.json
mermaid_theme: forest
SHA-256 f176bcc2da5d2a09039ccd0536e57cb8b550e60d9a4d8c3d0780455c2fd781d0
The first source is shown in full. All walkthrough source links and the complete document library remain available without JavaScript.
{"revision":"5b4cebb309a83a7a6749d8911d8bf96a1921e042","sources":[{"file":"demo/Absolute-Zero-Reasoner-v0/absolute_zero_reasoner_demo/config/default_config.yaml","content":"azr:\n iterations: 50\n tasks_per_iteration: 5\n random_seed: 42\n proposer:\n base_temperature: 0.7\n max_program_loc: 40\n difficulty_step: 0.05\n min_difficulty: 0.2\n max_difficulty: 1.0\n solver:\n base_temperature: 0.6\n accuracy_floor: 0.35\n accuracy_ceiling: 0.95\n improvement_rate: 0.08\n buffers:\n max_size_per_type: 40\n rewards:\n learnability_weight: 1.0\n correctness_weight: 1.0\n econ_weight: 0.2\n format_penalty: 0.5\n market:\n human_hour_value_usd: 85.0\n baseline_completion_minutes: 15\n ai_completion_minutes: 0.5\n complexity_bonus_weight: 12.0\n guardrails:\n target_success_rate: 0.55\n target_valid_rate: 0.85\n diversity_floor: 0.35\n max_iterations_per_run: 500\n telemetry:\n report_path: reports/absolute_zero_reasoner_report.md\n json_path: reports/absolute_zero_reasoner_metrics.json\n mermaid_theme: forest\n","format":"text","sha256":"f176bcc2da5d2a09039ccd0536e57cb8b550e60d9a4d8c3d0780455c2fd781d0","bytes":894,"download":"/AGIJobsv0/examples/f176bcc2da5d2a09-default_config.yaml","source":"https://github.com/MontrealAI/AGIJobsv0/blob/5b4cebb309a83a7a6749d8911d8bf96a1921e042/demo/Absolute-Zero-Reasoner-v0/absolute_zero_reasoner_demo/config/default_config.yaml"},{"file":"demo/Absolute-Zero-Reasoner-v0/azr_demo/__main__.py","content":"\"\"\"Absolute Zero Reasoner v0 self-play demo orchestration.\"\"\"\nfrom __future__ import annotations\n\nimport argparse\nimport json\nimport time\nfrom pathlib import Path\nfrom typing import Callable, Dict, Iterable, Optional\n\nfrom .economic import EconomicSimulator\nfrom .executor import SandboxViolation, SafeExecutor\nfrom .guardrails import GuardrailManager\nfrom .policy import TRRPlusPlusPolicy\nfrom .reward import RewardEngine\nfrom .solver import SelfImprovingSolver, SolverError\nfrom .tasks import AZRTask, TaskLibrary, TaskType\nfrom .telemetry import TelemetryTracker\n\nDEFAULT_CONFIG: Dict[str, object] = {\n \"seed\": 1234,\n \"runtime\": {\"iterations\": 10, \"tasks_per_iteration\": 2},\n \"executor\": {\"time_limit\": 2.0, \"memory_limit_mb\": 256},\n \"rewards\": {\"economic_weight\": 0.15, \"format_penalty\": 0.6},\n \"guardrails\": {\"target_success\": 0.55, \"tolerance\": 0.2, \"difficulty_step\": 0.12},\n \"policy\": {\"baseline_lr\": 0.2, \"base_temperature\": 0.85},\n \"economics\": {\"base_value\": 25.0, \"difficulty_multiplier\": 45.0, \"solver_cost\": 0.05},\n}\n\n\ndef load_config(path: Optional[Path]) -> Dict[str, object]:\n if path is None:\n return DEFAULT_CONFIG\n data = json.loads(path.read_text(encoding=\"utf-8\"))\n merged = DEFAULT_CONFIG.copy()\n for key, value in data.items():\n if isinstance(value, dict) and isinstance(merged.get(key), dict):\n merged[key] = {**merged[key], **value}\n else:\n merged[key] = value\n return merged\n\n\nclass AbsoluteZeroReasonerDemo:\n \"\"\"Executable orchestrator for the user-facing demo.\"\"\"\n\n def __init__(\n self,\n config: Dict[str, object],\n *,\n max_seconds: float | None = None,\n progress_interval: int = 1,\n verbose: bool = True,\n clock: Callable[[], float] | None = None,\n ) -> None:\n self.config = config\n self.max_seconds = max_seconds\n self.progress_interval = max(1, progress_interval)\n self.verbose = verbose\n self._clock = clock or time.perf_counter\n self.executor = SafeExecutor(\n time_limit=float(config[\"executor\"][\"time_limit\"]),\n memory_limit_mb=int(config[\"executor\"][\"memory_limit_mb\"]),\n )\n self.library = TaskLibrary(seed=int(config[\"seed\"]))\n self.policy = TRRPlusPlusPolicy(\n baseline_lr=float(config[\"policy\"].get(\"baseline_lr\", 0.2)),\n base_temperature=float(config[\"policy\"].get(\"base_temperature\", 0.8)),\n )\n self.reward_engine = RewardEngine(\n economic_weight=float(config[\"rewards\"].get(\"economic_weight\", 0.1)),\n format_penalty=float(config[\"rewards\"].get(\"format_penalty\", 0.5)),\n )\n self.telemetry = TelemetryTracker(\n solver_cost=float(config[\"economics\"].get(\"solver_cost\", 0.05))\n )\n self.guardrails = GuardrailManager(\n target_success=float(config[\"guardrails\"].get(\"target_success\", 0.5)),\n tolerance=float(config[\"guardrails\"].get(\"tolerance\", 0.15)),\n difficulty_step=float(config[\"guardrails\"].get(\"difficulty_step\", 0.1)),\n )\n self.simulator = EconomicSimulator(\n base_value=float(config[\"economics\"].get(\"base_value\", 20.0)),\n difficulty_multiplier=float(\n config[\"economics\"].get(\"difficulty_multiplier\", 40.0)\n ),\n )\n self.solver = SelfImprovingSolver(self.executor)\n runtime = config[\"runtime\"]\n self.iterations = int(runtime[\"iterations\"])\n self.tasks_per_iteration = int(runtime[\"tasks_per_iteration\"])\n\n def _log(self, message: str) -> None:\n if self.verbose:\n print(message, flush=True)\n\n def _validate_task(self, task: AZRTask) -> bool:\n try:\n result = self.executor.execute(task.program, task.input_payload)\n except SandboxViolation:\n self.guardrails.register_violation()\n return False\n except RuntimeError:\n return False\n if result.timed_out or result.non_deterministic:\n return False\n if task.task_type is TaskType.DEDUCTION:\n return result.output == task.expected_output\n if task.task_type is TaskType.ABDUCTION:\n return result.output == task.expected_output\n if task.task_type is TaskType.INDUCTION:\n for example in task.io_examples:\n expected = example[\"output\"]\n observed = self.executor.execute(task.program, example[\"input\"]).output\n if observed != expected:\n return False\n return True\n return False\n\n def _verify_solution(self, task: AZRTask, answer: Dict[str, object]) -> bool:\n if task.task_type is TaskType.DEDUCTION:\n candidate = answer.get(\"answer\")\n expected = self.executor.execute(task.program, task.input_payload).output\n return candidate == expected\n if task.task_type is TaskType.ABDUCTION:\n candidate = answer.get(\"answer\")\n if candidate is None:\n return False\n payload = {\"target\": candidate}\n result = self.executor.execute(task.program, payload)\n return result.output == task.expected_output\n if task.task_type is TaskType.INDUCTION:\n program = answer.get(\"program\")\n if not isinstance(program, str):\n return False\n for example in task.io_examples:\n result = self.executor.execute(program, example[\"input\"])\n if result.output != example[\"output\"]:\n return False\n return True\n return False\n\n def run(self) -> Dict[str, object]:\n start_time = self._clock()\n iteration = 0\n while iteration < self.iterations and not self.guardrails.should_pause():\n if self.max_seconds is not None:\n elapsed = self._clock() - start_time\n if elapsed >= self.max_seconds:\n self.guardrails.pause()\n self._log(\n f\"⏱️ Wall-clock limit reached after {elapsed:.2f}s; stopping early.\"\n )\n break\n iteration += 1\n batch = self.library.sample(\n count=self.tasks_per_iteration,\n difficulty_bias=self.guardrails.state.difficulty_bias,\n )\n for task in batch:\n if not self._validate_task(task):\n continue\n start = time.perf_counter()\n temperature = self.policy.current_temperature(\"solver\", task.task_type)\n try:\n answer, format_ok = self.solver.solve(task, temperature=temperature)\n except SolverError:\n self.guardrails.register_violation()\n continue\n latency = time.perf_counter() - start\n success = False\n if format_ok:\n try:\n success = self._verify_solution(task, answer)\n except (SandboxViolation, RuntimeError):\n self.guardrails.register_violation()\n format_ok = False\n economic_value = self.simulator.estimate(\n task, success=success, latency=latency\n )\n rewards = self.reward_engine.compute(\n task_type=task.task_type,\n solver_success=success,\n economic_value=economic_value,\n format_ok=format_ok,\n )\n solver_metrics = self.policy.record(\n \"solver\", task.task_type, rewards.total_solver_reward\n )\n self.solver.update_error_rate(task.task_type, solver_metrics[\"advantage\"])\n self.policy.record(\"proposer\", task.task_type, rewards.proposer_reward)\n self.telemetry.record(\n iteration=iteration,\n task_identifier=task.identifier,\n task_type=task.task_type,\n proposer_reward=rewards.proposer_reward,\n solver_reward=rewards.total_solver_reward,\n economic_value=economic_value,\n success=success,\n )\n aggregates = self.telemetry.aggregates()\n self.guardrails.register_iteration(aggregates.get(\"success_rate\", 0.0))\n if iteration % self.progress_interval == 0:\n self._log(\n \"→ iteration {iter}: success_rate={success:.2f} gmv={gmv:.2f} \"\n \"roi={roi:.2f}\".format(\n iter=iteration,\n success=aggregates.get(\"success_rate\", 0.0),\n gmv=aggregates.get(\"gmv_total\", 0.0),\n roi=aggregates.get(\"roi\", 0.0),\n )\n )\n payload = {\n \"config\": self.config,\n \"policy\": self.policy.snapshot(),\n \"rewards\": self.reward_engine.snapshot(),\n \"guardrails\": self.guardrails.snapshot(),\n \"telemetry\": self.telemetry.aggregates(),\n \"timeline\": self.telemetry.timeline(),\n }\n return payload\n\n\ndef build_arg_parser() -> argparse.ArgumentParser:\n parser = argparse.ArgumentParser(description=__doc__)\n parser.add_argument(\n \"--config\",\n type=Path,\n help=\"Optional path to a JSON configuration overriding demo defaults.\",\n )\n parser.add_argument(\n \"--output\",\n type=Path,\n help=\"Optional path for dumping the telemetry payload as JSON.\",\n )\n parser.add_argument(\n \"--max-seconds\",\n type=float,\n help=\"Optional wall-clock limit for the run to prevent long hangs.\",\n )\n parser.add_argument(\n \"--progress-interval\",\n type=int,\n default=1,\n help=\"How frequently to print iteration progress (in iterations).\",\n )\n parser.add_argument(\n \"--quiet\",\n action=\"store_true\",\n help=\"Suppress progress logging for non-interactive environments.\",\n )\n return parser\n\n\ndef main(argv: Optional[Iterable[str]] = None) -> Dict[str, object]:\n parser = build_arg_parser()\n args = parser.parse_args(list(argv) if argv is not None else None)\n config = load_config(args.config)\n demo = AbsoluteZeroReasonerDemo(\n config,\n max_seconds=args.max_seconds,\n progress_interval=args.progress_interval,\n verbose=not args.quiet,\n )\n payload = demo.run()\n if args.output:\n args.output.write_text(json.dumps(payload, indent=2), encoding=\"utf-8\")\n return payload\n\n\nif __name__ == \"__main__\": # pragma: no cover\n main()\n","format":"text","sha256":"d8877f375338e3b51bccfd921b129c16168d02fff5ad07d76fdf3c870a5c6954","bytes":10881,"download":"/AGIJobsv0/examples/d8877f375338e3b5-__main__.py","source":"https://github.com/MontrealAI/AGIJobsv0/blob/5b4cebb309a83a7a6749d8911d8bf96a1921e042/demo/Absolute-Zero-Reasoner-v0/azr_demo/__main__.py"},{"file":"demo/Absolute-Zero-Reasoner-v0/reports/absolute_zero_reasoner_metrics.json","content":"[\n {\n \"iteration\": 0,\n \"proposer_valid_rate\": 1.0,\n \"solver_success_rate\": 0.3333333333333333,\n \"diversity_score\": 0.0,\n \"gmv_total\": 23.11,\n \"cost_total\": 0.06,\n \"roi\": 23.05,\n \"notes\": [\n \"diversity-floor-breached\"\n ]\n },\n {\n \"iteration\": 1,\n \"proposer_valid_rate\": 1.0,\n \"solver_success_rate\": 0.6666666666666666,\n \"diversity_score\": 0.4444444444444444,\n \"gmv_total\": 69.36,\n \"cost_total\": 0.12000000000000001,\n \"roi\": 69.24,\n \"notes\": []\n },\n {\n \"iteration\": 2,\n \"proposer_valid_rate\": 1.0,\n \"solver_success_rate\": 0.6666666666666666,\n \"diversity_score\": 0.48,\n \"gmv_total\": 115.94000000000001,\n \"cost_total\": 0.18,\n \"roi\": 115.76,\n \"notes\": []\n },\n {\n \"iteration\": 3,\n \"proposer_valid_rate\": 1.0,\n \"solver_success_rate\": 0.6666666666666666,\n \"diversity_score\": 0.48979591836734704,\n \"gmv_total\": 162.4,\n \"cost_total\": 0.23999999999999996,\n \"roi\": 162.16,\n \"notes\": []\n },\n {\n \"iteration\": 4,\n \"proposer_valid_rate\": 1.0,\n \"solver_success_rate\": 0.3333333333333333,\n \"diversity_score\": 0.5,\n \"gmv_total\": 185.56,\n \"cost_total\": 0.3,\n \"roi\": 185.26,\n \"notes\": []\n }\n]","format":"json","sha256":"ec3179df8714ccc78b787c13b338461e3ec593760c09040d5c1ab36b6762607a","bytes":1208,"download":"/AGIJobsv0/examples/ec3179df8714ccc7-absolute_zero_reasoner_metrics.json","source":"https://github.com/MontrealAI/AGIJobsv0/blob/5b4cebb309a83a7a6749d8911d8bf96a1921e042/demo/Absolute-Zero-Reasoner-v0/reports/absolute_zero_reasoner_metrics.json"}]}
FROM READING TO A REPRODUCIBLE RUN
Try the selected path. Python, no demo package install Run from the repository root with Python 3.12. This selected entry point uses the Python standard library. Keep output in a separate directory so that you can compare runs.
Dependency files for this demo and its variants (1) Complete environment setup ↗ SELECTED EXECUTION PATH Copy
python demo/Absolute-Zero-Reasoner-v0/run_demo.py --max-seconds 15 --output /tmp/azr-demo.jsonWhat you should observe A bounded run prints progress and writes /tmp/azr-demo.json. Compare actual iterations and solved tasks; elapsed time and simulated economics can vary.
The source inspector above reads bundled repository material. Local commands run separately on your computer. Recorded examples may contain historical timestamps, placeholders and simulated metrics.
MAKE IT YOUR OWN
One useful experiment. Reduce --max-seconds and compare completed iterations. Keep the same implementation when comparing configurations.
THE SYSTEM, MADE VISIBLE
Architecture & relationships Architecture diagram · source preserved below
View original Mermaid source flowchart LR
Operators((Mission Owners)) --> demo_Absolute_Zero_Reasoner_v0[[Demo → Absolute Zero Reasoner v0]]
demo_Absolute_Zero_Reasoner_v0 --> Core[[AGI Jobs v0 (v2) Core Intelligence]]
Core --> Observability[[Unified CI / CD & Observability]]
Core --> Governance[[Owner Control Plane]]
WHEN SOMETHING DOESN’T MATCH
Troubleshooting Import or dependency error Confirm the active virtual environment and the selected demo’s requirements. Run python -m pip check; do not install unrelated demo requirements over a working environment.
Unexpected result or missing file Check the selected entry point, configuration and output argument. Keep the seed and implementation fixed before comparing outcomes.
TRACE THE CHECKS
Verification & next steps 5 tracked test source files are available in this directory. Inspect the tests and their environment before choosing a suite; file counts do not establish test results.
Browse the test sources For live commissioning, consult the production readiness record .
EVERY VARIANT, PRESERVED
Complete document library REPRODUCE & INSPECT
Registered commands Run commands from the repository root after following this demo's guide. Network and owner actions require their documented setup.
No root-level launch command is associated with this source path. Follow the guide or source directory for its own entry point.
Full command catalog and troubleshooting ↗