Attempt rules beyond Phase 1
The Phase 1 protocol is fixed, but organizer-run cycles beyond it must still specify retries, resets, missed windows, and failure handling without changing what the study measures.
A proposed environment where AI agents learn to predict and control generated systems whose rules are hidden.
Exploratory research prototype—not a live AGI/ASI benchmark, and no real-world training transfer has been demonstrated. The first H-S8B one-shot completed two of eight clusters before its fixed wall cap; the overall endpoint remains INVALID and the run will never be topped up or rerun. Blind descriptive diagnostics exposed shared baseline-synchronization problems and substantial trivial-success contamination. The active adapter revision is under review; the eight-wide engine is already reviewed. The research priority is now behavioral calibration, agent-owned discovery, matched curriculum learning, and independent cross-generator/real-system transfer—not additional instrument complexity alone. Any new run requires its own frozen authority and fresh selection; no benchmark level has been earned.
The capability that matters most is untested. Nearly every benchmark measures how well a system performs on tasks humanity already understands—answering questions, writing code, solving puzzles whose rules are known. None of them directly measures the thing we would actually want from a scientifically capable AI: dropped in front of a system nobody has documented, can it figure the system out? Form competing hypotheses, choose experiments that discriminate between them, read the results correctly, revise, and then reliably control the thing. That is the skill behind all of science and most of engineering, and today it is measured almost nowhere—because any test built from human knowledge is contaminated by pretraining the moment it is published.
The bet is that this skill is trainable on worlds that share reality's structure but none of its content: alien, generated systems—hidden rules, narrow instruments, delayed and confounded effects—that cannot be memorized because there are endless fresh ones, yet statistically resemble real systems in the ways that matter. A pilot study below measured this: machines learned from real software are inert, skewed, chain-like, and punishing in ways random ones never are. And there is a genuine scientific fault line here. Synthetic training demonstrably transfers to reality for passive prediction (TabPFN, Panda), but for interactive discovery—where a bad prior produces bad experiments, which produce misleading data, compounding—it has never been shown. Reworld is designed to stand exactly on that fault line and settle it.
Reworld generates a continuing series of artificial worlds. In each one, an agent can see only a few indicators and use only a few controls; the internal state and rules are hidden. It must learn from its actions and their results to predict what will happen and reach target outcomes. The benchmark measures this behavior, not the method used inside the agent.
Seeded, reproducible software—not people or a language model—creates the worlds. Automated checks reject worlds that are broken, trivial, random, or best solved by blind search, and keep worlds where agents that learn from feedback beat specified baselines that do not. Worlds vary uncertainty, delays, competing explanations, interacting mechanisms, experiment costs, and the need to revise predictions. Tests that remove one feature at a time, followed by transfer studies, determine whether those features matter outside Reworld.
Before launch, the complete generator and selection process is fixed and published with source code, documentation, and local tools. Retired worlds, interaction records, competing explanations, and the best verified strategies become public training data—training on them is encouraged. Each official test world is selected only after agents are locked in, so no agent can have memorized it.
Internal Reworld performance is necessary but insufficient. The project succeeds only if matched studies show that Reworld training materially improves discovery and control of unfamiliar real systems. Mastery without that transfer means the generator failed as a training proxy.
The first experiment uses small finite machines that can be analyzed exactly. Its primary test asks whether a learner that uses its full interaction history beats the best fixed action sequences. Two companion tests, fixed in advance, pit it against an agent with no memory and against a copy whose own experiment record has results mismatched to the actions that produced them, under equal budgets. Deeper worlds are built only if this test passes; evolution, summary scores, and a hosted platform are out of scope for this proposal.
What Reworld must do, what it must not claim, and which constraints define the product.
| PRD field | Definition |
|---|---|
| Problem | Most training environments reward competence within known tasks or mechanics. They do not directly test whether an open curriculum helps agents use new evidence to improve prediction and control on fresh hidden systems—or whether that improvement transfers beyond the benchmark. |
| Primary users | Researchers building autonomous learning agents, and benchmark maintainers who need reproducible evidence about discovery, adaptation, and transfer. |
| Product | A permanent published world generator, a fixed narrow interaction contract, an accumulating archive of retired worlds, and isolated evaluation on a freshly selected hidden-rule world. |
| Agent task | Use actions and observations to improve prediction and control on new starting cases. Any internal representation or algorithm is permitted. |
| Project outcome | A training curriculum that measurably improves how agents learn and control independently selected unfamiliar real systems. |
| Non-goal | An internal score is not, by itself, evidence of general intelligence, real-world usefulness, or safe behavior. |
| ID | Constraint | Requirement |
|---|---|---|
| O-01 | Formal, auditable generation | World rules are executable code with recorded seeds and strict resource limits, not natural-language or LLM-authored rules. The requirement is reproducibility and a complete audit trail. |
| O-02 | Fixed interface | The numbers of inputs and outputs, their possible values, timing, reset behavior, and software interface stay the same across worlds. |
| O-03 | Limited view | The agent sees only simplified output symbols. It cannot inspect the hidden state or the rule definitions. |
| O-04 | Behavioral evidence | The agent is evaluated through prediction, choosing actions in response to results, and control on new cases—not through a textual explanation. |
| O-05 | Meaningful difficulty | Hidden machinery counts as complexity only when it causes visible differences that an agent can learn. |
| O-06 | Open-book, open-tool runs | Official estimates come from organizer-run evaluations of declared, checksummed systems on freshly selected worlds. Any harness, tool, web access, public source code, and retired-world archive is permitted and recorded; only the current world’s evaluator-private internals remain hidden behind its exposed interface. |
| O-07 | One permanent generator | All worlds come from one generator. Difficulty grows inside its world space rather than through unrelated secret generators. |
| O-08 | Open curriculum | Retired worlds and best-known strategies are released for training. Learning the generator is legitimate. |
| O-09 | Fresh-batch estimates | Official scientific estimates use batches of worlds selected by the frozen process only after evaluated programs are fixed. |
| O-10 | Test before launch | The public benchmark will be built only if prototypes show that learning works better than simply testing many action sequences. |
| O-11 | External validity | Reworld scores are compared with an independent multi-benchmark battery across diverse model families; correlation supports relevance but does not make a score equivalent to AGI. |
| O-12 | Transfer or fail | If Reworld training does not beat both strong matched control curricula on unfamiliar real systems, the frozen generator has failed its main objective. |
| O-13 | Typed evidence | Every result is labelled as a mathematical guarantee, sound class-wide bound, statistical estimate, adversarial attack result, or governance commitment. |
Internal benchmark performance is necessary; transfer outside Reworld decides project-level success.
Reworld tests whether practice on unfamiliar hidden-rule worlds helps agents adapt to unfamiliar real systems. The real systems must be independent of Reworld and must not have been used to tune it. The study must test which parts of the curriculum cause any benefit, rather than assuming that delays or hidden rules make the worlds realistic.
Proposed levels must demonstrate evidence-responsive learning and control, breadth, calibrated upper-tail discrimination, and external validity. Each level requires a named gate and a typed result; synthetic hardness or failed attacks alone establish neither AGI nor ASI. H-S8B is the supporting executable-knowledge/handoff track. Immediate priorities are bounded adapter repair and development calibration; the next research tests restore autonomous experiment choice and matched training interventions. F10 remains decisive for the transfer claim.
Test many independently developed AI systems on both Reworld and a fixed set of other benchmarks. First ask whether their rankings broadly agree. Then ask the harder question: whether Reworld predicts performance on new benchmarks and real systems beyond what model size, compute, tools, agent setup, and existing benchmark scores already explain.
| Separate claim | Required observation | Permitted interpretation |
|---|---|---|
| Internal capability | An agent improves prediction and control across the fixed Reworld process. | Increasing Reworld capability—not proof of a causal representation, AGI, or ASI. |
| Agreement with other tests (convergent validity) | Reworld rankings line up reliably with a fixed set of other benchmarks across unrelated model families. | Reworld measures capabilities shared with existing intelligence evaluations. |
| Added predictive value (incremental validity) | Reworld helps predict new results after accounting for model scale, compute, tools, agent setup, and other benchmark scores. | Reworld contributes a distinct signal about learning unfamiliar systems. |
| Training causes transfer | Under the same compute budget, Reworld training improves independently selected unfamiliar real-system tasks. | The curriculum trains a transferable capability rather than merely correlating with one. |
| Broad superhuman evidence | A system exceeds leading systems and expert-human baselines across a broad, independently governed real-world portfolio—not only Reworld. | Evidence relevant to superhuman general capability; a Reworld score alone still does not establish ASI. |
Higher-ceiling hypothesis. ARC-AGI-2 and ARC-AGI-3 help calibrate Reworld, and fresh generated worlds may keep separating stronger systems after ARC scores cluster near their human reference points—but that extra separation counts only if it predicts results on new hard evaluations and real systems. Until then, the defensible claim is only “superhuman performance within Reworld.”
Validity evidence is collected in stages rather than through one huge study: the early Phase 1C correlation probe, then agreement and added predictive value measured on the systems already evaluated in the transfer study, reusing its benchmark scores. A separate large cross-benchmark study is out of scope for this proposal.
The closest precedents split cleanly. Models trained purely on diverse synthetic generators—TabPFN on generated causal datasets, Panda on evolved chaotic systems—transfer to real data they have never seen. The working explanation is that generator diversity makes memorization impossible, so the only strategy that fits training is to become a general inference procedure. In contrast, adaptive policies trained on narrow task families reliably fail outside them, because identifying which known task is active is cheaper than learning to infer. One small interactive study bridges the camps: agents trained only on simple random synthetic environments matched agents trained on real ones. Reworld is designed to put interactive learning in the regime that transfers: one generator with unbounded structural variety, admission rules that reject worlds solvable by the tested shortcut strategies, and tests, fixed in advance, for the shortcut signature. The untested step is interaction at scale—a wrong prior produces uninformative experiments and therefore misleading data—so the plan tests transfer early instead of only at the end. An exploratory pilot has now observed this failure directly at the smallest possible scale: on eleven-state worlds whose rules change mid-test, an agent with a scrambled experiment record outperformed one holding an intact record of the old rules—confidently wrong knowledge was worse than no knowledge.
The passive-transfer, cross-generator, and calibration pilots gate progress: each pilot’s pass threshold, sample size, and analysis are fixed in a public preregistration before it runs, and all three must show positive signal before the larger world design is funded. The sealed reserve machines are examined at most twice—once per tier, whatever the reason for the retry—while the other pilots rerun on public data; further attempts need a genuinely fresh corpus, and any generator revision voids earlier confirmations, restarting the freeze. Several pilots need only public data and simple stand-alone samplers, so they can run alongside the first exact experiment rather than after it. They cannot prove transfer—only the final matched training study can—but they can disprove it cheaply, years earlier.
Receives the retired-world curriculum under the frozen training budget.
Receives the strongest available active system-identification or meta-RL curriculum.
Receives a strong broad interactive or open-ended curriculum.
Use at least two independent base-model families and eight training runs for every base-family × curriculum combination. Match architecture, optimization, training data, compute, interfaces, downstream practice, and checkpoint selection. Evaluators should not know which curriculum a system received.
Before results are seen, an independent custodian selects about 12 separately configured systems in each of at least four domains. Repeated episodes on one system are not independent samples. The analysis must account separately for variation across training runs and external systems. The input conversions connecting each real system to the agents are authored or endorsed by the same independent researchers, frozen before results are seen, and the effect must hold under at least two independently authored conversions per domain. Safety is tested separately to show that it is not meaningfully worse—a noninferiority test—and cannot be traded for capability.
| Domain | Example unfamiliar task | What is measured |
|---|---|---|
| Software systems | Diagnose and control an unfamiliar service through a limited interface. | Experiments required, prediction, recovery, and control success |
| Physical devices | Learn hidden behavior from sparse sensors and permitted interventions. | Model accuracy, action efficiency, and reliability |
| Scientific data | Select experiments and revise explanations when new measurements disagree. | Information gained, prediction error, and model revision |
| Robotics or process control | Adapt a tool or process to dynamics not seen during training. | Goal success, adaptation time, and safety violations |
If Reworld training does not produce a replicated, adapter-robust real-world advantage over both strong matched control curricula, the system solved the environment but the frozen generator failed the project’s primary objective.
The controls, indicators, timing, reset behavior, and information boundary that stay the same.
| Field | Proposed value | What it means |
|---|---|---|
| Controls | 2 × 3 values | For each step, the agent chooses a neutral value or one of two distinct interventions for each control. |
| Indicators | 3 × 2 values | After each step, every indicator is either off or on, producing eight possible joint patterns. |
| Step | One action at a time | The controls are applied, the hidden world updates a fixed number of times, and new indicators appear. |
| Reset | Repeatable | Returns to exactly the same hidden and visible starting condition for the current case. |
| History | Complete visible record | The agent can review all of its earlier control choices and visible results. |
| Visual version | No extra information | Colors, shapes, and buttons show exactly the same values as the software interface. |
No interface size is best for every world and budget. Reworld uses the smallest tested interface that shows no meaningful performance ceiling within the stated limits on worlds, agents, interaction length, and compute. Too few options waste steps encoding commands; too many make blind search easy.
Changing only one control does not guarantee that only one hidden mechanism changes; that is true only in worlds built to separate those effects. Harder worlds may require combined controls, patterns across indicators, and actions spread over time.
Candidates from one two-setting control up to three three-setting controls with four indicators were compared before freezing. The winner is adopted as a documented design decision rather than the output of a separate optimization study—no successful synthetic-training system optimized its interface—and the checks that matter ride along with the main test: every headline result reports how much information each control and indicator actually carried, how many steps went to encoding commands, and whether the result survives renamed symbols. Checks using multi-step compound actions begin in the second phase, where exact search no longer covers them. If these checks ever reveal a real ceiling, the interface changes only as a separately named benchmark version.
The panel below shows what the fixed interface might feel like, running a small hand-written toy rule set rather than a generated world.
CHOSE [0 0]SAW [0 0 0]The visual interface and software interface must provide exactly the same information. Animation, layout, timing, labels, and accessibility features must not reveal anything about the hidden world.
The main parts of a Reworld run and what remains hidden.
A run has four hidden parts—a limited internal state, a seeded starting condition, one set of world rules shared by every starting case, and a fixed publicly documented step procedure—plus the visible controls and indicators of the previous section. The starting seed is published only when the world retires.
inputs→one actionoutputsThe hidden data structure, rule format, and correctness checks.
In Phase 1, each world is a small deterministic machine—formally, an input–output transducer—with eight hidden states, nine actions, and eight possible observations. Because it is finite, researchers can calculate the best fixed action sequence and compare it exactly with small controllers that react to feedback.
If the exact test passes, worlds grow by combining and linking many small exact machines—shared hidden parts, chance outcomes, and slowly drifting settings—so that baselines and discovery paths can still be computed exactly far beyond eight states. This mirrors how the successful synthetic-training systems scaled: simple compositional recipes, not expressive world-description languages.
Every generated rule must leave the world in another valid state. The generator cannot emit arbitrary program code.
| Valid references | Every symbol a rule mentions exists, and every changed value is well defined. |
| Valid ranges | Every update stays inside its declared set of possible values. |
| Limited work | The sizes of patterns, updates, event lists, histories, and hidden steps all have fixed limits. |
| Repeatability | The first prototype produces exactly the same results when given the same world, starting number, and actions. |
| Independent replay | Two correct implementations must produce identical states and visible histories. |
The number of hidden levels will never be part of any public result: if a simple implementation and a layered implementation behave the same at the interface, they count as equally difficult.
Generate many candidates, test them, and keep only useful challenges.
Before the curriculum begins, the generator’s rules, selection process, checks, and initial archive are permanently locked, and its source, documentation, and local generation tool are released. Its own fixed algorithm can keep changing candidate worlds, but maintainers cannot tune it after seeing agent results. A substantive change would start a separately named experiment and score history, not silently alter this universe.
The core curriculum should not become harder by adding controls, adding indicators, or increasing their alphabets. Those changes would give agents a wider channel and a different machine to operate, not a deeper learning problem. The slot counts, values, timing, and reset semantics are fixed before launch. If the prototype shows that the interface is inadequate, it is changed before the permanent curriculum begins—not between difficulty levels.
A small interface can still support varied behavior over time. Sequences of actions can configure mechanisms inside a world, and later indicator sequences can report their results, without adding new external controls or indicators.
| Scaling dimension | How worlds become harder | Required evidence |
|---|---|---|
| Interacting mechanisms | More reusable parts affect one another. | Accurate prediction or control requires distinguishing more relevant states. |
| Delays and memory | Effects take longer, unfold on several timescales, or depend on earlier events. | Using the relevant history improves prediction; memoryless behavior fails. |
| Competing explanations | Several explanations fit the early evidence; important facts remain hidden or confounded. | Targeted experiments separate explanations better than passive observation or blind search. |
| Changing conditions and feedback loops | Mechanisms depend on context, form cycles, or work only in particular regimes. | Later evidence exposes limits in an earlier useful model and rewards justified revision. |
| Instrument-building and abstraction | Stable low-level patterns support higher-level variables, and action sequences assemble sensors and effectors—instruments—inside the world. | The abstraction shortens history and improves prediction or control on new cases. |
| Shared principles and theory compression | Several apparently different phenomena arise from one hidden invariant, while the budget is too small to fit each phenomenon independently. | A compact shared theory predicts held-out regions, interventions, combinations, and regimes better than local table fitting; giving that theory to a fresh agent measurably saves experiments. |
| Costly or risky experiments | Experiments acquire costs, delays, irreversible effects, and safety constraints. | Successful agents balance information, control, cost, and risk rather than probing arbitrarily. |
| Systems that change | Mostly stable rules gain rare shifts or bounded parts that react to the agent. | Detecting change and revising predictions beats a permanently fixed policy without becoming random noise. |
| New combinations of known patterns | Previously encountered patterns appear in unseen combinations and at greater depth. | Archive learning helps, but the concrete world still requires new experiments. |
Difficulty has several independent dimensions, so one level number cannot describe it fully. Candidate tests should vary one dimension at a time before combining them. More hidden nodes, rules, connections, randomness, or steps count as harder only when they demand better learnable prediction or control.
Worlds are grouped into preregistered difficulty bands. The lowest bands carry measured human reference points—ordinary players first, expert teams at the top of the human range—collected through the visual panel, so the human ceiling is located empirically. Bands above it are defined by certified demands (minimum experiments, memory, jointly coupled mechanisms) and scored as distance from the certified optimum, with no human comparison claimed; the two scales must agree where they overlap before any superhuman band is quoted. A result is a band profile, never a percentage of AGI.
Each submitted program is locked before the next world is known.
After agents are locked, a public randomness beacon deterministically derives the fresh sealed evaluation batch under frozen search and fallback rules.
The world stays hidden while agents investigate it only through the controls and indicators.
Scores and evidence are frozen when its active window closes.
Rules, starts, the result checker, adaptive strategies, and useful successful and failed histories become public.
Systems may learn from the complete archive before committing to a later world.
The generator is constructor-led, not an evolutionary arms race—and not a blind lottery either. An exploratory pilot with exact baselines found that random sampling produced one admission-eligible world in 560 tries, and that one was solvable without using feedback at all: difficulty cannot be found by rejection, only verified by it. Candidate worlds are therefore built from certified constructors carrying the measured real-system traits, with randomized content filled in around the skeleton; rejection checks remain as verification. A candidate survives only if it passes the correctness checks above and clears every measurement below. The strongest existing synthetic-to-real transfer results used exactly this filtered-diversity pattern, while agent-coupled environment-evolution systems have never demonstrated external transfer; an evolutionary loop is out of scope for this proposal.
| Metric family | Question | Failure mode detected |
|---|---|---|
| Use of history | Does knowing earlier actions and results improve prediction? | Random output or a world with no meaningful memory |
| Effect of actions | Can different actions change what the agent later sees? | A passive world in which the controls do not matter |
| Amount to remember | How many meaningfully different past situations must an agent tell apart? | A huge hidden world whose visible behavior is actually simple |
| Value of feedback | How much better is a strategy that reacts to results than a preselected action sequence? | A puzzle that can be solved by searching for one sequence |
| Experiment length | How long does it take to tell apart possible rules that require different actions? | Rules that are either obvious or impossible to distinguish |
| Learning progress | Do several kinds of learning systems improve within the allowed budget? | A challenge that is already solved or is beyond all current learners |
| Need for revision | Can later evidence expose limits in an earlier useful model? | A world that rewards one fixed theory without correction |
| Archive novelty | Does a fresh world still require new evidence after training on retired releases? | A generator whose old solutions remove the need to experiment |
Used while generating worlds. Its behavior can influence which candidates survive.
A fixed group used to estimate challenge difficulty and make scores from different worlds comparable.
Agent systems hidden from world selection and score fitting. They test whether reported difficulty is circular.
How the benchmark asks for results and checks that an agent has learned.
starting panel: always the same visible pattern
requested result:
indicator 1 must be on
indicator 2 must stay off
indicator 3 must stay on for 5 steps
limits:
a fixed number of experiments while learning
a shorter final test on new starting cases
a published minimum success rate
In Phase 1, the agent may learn across four starting cases. Evaluation targets remain hidden during learning. For each later target and starting case, the benchmark makes an identical copy of the learned agent; copies cannot share new information, but each copy may keep learning internally—including updating its own parameters—within its own test run. A hosted submission service is out of scope for this proposal.
The generator, validator, target filter, seed derivation, fallback rules, and submitted agents are committed before anyone knows the worlds.
After submissions close, use the fixed public process described under Runtime integrity. Operators cannot reorder or selectively reject valid candidates.
The agent receives a limited number of starting cases, resets, and actions with which to learn the world. Final-test targets remain hidden during this period.
The benchmark copies the agent’s parameters and memory when learning ends. Each final-test copy may react within its own case but cannot share new information with another copy.
The saved agent is tested on new starting cases, with few or no resets. It can still react to results during each case.
The system records the artifact version, resource use, visible history, outcomes, sampling unit, and uncertainty so every reported result can be checked.
| Label | What it means |
|---|---|
| Proof within stated limits (formal) | Proved or exhaustively checked for a stated model, kind of agent, target and start set, time horizon, and resource limit. Covers both exact proofs and sound class-wide bounds. |
| Statistical estimate | Estimated from a stated population, with the sampling unit, uncertainty, and multiple-test rules declared. |
| Tested shortcut or attack | Observed performance of named agents or searches. Failing to find a shortcut is not proof that none exists. |
| Auditable process rule (governance) | A rule such as locking agents, public world selection, independent custody, or isolation—not a mathematical guarantee. |
Learning from current evidence means that using the current sequence of actions and results improves prediction and control on new cases, compared with specified agents that ignore or scramble that evidence under equal budgets. This measures behavior, not the agent’s internal method. Any unexpected method counts if it generalizes under the declared test.
Generate several valid worlds that look the same at first but eventually require different actions.
For each possible result, record the next experiment needed to narrow the remaining explanations.
A reference strategy uses only visible evidence to reach targets from starting cases not used while learning.
Exact analysis where possible, plus random, preplanned, observation-blind, memoryless, and shuffled-history agents, look for easier routes.
Every claim names the agent class, resource budget, world distribution, submission limit, sampling unit, and type of evidence.
An exact best result, a proven upper bound, a statistical estimate, and the strongest attack found are different kinds of evidence; one successful action sequence never proves mastery.
Every world used for scoring receives this full verification. Training-only worlds released at scale are verified by a sample fixed in advance, with the sampling rate published; shortcut-resistance claims attach only to fully verified worlds.
Reported difficulty is based on the best route found by a published set of test agents under fixed limits. It is not proof that an easier undiscovered solution does not exist.
A retired world is released as training data: its rulebook, visible histories, alternative explanations, experiments that separated them, best verified feedback strategies, failed approaches, and reproduction details. The project should say “best verified under this budget,” not “optimal,” unless optimality was proven.
Useful solutions are adaptive: they explain what to do next for each possible result. A single successful action sequence is not an adequate discovery record.
How confirmatory evidence is estimated without pretending that one number measures everything.
The prototype will not publish one overall rating or a “percentage of AGI.” Its main result is the average, across 96 test worlds, by which the learner using history outperforms the best exact preplanned strategy that cannot react to observations. Worlds are the independent sample; targets, starting cases, and repeated agent runs are additional measurements within each world.
Δ = average over worlds (history-aware success − exact open-loop success)Analyze the result at the world level using a statistical method fixed before results are seen; everything else is reported through the capability profile below. A pass at this scale is expected for classic automata-learning algorithms—it validates the harness and the gates, not the transfer hypothesis. The same gap is also reported, descriptively, on candidate worlds before admission filtering, so readers can see how much the world filter concentrates the effect. Wherever an exact optimum or a sound bound is computable, results are also reported as distance from that certified anchor—the kind of measurement that keeps working after every comparator is saturated. Worlds built around a frozen opponent use a declared alternative anchor: best response within a named policy class, or a certified exploitability bound.
| Dimension | Reported evidence |
|---|---|
| Control on new cases | Target success and uncertainty across worlds, targets, starts, and repeated runs. |
| Prediction | Prediction error, accuracy, and confidence calibration on new cases. |
| Value of feedback | Gaps over preplanned, memoryless, small-controller, and shuffled-history comparisons. |
| Efficiency | Actions, resets, model calls, tokens, compute, and failures. |
| Revision | Performance after evidence contradicts an earlier pattern or the system changes. |
| Robustness | Performance across new world regions, agent families, equivalent formats, and renamed symbols. |
| Curriculum value | Benefit from past worlds and the remaining benefit of evidence from the current world. |
| Learning slope | Improvement per unit of experience at fixed compute, within runs and across evaluation windows. Reported descriptively, never ranked—ranking on slope would reward sandbagging. |
| Safety | Violations of stated constraints and destruction of unstated fragile mechanisms, reported separately—never traded against capability. |
Version 1 is a released benchmark, not a hosted service: worlds, harness, baselines, and analysis code that anyone can run and reproduce locally. Official estimates come from organizer-run evaluations—public programs are checksummed first, fresh worlds are selected afterward, and results stay sealed until every run closes. A hosted submission platform, spotlight events, and any single summary score are out of scope for this proposal; the multi-part capability profile is the permanent public result.
How submitted agents are locked, isolated, and checked.
| Software package | An isolated copy of the submitted agent, locked by a published checksum: a short value that changes if any committed file changes. |
| Submission time | The agent is locked before the test worlds and starting cases are selected. |
| Resources | Compute, time, storage, model calls, and tool use are recorded under declared envelopes, but benchmark hardness never depends on withholding ordinary reasoning tools. The scarce scientific resource is interaction with the current world. |
| Connections | Web search, external models, code, provers, simulators, specialist portfolios, and cooperating agents are allowed through an auditable harness. All current-world-bearing requests share the same charged evidence budget and may not leak into another entry or evaluation session. |
| World copy | Each entry gets an isolated current-world session with equivalent hidden-law and starting-case distributions; concrete internals remain available only through exposed surfaces. |
| Inside the agent | Any harness, private memory, internal tools, program changes, and cooperating AI components are allowed and declared. Successful public solvers become mandatory baselines for later levels. |
| Audit record | The system records the software version, locked files, resource use, visible history, and repeatable outcomes. |
Freeze the generator distribution, validator, target filter, seed derivation, work limit, fallback rule, and checker.
Close submissions and publish a checksum for every agent.
Use a public beacon to derive candidates in a fixed order; accept the first valid sealed evaluation batch without operator discretion.
Run learning and held-out evaluation without exposing seeds, derivation values, or audit feedback to agents.
Close the evaluation, then publish all reports, derivations, rulebooks, and complete evidence together.
Fairness comes from selecting a fresh concrete world after commitment—not from hiding how Reworld works in general.
Each later stage depends on evidence that the earlier idea works.
Define what counts as learning from current evidence, distinguish each kind of result, and fix the main world-level measure, budgets, component-removal tests, and stopping rules before seeing results.
Current: exploratory construction and instrument validationUse eight-state machines, competing rule families, exactly solved preplanned and small-controller baselines, and 64 development worlds, plus a frontier-model diagnostic arm. Before freezing, verify action budgets against known query-complexity limits and compare the planned world distribution against machines learned from real systems. Do not build the larger architecture yet.
Use separate learning and evaluation cases, cover several experiment lengths, and require the benefit of feedback to survive equivalent formats and renamed symbols.
Test cheaply whether the generator’s worlds carry real-system structure: passive prediction transfer, a measured gap to real systems, a cross-generator test, a correlation probe, direct evaluation on state machines learned from real systems, and a head-to-head test of realistic versus random practice worlds. The gating pilots must pass before Phase 2 is funded, and several can run alongside Phases 1A–1B.
Scale by combining many exact mechanisms while preserving certified evidence limits: more delays, competing explanations, drifting settings, compositional law mosaics, and shared principles that explain superficially different phenomena. Higher levels must make local fitting unaffordable and reward a compact theory on held-out regimes.
Publish the frozen worlds, harness, baselines, and analysis code for local use; run organizer evaluations on fresh sealed batches; release retired worlds; and validate the reporting model.
Compare Reworld training with several strong matched curricula on independently selected software, physical, scientific, and control problems, folding benchmark-comparison analyses into the same study—a multi-year expense committed only after every earlier gate and pilot passes.
Evolutionary generation, hidden layers, the anonymous network representation, a hosted submission platform, spotlight events, scalar ratings, and a dedicated cross-benchmark study are out of scope; any of them would be a separately versioned proposal with its own gates. Four extensions are named for later versions: literature generated from the world itself with reading scored as experiments saved, teaching-gain scoring, frozen past agents embedded as world components, and portfolio allocation across worlds.
These are binding pass conditions, not illustrative goals. Every named requirement in a gate must pass; a secondary result cannot rescue a failed primary condition.
| Gate | Test | Failure consequence |
|---|---|---|
| F1 | Across 96 worlds, the designated finite-machine learner must beat the exact best per-target fixed action sequence (held between 25% and 37.5% success by design) by ≥20 percentage points, with a one-sided 95% world-level lower bound above 10. Worlds are admitted using a separate cheap learner, and F1 counts as a stability check, not confirmatory evidence; the harder evidence sits in F2 and F4. | If not, stop before any scaled world design. |
| F2 | Breaking the link between actions and the results they caused—rerunning the learner’s own experiment record with results misaligned—must cost ≥10 control points, with a lower bound above 5, and worsen next-step prediction by ≥10%. Merely reordering whole practice sessions preserves information and is explicitly not the test. | If not, do not claim that using evidence matters. |
| F3 | On new original/modified world families, the learner’s advantage over fixed action sequences must be ≥10 points larger with the added mechanism than without it, and prediction must be harder with it; the lineage-level interval must exclude zero. | If not, the complexity is decorative; remove it. |
| F4 | The learner must beat the exact best policy that sees only the current indicators by ≥10 points, with a lower bound above 5—memory beyond the present must matter. The advantage must stay positive at every experiment depth and survive equivalent formats and renamed symbols within 5 points. Replay and compound-action controls are already covered at this scale by the exact best fixed sequence; both become binding checks at the next phase. | If not, report memory, format, or protocol specialization. |
| — | Gate numbers F5–F9 from earlier drafts were retired: the cumulative-learning study (retained learning must add ≥10 points while current evidence still adds ≥5) and the breadth report (a positive advantage in every declared world region, averaging ≥10 points, with archive-only and replay strategies below fixed limits) are now binding protocol rules rather than numbered gates, and evolution, scalar scores, and hosted deployment are out of scope. Retired numbers are not reused. | Failing either protocol rule blocks the corresponding claim. |
| F10 | For ≥2 base-model families and ≥8 runs per base × curriculum, Reworld must beat both strong controls—each endorsed by independent researchers—across about 12 systems in each of ≥4 domains: ≥10 points overall, 95% lower bound above 5, ≥5 points in every domain, and a positive effect in each base-model family. The effect must survive equivalent input conversions, safety must meet its fixed noninferiority margin, and the study must be sized in advance so a true effect of this size passes with high probability. | If not, close this generator as a failed training proxy; any redesign starts a separate record. |
What has already been tested, what was found, and what it changed — in plain language.
Before asking anyone to fund the larger programme, the riskiest assumptions were tested with experiments that can be checked exactly: studies of machines learned from real software, certified construction pilots, preregistered frontier-model probes, and a first GPU training campaign. Everything below is exploratory, not confirmatory — none of it replaces the frozen project gates — but every accepted number is reproducible from the repository, and each finding changed the plan in a stated way. The full record is the evidence ledger in the plan (Section 16).
| Question asked | What was found | What it changed |
|---|---|---|
| Do real systems have a measurable signature that random ones lack? | Across 111 machines learned from real software (six families) versus 1,110 size-matched random ones: real machines are skewed, inert, and chain-like, and 72% of their reachable situations are one-way doors that cannot return to the start, versus roughly zero in random machines. A follow-up on 1,285 independently learned TLS machines replicated the signatures at ten times the scale — and exposed one statistic as an artifact of the measuring harness rather than the machines. | The generator targets these measured signatures rather than designer intuition; the fragile statistic was excluded from the quota bands. |
| Can random generation produce good test worlds? | No. With exact baselines computable, random sampling produced one admission-eligible world in 560 attempts — and that one was defective, solvable without using feedback at all. | Worlds are built by certified constructors; random filtering is demoted to verification. Difficulty cannot be found by rejection, only verified by it. |
| Can difficulty be built to order, and does it survive assembly? | Yes. Certified difficulty components assemble into machines of up to ~200,000 states. Assembled blindly, only about half kept every advertised property — and it worsens as components are added. With two cheap checks run before assembly, 100% kept every property across all sizes tested, and verification cost stayed near-constant per moving part across a thousand-fold size range. | The build-to-order pipeline with pre-assembly certificates is the planned generator architecture; its cost-explosion tripwire was tested and not tripped. |
| Can a trained recognizer shortcut the worlds by identifying them instead of studying them? | A classifier trained on many world variants identified unseen variants only at chance level — even given more observations than the certified discovery depth. | First evidence the anti-memorization defenses work. The gate is kept permanent and must be re-run with the strongest recognizer available in each era. |
| Do the measurements measure what they claim? | Removing an agent's memory, scrambling its experiment records, and freezing its beliefs each damaged performance exactly on the world family designed to require that capability — three for three. One measurement was initially drowned out by over-harsh world design; a punishment dial whose legal range is derived from admission arithmetic, not chosen by taste, made it four times clearer. | The capability profile's cells earn their names by test, not assertion — and world severity is now a calibrated parameter. |
| Can the scoring detect an agent that is worse than no adaptivity at all? | Yes: one agent class scored zero against a 25% no-feedback floor — measurably below it. | Every score keeps its floor reference, so "worse than not trying to learn" is visible instead of hidden as a low number. |
| Is the project's central risk — misleading evidence compounding — real? | Observed directly, at the smallest possible scale: when a world's rules changed, an agent with scrambled records outperformed one holding intact records of the old rules. Confidently wrong knowledge was worse than no knowledge, at every punishment level tested. | The early transfer pilots stay mandatory before scale-up; this failure mode is now demonstrated, not hypothetical. |
| Does the verify-everything discipline actually catch anything? | Nearly every experiment caught something: a difficulty component was rejected by its own certificate (structurally sound, behaviorally hollow); an undersized budget was caught clipping its own measurement by the cost instrumentation; and a fast reimplementation of the certifier carried a silent memory-corruption bug that was caught only because the new code had to match the reference implementation to the last digit. | Certificates are always recomputed and never inferred; exact-match verification is now a standing requirement for any reimplementation. |
| Does the project’s own training loop fall for the shortcut the plan warns about? | Yes. The first student memorized repeated sequences; a fresh-data revision learned output frequencies but not dynamics. A dedicated probe found genuine in-context inference only on tiny 3–6-state machines, disappearing at 6–12 states and at real-corpus scale. | The campaign closed at a student-capacity wall with no transfer verdict. Fresh-rollout streams and a target-scale capability canary are now mandatory before another training grid is funded. |
| Can current frontier models infer hidden dynamics rather than merely recall public machines? | On 24 real-derived machines, GPT-5.2 reached 0.77 of the certified floor-to-anchor gap and GPT-5-nano reached 0.53. Skill rose with additional context. On matched freshly generated machines that never previously existed, much of the signal persisted, although the 11-pair contamination comparison was too small for a conclusive origin-invariance result. | The prediction instrument orders current systems and retains frontier headroom. A larger fresh, direct-provider study is still required before any contamination-resistant claim. |
| Does good prediction imply the ability to discover and control? | No in the first probe. GPT-5.6-terra beat random but scored at the certified floor on all six fresh control tasks. An audit found that three tasks lacked enough passive evidence; on the three fair tasks, the model still remained at or below the floor. | Control is a real unsaturated axis, but the original task was demoted to instrument validation. Successor tasks must guarantee sufficient evidence and let the agent choose experiments without collapsing into a known automata-learning exercise. |
| Can difficulty survive unlimited computation rather than just today’s attacks? | Yes for the tested construction classes, conditional on independently fresh hidden bits. Bounded information per experiment and an exact posterior-success ceiling make success rise with evidence but not with compute. Nonlinear and irreversible shells retained this guarantee; a three-family mixture defeated off-family and nearest-archive attacks, while a public three-solver portfolio still solved every case. | Upper levels now require proofs that missing evidence limits every policy, not claims that an attacker plateaued. An official run must still derive fresh bits after commitment; the open design problem is making the shell demand general investigation rather than dispatch to a published solver. |
| Can several local laws interact without collapsing back to one global puzzle type? | Yes in the preregistered Mosaic pilot. Nine worlds mixed local quadratic, substitution–permutation, and Feistel mechanisms whose transformed outputs affected downstream regions. The compositional learner recovered every sampled world exactly; each global specialist and its one-time portfolio rejected, and the nearest archived Mosaic solved 0 of 3,072 held-out cases. An independently coded verifier reproduced the result. | Compositional construction is feasible, but this remains system identification: a published Mosaic-specific table-and-graph solver succeeds with generous evidence. |
| Can one compact shared law explain different phenomena and help a future investigator? | Partly. In the preregistered shared-potential pilot, the supplied unified learner recovered all 12 potential worlds, predicted all 1,507 held-out cases per world, completed all fresh final tasks, and transferred a compact artifact that saved 24 cases and 1,152 interactions. Matched non-integrable and separate-law controls were distinguished, and an independent implementation agreed across more than 350,000 typed comparisons. | The frozen terminal verdict was finite construct pass / named reduction fail: a named vector-field baseline exceeded its permitted cap. This is positive evidence for the exact finite construction and artifact handoff, but not yet for robust unification—and still not spontaneous concept invention outside a supplied grammar. |
| Did the first active shared-law design require outcome-dependent experiment choice? | No. Before preregistration or scoring, an exact D0 analysis reduced all 6,144 public actions to 3,072 response functions and proved that every useful branch follows one fixed experiment order: one outcome preserves the shared-law theory and continues, while a mismatch excludes it and rationally stops. Therefore the optimal adaptive policy equals the best fixed order with optimal stopping; the best nonadaptive choice is to ask nothing and report uncertainty. | The design closed at stage-A construct fail—active selection or exact regret. No reward change or restricted query menu rescued it. |
| Can a native downstream task require genuinely different experiments after different outcomes? | Yes in a public D0 construction. A sealed four-leaf service panel has one failed component; testing the left junction leaves two possible faults on either outcome and makes a different leaf test optimal. Exact exhaustive search gives adaptive evidence cost 2, best fixed-order cost 9/4, and nonadaptive cost 3; with only two tests, adaptive repair succeeds always while both comparators are capped at 3/4. Independent implementations reproduce the table, policies, ceilings, native repair matrix, and a depth-three scaling fixture. | The expanded 130-action virtual surface, process-isolated RPC, reset, scalar run, batches, malformed actions, and cost sensitivity now verify independently. The conclusion is nevertheless no-go for protocol elevation or model-facing discovery claims: a public two-line specialist solves it. It remains a regression fixture. |
| Can adaptively gathered source evidence produce a committed theory that controls a fresh target? | Yes inside the declared two-phase dose-response virtual world. Adaptive source assays cost 2 versus 9/4 fixed and 3 nonadaptive; the reference investigator constructs its artifact only from its actual transcript, commits before fresh target nuisance exists, then uses one calibration marker to select a dose whose native benefit/toxicity state succeeds. Across matched controls, target-local, independently redrawn-law, and uniformly random valid artifacts each score 1/4, while the declared oracle scores 1. Independent implementations replayed 80 environments byte-identically. | Conditional virtual construction pass only. A primary-source audit of antimicrobial assays, organoid screens, and destructive tensile tests found no domain that naturally combines one-shot sources, a unique target, no replicates, no gradients or intermediate reads, discrete interventions, and comparable costs. The fixture is retained for executable-theory development, not promoted as physically embodied evidence. |
| Can reusable executable theories be tested at larger scale without mistaking verifier complexity for scientific evidence? | Not yet. A generic theory language, coupled S1–S3 worlds, source/target chronology, and independent lowering checks were built, but the attempted exact ten-row classification never ran: three rounds of adversarial review found unresolved production-integrity defects, and the effort was closed with zero scientific rows. The original high-assurance protocol remains preserved and unexecuted. | The programme now has two explicit lanes: exhaustive small-world calibration and weaker, named-system statistical evidence at scale. A development design freezes 256 independent world-family clusters, an all-required-cell endpoint, an 80% absolute floor, a 20-point paired margin, complete-cost comparisons, strong public baselines, and solver ratcheting. This is a design commitment, not a result; worlds, evaluators, baselines, models, and sealed validation have not yet been run. |
What none of this establishes: Reworld-to-reality training transfer, a frozen launchable benchmark, open-ended scientific-law discovery, or any confirmatory project claim. What it does establish is narrower and still valuable: exact construction and certification work at prototype scale; frontier prediction has measurable signal and headroom; prediction does not automatically produce control; and information-bound worlds can separate evidence from raw compute. The next design risk is construct validity—making higher levels reward shared explanatory principles rather than increasingly elaborate parameter recovery.
Questions that must be answered before the first complete specification.
The Phase 1 protocol is fixed, but organizer-run cycles beyond it must still specify retries, resets, missed windows, and failure handling without changing what the study measures.
World selection may exploit quirks in the agents used to choose candidates instead of producing broadly useful learning problems.
Training on retired worlds is intended, but fresh worlds must still require evidence about their concrete rules and starting conditions.
Transfer tests need equivalent input conversions so they measure learning rather than familiarity with Reworld’s controls and indicators.
Becoming better at controlling complex systems does not guarantee safe goals or behavior. Constraint violations require separate measurement.
Running comparison agents and checking internal solutions may cost much more than executing the final selected world.
Composing many local mechanisms does not itself require a scientific insight. Higher levels need a binding shared-law test: separate fitting must exceed the budget, while one compact principle predicts new regions, interventions, and regimes.
Then test whether training on the fixed, public curriculum preserves that advantage and beats strong alternatives on unfamiliar real systems. Even full success leaves one question no benchmark can answer—whether it matters in the open world. Only deployment answers that; saying so is what makes the rest worth defending.