Observe
Receive a few anonymous symbols. The latent state and rulebook remain hidden.
Reworld asks a machine to do something closer to science than question answering: enter an unfamiliar world, run experiments, infer its hidden causal structure, and learn to control it.
Current status: specification and falsification design. No benchmark score is live yet.
Most benchmarks present tasks in familiar representations and ask for an answer. Reworld presents a continuing causal system and asks the agent to discover what questions are worth asking.
The surface is intentionally simple and stable. Almost everything below it—the entities, relations, operators, memory, delays, and layered interactions—can change with each world.
Receive a few anonymous symbols. The latent state and rulebook remain hidden.
Manipulate fixed controls, reset instances, and design interventions that separate hypotheses.
Compress the evidence into a predictive causal model—not necessarily the simulator’s own representation.
Demonstrate understanding by reliably producing progressively deeper target observations.
Every world uses the same narrow instrument panel: a handful of manipulation slots, observation slots, a step operation, reset, target, and visible transcript.
The panel shown here is an illustrative toy, not a generated Reworld instance.
M [0 0 0 0]O [0 0 0 0 0 0]Higher layers are not decorative labels. They count only if they form grounded, low-error abstractions that improve prediction, compression, or control.
The proposal does not use an LLM to invent mechanics. It uses a bounded formal grammar, seeded mutation, deterministic validation, simulation, and a suite of search and learning systems.
retain(W)
=
Each world presents an ordered ladder of experimentally certified targets. The public score is the deepest calibrated level controlled before that world rotates.
Illustrative values. Reworld has no live leaderboard yet.
A level counts only when an observation-contingent policy succeeds across fresh instances and declared side constraints.
With one fixed interaction window and resource envelope, inefficient learning already means reaching less of the world.
One exceptional or weak world cannot dominate. World-specific solutions age out as the laws rotate.
Discovery happens under a budget. Success is then tested on fresh, post-commitment instances.
Agent artifacts are hashed before reveal, isolated, resource-bounded, and run without network or human intervention.
Difficulty comes from behavioral response curves—not latent rule count or a claim about hidden bits.
World selection, rules, instances, and level definitions are committed before activation and auditable afterward.
The public benchmark is downstream of a smaller question: can formal worlds be found where feedback-driven learning reliably beats blind sequence optimization?
Specify steps, seeds, budgets, metrics, baseline suites, and falsification thresholds.
NOWBuild tiny, exactly analyzable worlds with a hand-parameterized compositional generator.
Compare feedback policies, model learners, and strong open-loop search under equal budgets.
Add quality-diversity evolution only if the hand-built generator produces a viable difficulty band.
Retain abstractions only when they measurably improve prediction, compression, or control.
Build the fixed shell, hosted evaluator, commitments, rotation, and longitudinal validation.
Observation-contingent policies succeed where strong open-loop optimizers fail.
Prediction and control improve materially with evidence rather than remaining noise.
Accepted world complexity creates new predictive distinctions at the surface.
Methods learned on earlier worlds accelerate adaptation to held-out generator families.
Independent anchor systems induce a stable enough ordering for one depth score.
Evolution produces a better frontier than the hand-parameterized generator.
Reworld does not yet establish superiority to these systems. Its proposed contribution is the combination of generated causal laws, fixed narrow embodiment, operational control targets, one rotating public world, and explicit falsification gates.
The surface borrows the legibility of a game, but the purpose is AGI evaluation and research. Humans may inspect the interface and provide reference data; official scored runs are autonomous agent executions.
No. Candidate worlds come from a formal typed grammar, seeded mutation, deterministic validation, simulation, and selection against observable and learnability criteria.
No. Only the induced action/observation process matters. Hidden complexity that never creates a learnable surface distinction is discarded.
The visible panel, world laws, and task stay the same. Small hidden instance variation ensures that a feedback policy—not one memorized action string—is required to demonstrate control.
No benchmark is independent of compute. Reworld’s target is narrower: under a fixed execution envelope, model-building and informative experiments should outperform blind enumeration by a large, measured margin.
Each world remains finite and bounded so validity can be checked. The intended frontier is practically unbounded relative to current systems, not mathematically infinite.
No. It would measure active system identification and control under severe partial observation on a held-out generator distribution. Broader AGI claims require demonstrated transfer beyond that distribution.
Give it a body. Give it observations. Let it intervene. Then measure how much causal depth it can make intelligible.