A project of ISTC-CNR
Istituto di Scienze e Tecnologie della Cognizione

A project of ISTC-CNR

Many planning experiments. One shared language.

Planning 101 brings published experiments on human planning into a single, validated data model. Each task is described as a Markov decision process and each move is recorded as a Step, so an analysis written once can run across studies, for human participants and artificial agents alike.

Every move is a Step Loading map…
s_pre—
action—
s_post—
Make a move to record a Step

Click a dot joined to the token.

A simplified map in the style of ThinkAhead (Eluchans et al., 2025). The state and action strings use the same canonical format as the Planning 101 data model.

9published experiments selected
5already with conversion rules
8entities in one data model
3tiers of validation

Why Planning 101

Planning has been studied in many ways. The data rarely fits together.

Researchers have studied how people plan with graph navigation, two-stage decision games, subway routing, mazes and board games. Each dataset is usually released in its own laboratory's format and processed with study-specific scripts, where the conversion from raw data to analysed variables is written inline rather than documented. Combining studies then runs into two obstacles that are hard to detect after the fact.

Silent information loss

A dropped column, an unrecorded exclusion or a filtered timeout leaves no trace in the processed file. The file looks complete, so a later analyst cannot see what is missing.

Incompatible states

The same logical state, written two different ways in two studies, does not match when the data are joined. Paradigms that share a structure cannot be compared on it.

Planning 101 addresses the representation first. Cross-paradigm modelling, comparisons between human and model behaviour, and a review of which aspects of planning have and have not been studied all depend on the data sitting in one validated, queryable structure.

How it works

From a published dataset to a harmonised one, in three steps.

Every conversion decision is written down, versioned and checked, rather than left in a notebook cell.

STEP 01

Describe the task

Each paradigm is specified as a Markov decision process: its states, actions, transitions, rewards, and what the participant can observe.

Variants that share a structure share a Task. The magic-carpet and spaceship versions of the two-step task, for example, differ only in their cover story, which is kept as a recorded parameter rather than discarded.

STEP 02

Write the rules down

A YAML manifest states how a source dataset maps into the schema: which files, which columns, which exclusions, and which published result the converted data must reproduce.

The mapping is explicit, reviewable and versioned, in the spirit of BIDS in neuroimaging.

STEP 03

Validate, then write

Schema and chain checks first; then, where a task's dynamics are known, checks that every move is legal and every outcome reachable. A dataset that fails them is not written.

Reproduction targets re-derive a published result from the converted data. Output is Parquet or CSV with a provenance record: schema and manifest versions, source commit, content hash.

manifests/thinkahead_trial_events_json.yaml · excerpt
streams:
  - role: steps                 # each event becomes a Step
    record_path: "TrialEventData"
    state_fields: ["touchedNode"]
    onset: {kind: field, field: "time", unit: s}
  - role: trace                 # pointer samples
    record_path: "UserTrajectory"
    channels: {x: "x", y: "y", z: "z"}
  - role: episode_extras        # trial summary
    fields: {trial_score: "TrialScore"}
In the committed example, one trial file becomes22 Stepspointer Trace rowsEpisode extras

Rules every conversion follows

Each one addresses a documented failure mode, and is enforced in code or checked in review.

  • No data silently dropped. Invalid steps are flagged, not removed; filtering is an analysis-time choice.
  • Nothing fabricated. Values are read from the source; gaps stay empty or use a documented sentinel.
  • Hidden variables stay hidden. Latent parameters, such as drifting reward probabilities, are kept out of the participant's observable state.
  • Every derived value is versioned. Unit conversions and derived clocks are recorded in the provenance.
  • Sources are pinned. Each conversion records the source commit and the manifest's content hash.

The data model

Three layers, eight entities.

Each Step belongs to one Episode, each Episode to one Session, each Session to one Subject and one Experiment, and each Experiment to one Task. Because these links are enforced, analysis code written against them holds for any converted dataset.

Specification
TaskThe abstract MDP: states, actions, transitions, rewards, observability.
Deployment
ExperimentA concrete instantiation of a Task.
SubjectA human participant or a synthetic agent.
SessionOne Subject's engagement with one Experiment.
Trace
EpisodeOne trajectory from an initial to a terminal state.
StepOne transition, from state to state through an action.
TraceSamples within a Step, such as pointer or gaze.
MeasureProbes on the task that are not moves, such as subgoal reports.
Anatomy of a Stepone row of the Step table
step_index
3
t_step_onset_ms
4120.0
s_pre
{"node":"n_4","remaining":["g_2","g_7"]}
decision
nullKept apart from the action, and never imputed at conversion.
action
{"to":"n_5","type":"move"}
s_post
{"node":"n_5","remaining":["g_2","g_7"]}
reward
0.0
valid
trueInvalid steps are marked, not removed.
extras
{…}Paradigm-specific fields that do not generalise.

Humans and agents share one table. A planning agent is a Subject with subject_type = "agent", so comparing human and model behaviour reads from one structure instead of two parallel ones.

The experiments

Nine published studies, each adding something new to the format.

The selection, agreed in September 2026, starts from a homogeneous initial group and adds paradigms in order of coding effort. Five already have conversion rules; four are next.

Rules readyP1

01ThinkAhead

Collect every gem on a map of dots joined by lines. Some maps are solved at a glance, others need several moves of look-ahead. Measures how far ahead people plan as difficulty grows.

Eluchans et al. (2025), Royal Society Open ScienceDOI ↗
Rules readyP1

02Two-step task

Two choices in a row: a magic carpet or spaceship that usually leads to one place, then an option that pays off with a slowly drifting probability. Tests whether people build and use a map of the game, and how much the instructions matter.

Feher da Silva & Hare (2020), Nature Human BehaviourDOI ↗
Rules readyP2

03Reduced two-step task

The same two-move game with one rule simplified: the first choice always leads to the same place. Asks when it pays to reason ahead instead of repeating what worked last time.

Kool, Cushman & Gershman (2016), PLOS Comput. Biol.DOI ↗
Rules readyP2

04Chunking

Travel through a network of places grouped into neighbourhoods. Tests whether people split the map into neighbourhoods and plan between them first, then within each.

Tomov et al. (2020), PLOS Comput. Biol.DOI ↗
Rules readyP3

05Task decomposition

Learn an eight-place map by moving through it, then name the stop you would aim for first. Compares the declared subgoals with the paths people actually take.

Correa et al. (2023), PLOS Comput. Biol.DOI ↗
PlannedP3

06Value-guided construal

Steer through mazes to a goal, then report which obstacles you noticed. Shows that people plan with a simplified picture that keeps only the obstacles relevant to the route.

Ho et al. (2022), NatureDOI ↗
PlannedP4

07Navigation with rollouts

Find a hidden reward in a maze whose edges wrap around, then return to it from somewhere else. Measures how long people pause to think before each move.

Jensen, Hennequin & Mattar (2024), Nature NeuroscienceDOI ↗
PlannedP4

08Mouselab-MDP

Uncover values hidden in a web of boxes, where every click costs, then choose a three-move path. Observes which information people pay for before they commit.

Callaway et al. (2022), Nature Human BehaviourDOI ↗
PlannedP5

09Four-in-a-row

Two players compete on a board of four rows by nine columns, in some sessions with eye tracking. Measures how many moves ahead players think and how this grows with practice.

van Opheusden et al. (2023), NatureDOI ↗

“P1–P5” is the proposed implementation phase, by increasing coding effort. Each dataset remains the work of its original authors and is subject to its own licence.

The experiments site

See it working on real data.

The experiments site is an interactive application built on the Planning 101 pipeline. Browse the selected studies, check how each conversion rule turns original rows into Episodes and Steps, map a dataset of your own, and replay recorded sessions.

Open the experiments
experiment.planning101.app

Experiments Explorer

The selected studies with their papers, original datasets and rule status. Browse any conversion episode by episode, original rows beside the converted Steps.

Custom Mapping

Upload a CSV or JSON file in any layout and map its columns visually to subject, state, action, reward and timing. Save the result as a new rule.

Upload & Convert

Convert files with an existing rule, several at once with one per subject, and compare the result with the original.

Build database

Combine converted datasets into one consolidated table to query and chart.

Simulators

Replay a participant's recorded Magic Carpet session as they played it, trial by trial.

Start exploring

No installation needed.

Who we are

A project of ISTC-CNR.

Planning 101 is a project of the Institute of Cognitive Sciences and Technologies (ISTC) of the National Research Council of Italy (CNR).

Consiglio Nazionale delle Ricerche ISTC-CNR, Istituto di Scienze e Tecnologie della Cognizione
ISTC
Istituto di Scienze e Tecnologie della Cognizione
Institute of Cognitive Sciences and Technologies
CNR
Consiglio Nazionale delle Ricerche
National Research Council of Italy

The team

Giovanni Pezzulo
Mattia Eluchans
Gianguglielmo Calvi

The experiments come from the authors of the original studies, credited above. Planning 101 builds on their published data and code.

Ready when you are

Start exploring.

Open the experiments site to browse the studies, inspect the conversions and replay recorded sessions.

Go to the experiments
experiment.planning101.app