How should future neural reasoning systems implement extended computation? Recursive Reasoning Models (RRMs) offer a promising alternative to autoregressive sequence extension by performing iterative latent-state refinement with shared transition functions. Yet existing RRMs are largely deterministic, following a single latent trajectory and converging to a single prediction. We introduce Generative Recursive reAsoning Models (GRAM), a framework that turns recursive latent reasoning into probabilistic multi-trajectory computation. GRAM models reasoning as a stochastic latent trajectory, enabling multiple hypotheses, alternative solution strategies, and inference-time scaling through both recursive depth and parallel trajectory sampling. This yields a latent-variable generative model supporting conditional reasoning via p(y|x) and, with fixed or absent inputs, unconditional generation via p(x). Trained with amortized variational inference, GRAM improves over deterministic recurrent and recursive baselines on structured reasoning and multi-solution constraint satisfaction tasks, while demonstrating an unconditional generation capability.
Existing RRMs are fundamentally deterministic: given the same input and initialization, they follow a single latent trajectory and converge to a single prediction. A capable reasoning system should be not only deep (repeated refinement) but also wide (multiple parallel trajectories). GRAM treats the reasoning process itself as a stochastic latent trajectory: at each recursion step, the model samples a transition conditioned on the input and the current reasoning state, rather than deterministically updating to a single next state.
Deterministic vs. probabilistic recursive reasoning. (a) Prior RRMs are deterministic — all runs collapse to an identical trajectory, converging to a single solution. (b) GRAM explores diverse trajectories that reach multiple valid solutions, naturally enabling parallel inference-time scaling.
GRAM architecture. A single stochastic latent transition. After K low-level refinements via fL, the high-level update fH produces a deterministic proposal ut, to which learnable stochastic guidance εt is added: ht = ut + εt. The mean encodes a state-dependent direction; the variance controls the amount of exploration.
GRAM beats every deterministic recursive baseline, even with a single sample. It reaches 86.6% against TRM's 83.6% on Sudoku-Extreme, 50.3% against 44.6% on ARC-AGI-1, and 9.5% against 7.8% on ARC-AGI-2. Sampling 20 trajectories in parallel pushes Sudoku-Extreme to 96.1%.
Stochastic guidance improves reasoning. Looped TF, HRM, and TRM learn from a single fixed path per input. GRAM's stochastic transitions expose it to diverse intermediate states during training, which we credit for its lead even when it draws only one sample at test time.
Parallel sampling provides a new inference-time scaling axis. GRAM supports two complementary axes of inference-time scaling: depth, by varying the number of recursive transitions, and width, by sampling multiple latent reasoning trajectories in parallel. Candidates are aggregated by majority voting, which adds no parameters or training, so the gains are attributable to parallel sampling itself.
(Left) Recursion depth and parallel sampling. All models benefit from longer recursion, while GRAM additionally scales with the number of samples N. (Right) Width beats depth at equal compute. With nearly the same FLOPs (8.48 vs. 8.19 TFLOPs per puzzle), 20 parallel samples at 16 iterations reach 96.1%, while TRM's 320 sequential iterations reach 92.1%. Since the samples run in parallel, GRAM also takes 184 ms per puzzle against 2,755 ms for TRM. Curves show the mean over 3 seeds with one standard deviation.
Deterministic recursion fails on multi-solution tasks. A deterministic model must squeeze several valid answers into one output. On 8×8 N-Queens, deterministic baselines reach at most 78.7% accuracy (GRAM 99.5%) and recover at most 26.7% of valid solutions, while sampling models recover 84.8% or more.
Recursive refinement yields sharper constraint satisfaction. Generative models (AR, MDLM) also cover many solutions, but their samples break constraints more often in most settings. GRAM reaches 99.5% on 8×8 N-Queens and leaves under a third of AR's conflict edges on Graph Coloring, and it still comes out ahead on most tasks when every generative model draws 20 samples.
The gap grows with the number of solutions. Grouping 8×8 N-Queens inputs by their number of valid solutions, deterministic models lose accuracy as solutions increase, whereas GRAM stays flat even with a single sample. More valid solutions mean more conflicting targets, which a deterministic model must collapse into a single output, while GRAM absorbs this ambiguity in its latent trajectory.
Accuracy by number of valid solutions. Single-sample accuracy on 8×8 N-Queens, grouped by the number of valid solutions per input. Deterministic recursive baselines lose accuracy as solutions increase, whereas GRAM stays flat.
We visualize latent trajectories during recursive computation on Sudoku by projecting the high-level state into 2D via PCA. TRM follows a single deterministic path with no mechanism to escape suboptimal regions. GRAM samples diverse trajectories that explore different regions of latent space — while some become trapped in local minima, others successfully navigate toward the global optimum, enabling reliable solution discovery through parallel exploration.
TRM: Single deterministic path.
GRAM (50 samples): Diverse stochastic trajectories.
GRAM extends from conditional reasoning to unconditional generation. By replacing the input with an empty conditioning signal, the same recursive process defines an unconditional generative model p(x).
Sudoku — from empty board to valid solution
MNIST — from noise to recognizable digit
Generative behavior beyond reasoning. GRAM produces valid boards with 99.05% validity using 10.9M parameters and 16 supervision steps, surpassing D3PM baselines that use up to 55.1M parameters and 1000 denoising steps. The generation trajectories below suggest that refining a complete candidate suits hard-constraint generation better than filling cells progressively.
Generation trajectories of D3PM and GRAM on unconditional Sudoku. Each row shows successive states of a single generation from an empty grid, D3PM over denoising steps and GRAM over recursion steps. GRAM proposes a complete board at its first step and corrects it in place, while D3PM fills cells progressively. Incorrect digits are highlighted in red.
| Method | #Params | Steps | Validity (%) |
|---|---|---|---|
| D3PM (Big) | 55.1M | 1000 | 79.18 |
| D3PM (Small) | 15.9M | 1000 | 21.88 |
| GRAM (Ours) | 10.9M | 16 | 99.05 |
Visualization of the generation process and samples. GRAM progressively refines the generated image through recursive latent updates, correcting initial errors.
The deterministic baseline TRM exhibits mode collapse (FID 303.29), whereas GRAM samples diverse digits from its prior, outperforming VAE with only 16 steps and surpassing D3PM on both IS and FID with 256 steps, still far fewer than D3PM's 1,000 steps. Inference-time scaling transfers to generation: increasing recursion at inference improves quality monotonically (IS 1.89 → 2.04, FID 77.79 → 73.34 from 16 to 256 steps), even though training uses only 16 steps. This shows that GRAM can use additional recursive computation at inference to improve generation quality.
Stochastic guidance helps any recursive architecture. SG improves recursion whether it is flat or hierarchical. Adding SG alone lifts the flat Looped TF baseline from 61.7% to 74.0% on Sudoku and from 68.1% to 87.3% on N-Queens (N=20). Adding it on top of deep supervision and hierarchical recursion turns TRM into GRAM, the best model on both tasks (96.1% / 100.0%).
Stochasticity explores, guidance steers. Without stochasticity (guide only), GRAM fails completely, with 0.0% on both tasks, so exploration is what recursion relies on. Without guidance (stochasticity only), Sudoku stays at 96.0% but N-Queens drops to 62.1%. When many valid solutions compete, exploration needs a direction, and guidance provides it.
Randomness alone is not enough. Simply adding noise to TRM, by sampling its outputs or randomizing its initial state, stays below GRAM on both tasks (at most 91.2% on Sudoku and 71.4% on N-Queens). The gains come from learning the stochastic transitions variationally, not from injecting randomness.
@misc{baek2026generativerecursivereasoning,
title={Generative Recursive Reasoning},
author={Junyeob Baek and Mingyu Jo and Minsu Kim and Mengye Ren and Yoshua Bengio and Sungjin Ahn},
year={2026},
eprint={2605.19376},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2605.19376},
}