AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

LLaDA-PRM: A Bidirectional Step-Level Reasoning Evaluator

arXiv · AI, language, vision and robotics · article · Sep 19, 2026 · UTC

Step-level reasoning evaluators are commonly based on autoregressive language models, whose causal attention restricts each step representation to the problem, previous steps, and the current step. Yet, when the complete solution is available, the validity of an earlier step may become clearer only through its downstream consequences. We validate this hypothesis through a controlled 54-run comparison of causal and bidirectional LLaDA evaluators at 1B--3B scale, changing only the self-attention mask, and find bidirectional attention yields consistent improvements. Building on this finding, we i

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.