Back to News
Research2026-04-17
From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space
Source: Arxiv CS.AI
arXiv:2604.14142v1 Announce Type: cross Abstract: While reinforcement learning with verifiable rewards (RLVR) significantly enhances LLM reasoning by optimizing the conditional distribution P(y|x), its potential is fundamentally bounded by the base model's existing output distribution. Optimizing...
arxivpapersrl