Papers
arxiv:2609.03241

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

Published on Sep 3
· Submitted by
Kishan Panaganti
on Sep 8
#2 Paper of the day
Authors:
,
,
,

Abstract

FlowBalance improves reasoning models via verifier-calibrated self-guidance using trajectory-level score reweighting and profile-based trajectory balance.

A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on a narrow solution mode. We introduce FlowBalance, a verifier-grounded self-improvement method that learns a normalized distribution over complete responses. For each on-policy trajectory, a frozen training-time view of the same policy uses privileged context to produce token-level log-probability gains, which are aggregated into a trajectory-level self-guidance score. FlowBalance calibrates this score with the verifier-derived group advantage: guidance is retained on positive-advantage trajectories, reversed on negative-advantage trajectories, and disabled when the rollout group provides no outcome preference. The resulting energy exponentially reweights a reference policy, and profiled trajectory balance fits the normalized target with one log-partition estimate per rollout group. This realizes outcome-calibrated self-guidance via trajectory balance, without a separate token-level imitation loss. Our analysis establishes within-group contrast preservation, a minimum-change reverse-KL characterization, monotonic verifier control of target reward, and an exact correction against false-positive self-guidance on rejected responses. On mathematical reasoning, FlowBalance improves average performance over FlowRL on both Qwen3-4B and Qwen3-8B, while also improving training speed and stability, avoiding direct OPSD's response-length collapse, and exhibiting higher correct-strategy diversity in a controlled AIME24 diagnostic.

Community

FlowBalance is a verifier-grounded self-improvement method for reasoning models. It addresses one known failure mode of OPSD on thinking models: privileged feedback can reinforce locally plausible but verifier-wrong traces. FlowBalance combines sparse outcome rewards with dense self-feedback while flipping misleading teacher support on failures.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.03241
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.03241 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.03241 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.03241 in a Space README.md to link it from this page.

Collections including this paper 2