RL-Forgetting-Experiments-3/qwen2.5-3b-kk-sft-ordered-lr1e5-step3175 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-kk-sft-ordered-lr1e5-step3175 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-kk-sft-shuffled-lr1e5-step3175 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-kk-sft-shuffled-lr1e5-step3175 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-math-sft-shuffled-lr1e5-step1072 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-math-sft-ordered-lr1e5-step1072 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-math-sft-shuffled-lr1e5-step1072 Text Generation • 3B • Updated about 10 hours ago
RL-Forgetting-Experiments-3/qwen2.5-3b-math-sft-ordered-lr1e5-step1072 Text Generation • 3B • Updated about 10 hours ago
Rethinking Diverse Human Preference Learning through Principal Component Analysis Paper • 2502.13131 • Published Feb 18, 2025 • 37
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning Paper • 2505.24846 • Published May 30, 2025 • 15
Improving Data Efficiency for LLM Reinforcement Fine-tuning Through Difficulty-targeted Online Data Selection and Rollout Replay Paper • 2506.05316 • Published Jun 5, 2025 • 1