KV Cache Steering for Inducing Reasoning in Small Language Models Paper โข 2507.08799 โข Published Jul 11, 2025 โข 40
view article Article I trained a Language Model to schedule events with GRPO! anakin87 โข Apr 29, 2025 โข 95
view article Article A failed experiment: Infini-Attention, and why we should keep trying? +1 neuralink, lvwerra, thomwolf โข Aug 14, 2024 โข 76