Sleeping
Agents
Wechat Style Sft
🌍
QLoRA fine-tuning demo on WeChat essays for style adaptation
None defined yet.
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space