AI & ML interests
RL Environments at Scale
Recent Activity
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 16 -
HuggingEnvs/watercolour-reference-pool
Viewer โข Updated โข 178 โข 45 -
HuggingEnvs/watercolour-rollouts-hps-only
Viewer โข Updated โข 470 โข 40 -
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
-
HuggingEnvs/watercolour-grpo-hps-only
Reinforcement Learning โข Updated โข 16 -
HuggingEnvs/watercolour-reference-pool
Viewer โข Updated โข 178 โข 45 -
HuggingEnvs/watercolour-rollouts-hps-only
Viewer โข Updated โข 470 โข 40 -
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning โข Updated
Deterministic data-analysis agent tasks from the jupyter-agent dataset โ verified answers, no LLM judge. Harbor env suites, plain dataset & SFT.
A curated collection of articles, guides, tutorials, slides, and resources for learning how to build, train, and evaluate RL environments for Agents