Paint with Code
An RL environment where the agent paints by writing p5.brush sketches, rewarded by an aesthetic preference model looking at the render.
Reinforcement Learning • Updated • 16Note GRPO adapter for Qwen3.5-35B-A3B. Reward is 0.90 HPSv3 with the pairwise judge off.
HuggingEnvs/watercolour-reference-pool
Viewer • Updated • 178 • 45Note The 178 paintings that define the reward, with the sketch that produced each one.
HuggingEnvs/watercolour-rollouts-hps-only
Viewer • Updated • 470 • 40Note Every rollout of the HPS-only run: 470 paintings, sketches and rewards, by step.
HuggingEnvs/watercolour-grpo-judge-led
Reinforcement Learning • UpdatedNote GRPO adapter for Qwen3.5-35B-A3B. The original mix: pairwise judge at 0.60, HPSv3 at 0.30.
HuggingEnvs/watercolour-grpo-hps-led
Reinforcement Learning • UpdatedNote GRPO adapter for Qwen3.5-35B-A3B. The middle mix: HPSv3 at 0.60, pairwise judge at 0.30.
HuggingEnvs/watercolour-rollouts-judge-led
Viewer • Updated • 861Note Every rollout of the judge-led run: 861 paintings, sketches and rewards, by step.
HuggingEnvs/watercolour-rollouts-hps-led
Viewer • Updated • 872Note Every rollout of the hps-led run: 872 paintings, sketches and rewards, by step.
Watercolour Environment Server
🎨Note The OpenEnv environment: renders each sketch and computes the reward. Duplicate it to train.
Watercolour HPSv3
🎨Note The HPSv3 scorer behind the quality term. Duplicate it on a100-large for a run.
Watercolour Trackio Judge Led
🎯Show interactive tracking visualizations
Note Live training curves of the judge-led run, per step and per metric.
Watercolour Trackio Hps Led
🎯Display track information and visualizations
Note Live training curves of the hps-led run, per step and per metric.
Watercolour Trackio Hps Only
🎯Show interactive tracking visualizations
Note Live training curves of the hps-only run, per step and per metric.