How you find out whether an agent is any good: held-out splits, trajectory-level scoring, and leaderboards that publish their method.
👋 Open to Work
Yauhen Bichel
YauhenBichel
·
AI & ML interests
Create new models, create harnesses, use projects with AI agentic apps, platform engineering
Recent Activity
new activity about 5 hours ago
Qwen/Qwen2.5-Coder-32B-Instruct:Laptop helper measurement: 1/10 then 9/10 after stripping markdown fences new activity about 5 hours ago
YauhenBichel/python-vibe-0.5b:What these adapters were measured against new activity about 5 hours ago
Qwen/Qwen2.5-Coder-0.5B-Instruct:Laptop helper measurement: style prior, not daily work (0/4 vibe)