Spaces:
Running
Running
Reproducing ClawBench V2 scores: corpus and evaluation protocol
#1
by Reacherx - opened
Maintainer note for users of this leaderboard Space. ClawBench V2 contains 130 live-website tasks; the public repository currently exposes V1+V2 task definitions (283 tasks total), while leaderboard rows identify the corpus and harness used. Scores are reported with both request-interception evidence and the agentic judge; please cite the exact corpus, harness, model, judge mode, and commit when adding a run.
Reproduction resources:
- Code and task schemas: https://github.com/TIGER-AI-Lab/ClawBench
- Paper: https://arxiv.org/abs/2604.08523
- Project and run instructions: https://claw-bench.com
- Trace datasets: https://huggingface.co/collections/TIGER-Lab/clawbench
This note is informational and does not alter leaderboard results. New results should follow the repository submission instructions and include enough metadata for an independent rerun.