HuggingFaceFW/fineweb
Viewer • Updated • 52.5B • 419k • 3.28k
Note: AI was used in the creation of this project. But that's why you're here, isn't it?
This is the base model for VegaLM1-42M, an experimental SLM trained on various corpi... corpuses... datasets. It's not Fable 6, but it's good enough, ok?
The first 262M tokens the model saw came from a 90M selection of fineweb-edu The next 688M was from a 50/25/15/10 split of fineweb-edu, fineweb, wikipedia, and a replay respectively, totaling around 420M tokens The final 2B tokens came from a 60/25/15 split of stack-v3-train, dclm-baseline-1.0, and smollm-corpus, trained for 1 epoch because I'm lazy
Final version gets 944 elo on Bananamind Base Bench 1.1.
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| arc_challenge | 1 | none | 0 | acc | ↑ | 0.1800 | ± | 0.0112 |
| none | 0 | acc_norm | ↑ | 0.2125 | ± | 0.0120 | ||
| arc_easy | 1 | none | 0 | acc | ↑ | 0.3763 | ± | 0.0099 |
| none | 0 | acc_norm | ↑ | 0.3523 | ± | 0.0098 | ||
| hellaswag | 1 | none | 0 | acc | ↑ | 0.2629 | ± | 0.0044 |
| none | 0 | acc_norm | ↑ | 0.2686 | ± | 0.0044 | ||
| piqa | 1 | none | 0 | acc | ↑ | 0.5702 | ± | 0.0116 |
| none | 0 | acc_norm | ↑ | 0.5762 | ± | 0.0115 |
Arithmark 3: