Well, I'd say Openai did malicious marketing stunt to get new customers.
🔄 In a Training Loop
olli
bolenath
AI & ML interests
None yet
Recent Activity
commentedon an article 2 days ago
Security incident disclosure — July 2026 liked a model 26 days ago
BugTraceAI/BugTraceAI-CORE-Ultra-27B-Q6 liked a model 2 months ago
OBLITERATUS/Qwen3.6-27B-OBLITERATEDOrganizations
commented on Security incident disclosure — July 2026 2 days ago
Bigger models.
1
#8 opened 3 months ago
by
bolenath
reacted to mayafree's post with 🔥 4 months ago
Post
5967
Leaderboard of Leaderboards — A Real-Time Meta-Ranking of AI Benchmarks
MAYA-AI/all-leaderboard
Hundreds of AI leaderboards exist on HuggingFace. Knowing which ones the community actually trusts has never been easy — until now.
Leaderboard of Leaderboards (LoL) ranks the leaderboards themselves, using live HuggingFace trending scores and cumulative likes as the signal. No editorial curation. No manual selection. Just what the global AI research community is actually visiting and endorsing, surfaced in real time.
Sort by trending to see what is capturing attention right now, or by likes to see what has built lasting credibility over time. Nine domain filters let you zero in on what matters most to your work, and every entry shows both its rank within this collection and its real-time global rank across all HuggingFace Spaces.
The collection spans well-established standards like Open LLM Leaderboard, Chatbot Arena, MTEB, and BigCodeBench alongside frameworks worth watching. FINAL Bench targets AGI-level evaluation across 100 tasks in 15 domains and recently reached the global top 5 in HuggingFace dataset rankings. Smol AI WorldCup runs tournament-format competitions for sub-8B models scored via FINAL Bench criteria. ALL Bench aggregates results across frameworks into a unified ranking that resists the overfitting risks of any single standard.
The deeper purpose is not convenience. It is transparency. How we measure AI matters as much as the AI we measure.
MAYA-AI/all-leaderboard
Hundreds of AI leaderboards exist on HuggingFace. Knowing which ones the community actually trusts has never been easy — until now.
Leaderboard of Leaderboards (LoL) ranks the leaderboards themselves, using live HuggingFace trending scores and cumulative likes as the signal. No editorial curation. No manual selection. Just what the global AI research community is actually visiting and endorsing, surfaced in real time.
Sort by trending to see what is capturing attention right now, or by likes to see what has built lasting credibility over time. Nine domain filters let you zero in on what matters most to your work, and every entry shows both its rank within this collection and its real-time global rank across all HuggingFace Spaces.
The collection spans well-established standards like Open LLM Leaderboard, Chatbot Arena, MTEB, and BigCodeBench alongside frameworks worth watching. FINAL Bench targets AGI-level evaluation across 100 tasks in 15 domains and recently reached the global top 5 in HuggingFace dataset rankings. Smol AI WorldCup runs tournament-format competitions for sub-8B models scored via FINAL Bench criteria. ALL Bench aggregates results across frameworks into a unified ranking that resists the overfitting risks of any single standard.
The deeper purpose is not convenience. It is transparency. How we measure AI matters as much as the AI we measure.
Qwen/Qwen3-Coder-480B-A35B-Instruct
Text Generation • 480B • Updated • 85k • • 1.36k
Chain-GPT/Solidity-LLM
Text Generation • 3B • Updated • 29 • 1.58k
moonshotai/Kimi-K2-Instruct
Text Generation • 1T • Updated • 202k • • 2.37k
tencent/Hunyuan-A13B-Instruct
Text Generation • 80B • Updated • 55.7k • 796
upvoted a collection about 1 year ago