Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
E.M. Rabani
PRO
SkyMind
23
14
369
Follow
tegridydev's profile picture
projectlosangeles's profile picture
RiverRider's profile picture
12 followers
·
82 following
radsci
radsci
radsci
radsci
AI & ML interests
Neuromorphic NNs
Recent Activity
liked
a model
about 22 hours ago
Qwen/Qwen3.8-Flash-Next-FP8
reacted
to
Bc-AI
's
post
with 🔥
1 day ago
Building a 6.58B sparse MoE model from scratch on a single GPU. Hey everyone! Today is day 1 of training Smilyai-Lab's new model I call G1-MINI. It's basically the smaller version of our planned model G1 which will be 20B and activate about 2B per token. MINI activates about 1.16B params per token and is currently training right now. If no errors spring up now, I'd say i can launch sometime around September 15th~ish. Thanks to our beta testers: @guardamarcos @ProCreations @juiceb0xc0de @Timmy6767 @Sbui503 @atom77777 @Fishtiks @smartdigitalnetworks @EmetTheGolum @smilyai-large-team @MUK-IS-GOAT @Bc-AI @Banaxi-Tech @vovaRL @Datdanboi25
reacted
to
RiverRider
's
post
with 🔥
1 day ago
Where the Hivemind Comes From: Geometry, Tuning and Format, Separated on Open Weights “First, representations are mutually recoverable. On 12 open-weight models from 8 labs, a ridge map from one model's hidden states to another's retrieves the right held-out item 0.9181 of the time across lab boundaries, against a shuffled floor of 0.00101 and a self-map ceiling of 0.999. Shared corporate lineage is worth only 0.0357 of that.” “Second, base models do not reproduce the reported level. Under the original study's own sampling settings, our base models reach intra-model 0.3644 and inter-model 0.3401 on a floor of 0.0993 that matches theirs, and zero of 720 model-prompt cells clear 0.8. The floors agree while the signal differs by more than a factor of two, so this is not a scale artifact.” “Third, and decisively, we recover their level and isolate its cause. Using six matched base/instruct pairs, holding pretrained weights, prompts, decoding and scorer fixed, instruction tuning alone raises intra-model similarity by 0.0786. The same tuned weights prompted through the model's own chat template raise it by 0.3623, reaching 0.7272, with four of six models exceeding 0.80 and reproducing the band reported for frontier systems from models of 0.6B to 2B. The prompt format does roughly 4.6 times the work of the tuning.” paper attached 🧾 https://huggingface.co/blog/RiverRider/where-the-hivemind-comes-from-geometry-tuning-and
View all activity
Organizations
None yet
models
2
Sort: Recently updated
SkyMind/Agents-A1-Q6_K-GGUF
Text Generation
•
35B
•
Updated
Jul 22
SkyMind/Agents-A1-Q5_K_M-GGUF
Text Generation
•
35B
•
Updated
Jul 17
•
10
datasets
0
None public yet