Overaddicted-500K / README.md
DedeProGames's picture
NanoDex 500k · 1,499,987,968 fineweb-edu tokens · loss 3.5919
715f9f9 verified
|
Raw
History Blame Contribute Delete
2.03 kB
metadata
license: odc-by
datasets:
  - HuggingFaceFW/fineweb-edu
language:
  - en
library_name: transformers
pipeline_tag: text-generation
tags:
  - nanodex
  - tiny-lm
  - pretrained-from-scratch

Overaddicted-500K

A 492,192-parameter decoder-only language model pre-trained from scratch on fineweb-edu, using the NanoDex Trainer Space.

Architecture

A standard LlamaForCausalLM decoder-only transformer — SiLU MLP, RMSNorm, rotary position embeddings, grouped-query attention, tied embeddings, no biases — scaled down in width and depth to fit the parameter budget.

Parameters 492,192
Hidden size 96
Layers 3
Attention heads 6 (KV: 2)
FFN size 256
Context length 512
Vocab 2,048 (custom BPE trained on fineweb-edu)

Training

Tokens seen 1,499,987,968
Steps 11,444
Tokens / step 131,072
Optimizer AdamW(0.9, 0.95) wd=0.1 clip=1.0
LR schedule warmup 2% + cosine to 10% (peak 4e-03)
Final loss 3.5919 (ppl 36.3)
Wall time 33.9 min
Trained by @DedeProGames

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("DedeProGames/Overaddicted-500K")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/Overaddicted-500K")

ids = tok("The mitochondria is", return_tensors="pt").input_ids
print(tok.decode(model.generate(ids, max_new_tokens=60, do_sample=True,
                                temperature=0.8, top_k=50)[0]))

Caveats

This is a nano-scale research artifact. At this parameter count and token budget the model learns word shapes, common collocations and a little syntax — it is not a useful assistant and its output is not factual. It exists to make "pre-train a transformer from scratch" something you can actually watch happen.