--- license: apache-2.0 language: - en library_name: transformers pipeline_tag: text-generation tags: - custom-architecture - rope - gqa - swiglu - rmsnorm - safetensors --- # FastLLM (150M) — Modern Causal Language Model **FastLLM** is a ~150M parameter, decoder-only causal language model built completely from scratch in PyTorch and fully integrated with Hugging Face `transformers`. It incorporates state-of-the-art LLM architectural choices—**Grouped-Query Attention (GQA)**, **SwiGLU MLPs**, **RMSNorm**, and **Rotary Position Embeddings (RoPE)**—and natively saves weights in the zero-copy **Safetensors** format. --- ## Model Details * **Developed by:** devoppro * **Model Type:** Decoder-only Causal Language Model * **Architecture:** Custom Transformer (`ModernLLMForCausalLM`) * **Parameter Count:** ~150,000,000 (150M) * **Tokenizer:** Qwen 2.5 BPE Vocabulary (`vocab_size`: 151,936) * **Precision:** Mixed Precision (`FP16`) * **Storage Format:** `.safetensors` * **Repository:** `devoppro/FastLLM` --- ## Architectural Specifications | Parameter | Configuration | | :--- | :--- | | **Hidden Size ($d_{\text{model}}$)** | 768 | | **Intermediate Size (SwiGLU)** | 2048 | | **Hidden Layers** | 12 | | **Query Heads** | 12 | | **Key/Value Heads (GQA)** | 4 (3:1 Query-to-KV ratio) | | **Max Context Length** | 2048 tokens | | **Normalization** | RMSNorm ($\epsilon = 10^{-6}$) | | **Positional Embedding** | Rotary Embeddings (RoPE, $\theta = 1000000.0$) | --- ## Training Data Mixture The model was pre-trained using dynamic stream interleaving across four high-quality datasets: