custom
code
sovereign-compute
File size: 4,031 Bytes
2fa0e9a
 
 
 
 
 
 
 
 
ce25ab4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
---
license: other
license_name: sovereign-source-license-v2
library_name: custom
tags:
- code
- sovereign-compute
---

# Assembly Bite β€” ML Macromodel Corpus

[![License: BSL-1.1](https://img.shields.io/badge/license-BSL--1.1-orange?style=flat-square)](LICENSE)
[![License: AGPL-3.0](https://img.shields.io/badge/license-AGPL--3.0-blue?style=flat-square)](LICENSE)
[![Patent Pending](https://img.shields.io/badge/patent-pending-red?style=flat-square)]()
[![sm_89](https://img.shields.io/badge/GPU-sm__89_Hopper-76b900?style=flat-square)]()
[![PTX/SASS](https://img.shields.io/badge/kernels-PTX%2FSASS-yellow?style=flat-square)]()
[![ΞΈ](https://img.shields.io/badge/ΞΈ-89%2F2462-gold?style=flat-square)]()

**Author:** Ahmad Ali Parr  
**Trust:** Bel Esprit D'Accord Irrevocable Trust Β· EIN 42-697643

> Low-level pseudo-assembly language for representing ML models at the instruction level. Token matchers, tree matchers, full transformer macromodels with multiplicity β€” plus production SASS/PTX kernels for sm_89 (RTX 4090).

---

## What Is Assembly Bite

Assembly Bite is Ahmad's custom pseudo-assembly language for expressing ML model computations at the byte-code level. Four instruction types:

- `.DATA` β€” memory layout declarations (`.word`, `.repl`, `.float`)
- `.CODE` β€” instruction stream (`LOAD`, `STORE`, `CALL`, `MATMUL`, `ADD`, `SUB`, `CMP`, `JEQ`, etc.)
- Subroutine calls for primitive operations (`MATMUL`, `SOFTMAX`, `RELU`, `LAYER_NORM`, `ADD_BIAS`)
- No invented syntax β€” grounded in standard CS algorithms

---

## Contents

```
examples/
  token-matcher/token_matcher.asm     β€” Literal token ID sequence matcher (sliding window)
  tree-matcher/tree_matcher.asm       β€” DFS path matcher on [id, child, sibling] trees
  transformer/transformer_macromodel.asm β€” Full transformer: 24L Γ— 12H Γ— 768D Γ— M=4

python/
  deberta_encoder.py                  β€” DeBERTa-v3 encoder wrapper + instruction token
  gguf_dag_pipeline.py                β€” GGUF load + networkx DAG pipeline

sass/
  flash_attention.ptx                 β€” Flash attention paged (sm_89 Hopper)
  dequant_q4k.ptx                     β€” GGUF Q4_K dequant PTX (sm_89)
  dequant_q4k.sass                    β€” GGUF Q4_K dequant SASS
  cuda_kernels.c                      β€” Host launch wrappers
  mamba_bind.h                        β€” Mamba SSM + GGUF binding header
  qemu_arm64_holyc.HC                 β€” HolyC QEMU ARM64 integration
```

---

## Transformer Macromodel β€” Multiplicity Architecture

The key insight in `transformer_macromodel.asm`: each neuron block has **M copies** (default M=4). The outer loop is `LΓ—AΓ—M` β€” layer Γ— head Γ— multiplicity. Each copy computes Q/K/V independently. Results are summed across copies before the next layer.

```
Parameters: L=24, A=12, D=768, H=3072, M=4
Weight layout: W_Q[L][A][M][D][D] = 24Γ—12Γ—4Γ—768Γ—768
Per copy: full attention + FFN + residual + layer norm
Aggregate: LAYER_ACC += LAYER_OUT for each M copy
```

This is distinct from standard multi-head attention β€” it's multiplicity within each head, not across heads.

---

## SASS/PTX Kernels β€” sm_89 (RTX 4090)

### flash_attention.ptx
Full paged flash attention with TMA async copy and WMMA tensor core tiles.

### dequant_q4k.ptx / dequant_q4k.sass  
GGUF Q4_K block dequantization. 32-value blocks β†’ FP16. Each thread processes 8 values (4 packed bytes). 36 bytes per block layout: `[32 bytes packed][2 bytes scale][2 bytes min]`.

```
Grid: ceil(num_blocks / 256) blocks
Block: 256 threads
Dequant: q_val * scale + min β†’ FP16
```

---

## Build

```bash
# PTX β†’ Cubin (requires CUDA 12.x + SM89 GPU)
ptxas -arch=sm_89 sass/dequant_q4k.ptx -o dequant_q4k.cubin
ptxas -arch=sm_89 sass/flash_attention.ptx -o flash_attention.cubin

# Host wrappers
nvcc -arch=sm_89 sass/cuda_kernels.c -o kernels.so

# Python
pip install torch transformers llama_cpp_python networkx
python python/deberta_encoder.py
```

---

Β© 2026 Bel Esprit D'Accord Irrevocable Trust Β· Patent Pending Β· ΞΈ = 89/2462