--- library_name: transformers tags: - RNA - language-model - UTR - genomics - biology license: gpl-3.0 --- # UTR-LM-MLMSS Minimal HuggingFace port of the **MLM + secondary structure** variant of [UTR-LM](https://github.com/a96123155/UTR-LM) -- an ESM2-style 5' UTR RNA language model pretrained on endogenous sequences from five species and a large synthetic library. ## Architecture | Parameter | Value | |---|---| | Layers | 6 | | Attention heads | 16 | | Embedding dimension | 128 | | FFN hidden dimension | 512 (GELU) | | Vocabulary size | 10 | | Positional encoding | RoPE (base=10000) | | Normalization | LayerNorm | | Architecture | ESM2-style pre-LN Transformer with GELU FFN | | Max sequence length | 1024 tokens (1022 nucleotides + `` / ``) | **Vocabulary:** `` (0), `` (1), `` (2), `A` (3), `G` (4), `C` (5), `T` (6), `` (7), `` (8), `` (9) ## Pretraining - **Objective:** Masked language modeling + per-token secondary structure prediction (3-class: unpaired, stem, loop) - **Data:** Endogenous 5' UTRs from five species (human, mouse, zebrafish, *Drosophila*, yeast) combined with the Cao et al. random 5' UTR synthetic library - **Source checkpoint:** `ESM2SS_FS4.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_lr1e-05_structureweight1.0_MLMLossMin_epoch200.pkl` Only one `ESM2SS` (secondary structure only, no MFE regression) checkpoint was available; no selection decision was required. ## Parity Verification All 7 representation levels (embedding + 6 transformer blocks) were verified to be bit-exact (max absolute difference = 0.00) against the original `ESM2SS_FS4.1_fiveSpeciesCao_6layers_16heads_128embedsize_4096batchToks_lr1e-05_structureweight1.0_MLMLossMin_epoch200.pkl` weights. Verified on GPU with PyTorch 2.7.1 / CUDA 12.9. ## Related Models See the full [UTR-LM collection](https://huggingface.co/collections/Taykhoom/utr-lm-6a173a96ae7c070c3a84ebb4). | Model | Pretraining Objective | Notes | |---|---|---| | [UTR-LM-MLM](https://huggingface.co/Taykhoom/UTR-LM-MLM) | MLM | Base model | | [UTR-LM-MLMSI](https://huggingface.co/Taykhoom/UTR-LM-MLMSI) | MLM + MFE regression | Recommended for TE / EL tasks | | **[UTR-LM-MLMSS](https://huggingface.co/Taykhoom/UTR-LM-MLMSS)** | MLM + secondary structure | This model | | [UTR-LM-MLMSISS](https://huggingface.co/Taykhoom/UTR-LM-MLMSISS) | MLM + MFE + secondary structure | Recommended for MRL tasks | ## Usage ### Embedding generation ```python import torch from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True) model = AutoModel.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True) model.eval() sequences = ["ATGCATGCATGC", "GCTAGCTAGCTAGCTA"] enc = tokenizer(sequences, return_tensors="pt", padding=True) with torch.no_grad(): out = model(**enc) # CLS token embedding (position 0) - recommended for sequence-level tasks cls_emb = out.last_hidden_state[:, 0, :] # (batch, 128) # All-token embeddings token_emb = out.last_hidden_state # (batch, seq_len, 128) # Intermediate layer representations out_all = model(**enc, output_hidden_states=True) layer3_emb = out_all.hidden_states[3] # after layer 3, shape (batch, seq_len, 128) ``` ### MLM logits ```python import torch from transformers import AutoTokenizer, AutoModelForMaskedLM tokenizer = AutoTokenizer.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True) model = AutoModelForMaskedLM.from_pretrained("Taykhoom/UTR-LM-MLMSS", trust_remote_code=True) model.eval() enc = tokenizer(["ATGCATGC"], return_tensors="pt") with torch.no_grad(): logits = model(**enc).logits # (1, seq_len, 10) ``` ### Faster attention backends ```python # SDPA (PyTorch 2.0+) model = AutoModel.from_pretrained( "Taykhoom/UTR-LM-MLMSS", trust_remote_code=True, attn_implementation="sdpa", ) # Flash Attention 2 (requires flash-attn) model = AutoModel.from_pretrained( "Taykhoom/UTR-LM-MLMSS", trust_remote_code=True, attn_implementation="flash_attention_2", dtype=torch.bfloat16, ) ``` ### Fine-tuning The model follows standard HF conventions and can be fine-tuned with any Trainer-compatible setup. For sequence regression tasks, use the CLS token embedding as input to a prediction head (as done in the original UTR-LM paper). ## Implementation Notes The source tokenizer uses the DNA-style `A/G/C/T` alphabet. Convert `U` to `T` before tokenization when supplying RNA-spelled sequences; a literal `U` otherwise maps to ``. Secondary structure was an auxiliary prediction target, not an input channel; this minimal port preserves the backbone and MLM head but omits the auxiliary structure head. The original UTR-LM implementation uses eager scaled dot-product attention. This port additionally supports `attn_implementation="sdpa"` and `attn_implementation="flash_attention_2"`. ## Citation ```bibtex @article{chu2024_utrlm, title = {A 5' {UTR} Language Model for Decoding Untranslated Regions of {mRNA} and Function Predictions}, author = {Chu, Yanyi and Yu, Dan and Li, Yupeng and Huang, Kaixuan and Shen, Yue and Cong, Le and Zhang, Jason and Wang, Mengdi}, journal = {Nature Machine Intelligence}, volume = {6}, number = {4}, pages = {449--460}, year = {2024}, doi = {10.1038/s42256-024-00823-9} } ``` ## Credits Original model and code by Yanyi Chu et al. Source: [UTR-LM GitHub repository](https://github.com/a96123155/UTR-LM). Hugging Face port maintained by Taykhoom Dalal. ## License GPL-3.0, following the original UTR-LM repository.