File size: 1,979 Bytes
a90f6d5
fdb7961
ba01b1a
 
 
 
 
 
 
 
 
 
 
67349c2
 
 
 
 
 
 
 
ba01b1a
 
67349c2
a90f6d5
9574bbb
a90f6d5
67349c2
a90f6d5
67349c2
a90f6d5
67349c2
 
 
 
 
 
a90f6d5
67349c2
a90f6d5
67349c2
 
 
 
 
 
 
 
9574bbb
67349c2
fdb7961
67349c2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9574bbb
fdb7961
9574bbb
67349c2
 
 
 
 
a90f6d5
 
67349c2
a90f6d5
67349c2
 
 
 
 
 
 
a90f6d5
67349c2
a90f6d5
fdb7961
a90f6d5
67349c2
a90f6d5
67349c2
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
language:
- my
- en
license: apache-2.0
library_name: transformers
pipeline_tag: text-generation
tags:
- myanmar
- burmese
- llm
- chat
- instruction-following
- conversational
- autoregressive
base_model: MiniMaxAI/MiniMax-M2.7
datasets:
- amkyawdev/myanmar-v3-clean
- amkyawdev/burme-coder-max
- amkyawdev/mm-llm-coder-agent-dataset
- saillab/alpaca-myanmar_burmese-cleaned
---

# πŸ‰ Myanmar Ghost

**Advanced Myanmar Language Model (LLM)**

Fine-tuned on MiniMax-M2.7 with QLoRA for Myanmar language understanding.

## πŸ’¬ Features

- πŸ—£οΈ **Myanmar Chat** - Natural conversation in Burmese
- πŸ“ **Instruction Following** - Follow complex Myanmar instructions  
- πŸ’» **Code Generation** - Write Myanmar code and documentation
- 🌐 **Translation** - Myanmar ↔ English
- πŸ“– **Summarization** - Summarize Myanmar text
- ❓ **QA** - Answer questions in Myanmar

## πŸ“Š Training Data

| Dataset | Samples |
|---------|---------|
| myanmar-v3-clean | 877,706 |
| burme-coder-max | 1,000,000 |
| mm-llm-coder-agent | 4,000,020 |
| alpaca-myanmar | 41,601 |

**Total: ~6M instruction samples**

## πŸš€ Quick Start

```python
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

# Load model
model_name = "amkyawdev/myanmar-ghost"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    load_in_4bit=True,
    device_map="auto"
)

# Generate
prompt = """### Instruction:
မြန်မာစာမေးပွဲထကြောင်း ရှင်းပါ

### Response:
"""

inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256, temperature=0.7)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```

## πŸ“‹ Requirements

```
torch>=2.0.0
transformers>=4.40.0
bitsandbytes>=0.40.0
peft>=0.4.0
accelerate>=0.20.0
```

## πŸ“œ License

Apache 2.0

## πŸ‘€ Author

**Aung Myo Kyaw (amkyawdev)**