Mayank Maheshwari
mayankiit04
AI & ML interests
None yet
Recent Activity
View all activity
Organizations
None yet
Quick share: Got 30–40 t/s on the RTX 5060 Ti 16GB
12
#61 opened 20 days ago
by
QilinWan
Small gain with MTP
7
#59 opened 20 days ago
by
OmarColocci
can we have a gguf varity where ngram layer 2 is at q8?
#53 opened 23 days ago
by
mayankiit04
How can I improve the prefilling speed for this model?
14
#49 opened 24 days ago
by
BipedalBit
With Qwen and GLM now in sub 300B local machine level, this is a big model.
👍 1
1
#10 opened 25 days ago
by
mayankiit04
Is the n-gram chunk embedded in the gguf(s), and is it ~51GB independent of quantization?
14
#35 opened 27 days ago
by
dagb
How to keep n-gram table on fast nvme ssd
20
#23 opened 28 days ago
by
mayankiit04
here qwen3.8-flash-next is 125B but unsloth has 180B ... how come?
3
#17 opened 28 days ago
by
mayankiit04
Are you serious, Chatgpt puts this model at estimate 59 score AA above opus 4.8!!
👍 1
#3 opened about 1 month ago
by
mayankiit04
ArtificialAnalysis score of 52 outscore GLM 5.2 , Opus 4.6 and touches Opus 4.7!!!!
🔥 1
10
#143 opened about 1 month ago
by
mayankiit04
Qwen 3.8-27b context usage is about 10x of 3.6-27b
👍 4
8
#45 opened about 1 month ago
by
WhiteDan64
Does this preserve the scientific and reasoning mind of Qwen3.6 models or does that get scarificed for agentic coding
➕ 1
2
#22 opened about 2 months ago
by
mayankiit04
Stable MTP first release!
❤️ 15
11
#6 opened 4 months ago
by
danielhanchen
Is it possible to only download the mtp gguf (<1GB one) to use with existing ggufs?
4
#3 opened 5 months ago
by
CHNtentes
Does this perform in comparision to base b16 quantized model ?
1
#1 opened 5 months ago
by
mayankiit04
Will there be a smaller model like Qwen3.5 122 or Nemotron 3 super
➕ 7
7
#9 opened 5 months ago
by
mayankiit04