Ming
DarrenChen
AI & ML interests
None yet
Organizations
None yet
How do I use this embedding model?
#1 opened 6 months ago
by
DarrenChen
Slow inference speed
👍 1
#25 opened 6 months ago
by
DarrenChen
Loading weigts error when running MiniMax-2.1 with sglang using pipeline parallelism
4
#18 opened 6 months ago
by
tuo02
GroupQueryAttention operation is not supported
1
#1 opened 10 months ago
by
DarrenChen
ONNX on mobile
#1 opened 10 months ago
by
DarrenChen
Are there any other frameworks tested besides transformers that can be deployed?
3
#5 opened 12 months ago
by
DarrenChen
How did you do it?May I ask what quantization method is being used?
#1 opened about 1 year ago
by
DarrenChen
How was this made? Quant configuration? Have you deployed this with SGLANG or vllm ?
1
#1 opened about 1 year ago
by
chriswritescode