A newer version of the Gradio SDK is available: 6.28.0
metadata
title: Kimi K3 AirLLM Demo
emoji: 🤖
colorFrom: purple
colorTo: blue
sdk: gradio
sdk_version: 5.38.0
python_version: '3.10'
app_file: app.py
pinned: false
Kimi K3 + AirLLM
A Gradio demonstration of devoppro/Kimi-K3 using AirLLM for memory-efficient inference.
The model cache is stored on the attached persistent /data volume.
Model
devoppro/Kimi-K3
Backend
AirLLM + Transformers
Storage
The Space expects a persistent volume mounted at:
/data
The model can require a very large amount of persistent storage.
First startup can take a long time because the Kimi K3 checkpoint consists of many large safetensor shards.