llms Collection by Xodan69 Aug 12, 2024 - meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
Papers - mPLUG Collection by matlok Aug 12, 2024 - mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33 mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding Paper • 2403.12895 • Published Mar 19, 2024 • 32
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding Paper • 2403.12895 • Published Mar 19, 2024 • 32
Papers - Image - Multi-Image Collection by matlok Aug 19, 2024 - mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33 xGen-MM (BLIP-3): A Family of Open Large Multimodal Models Paper • 2408.08872 • Published Aug 16, 2024 • 101
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models Paper • 2408.08872 • Published Aug 16, 2024 • 101
Flux Collection by Moehawk Oct 19, 2024 - black-forest-labs/FLUX.1-dev Text-to-Image • 12B • Updated Jun 27, 2025 • 668k • • 14.3k pyannote/speaker-diarization Automatic Speech Recognition • Updated May 10, 2024 • 339k • 1.31k
pyannote/speaker-diarization Automatic Speech Recognition • Updated May 10, 2024 • 339k • 1.31k
Text Collection by AlexanderGreat Aug 12, 2024 - meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
VIDEO CON IMAGEN Collection by AndreBoi Aug 12, 2024 - Running on Zero Agents 3.78k Live Portrait 🤪 3.78k Apply the motion of a video on a portrait
Papers - Benchmark - Distractions Collection by matlok Aug 12, 2024 - mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
To be Used Soon Collection by Simbanator Oct 4, 2024 - black-forest-labs/FLUX.1-dev Text-to-Image • 12B • Updated Jun 27, 2025 • 668k • • 14.3k alibaba-pai/CogVideoX-Fun-5b-InP Image-to-Video • Updated Dec 11, 2025 • 46 • 27 timbrooks/instruct-pix2pix Image-to-Image • 0.9B • Updated Jul 5, 2023 • 25.9k • 1.18k
CodeMix Collection by Kartikeya Aug 12, 2024 - cjvt/roberta-en-hi-codemixed Fill-Mask • Updated Jan 8, 2023 • 9
llms Collection by Xodan69 Aug 12, 2024 - meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
VIDEO CON IMAGEN Collection by AndreBoi Aug 12, 2024 - Running on Zero Agents 3.78k Live Portrait 🤪 3.78k Apply the motion of a video on a portrait
Papers - mPLUG Collection by matlok Aug 12, 2024 - mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33 mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding Paper • 2403.12895 • Published Mar 19, 2024 • 32
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding Paper • 2403.12895 • Published Mar 19, 2024 • 32
Papers - Benchmark - Distractions Collection by matlok Aug 12, 2024 - mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
Papers - Image - Multi-Image Collection by matlok Aug 19, 2024 - mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33 xGen-MM (BLIP-3): A Family of Open Large Multimodal Models Paper • 2408.08872 • Published Aug 16, 2024 • 101
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Paper • 2408.04840 • Published Aug 9, 2024 • 33
xGen-MM (BLIP-3): A Family of Open Large Multimodal Models Paper • 2408.08872 • Published Aug 16, 2024 • 101
To be Used Soon Collection by Simbanator Oct 4, 2024 - black-forest-labs/FLUX.1-dev Text-to-Image • 12B • Updated Jun 27, 2025 • 668k • • 14.3k alibaba-pai/CogVideoX-Fun-5b-InP Image-to-Video • Updated Dec 11, 2025 • 46 • 27 timbrooks/instruct-pix2pix Image-to-Image • 0.9B • Updated Jul 5, 2023 • 25.9k • 1.18k
Flux Collection by Moehawk Oct 19, 2024 - black-forest-labs/FLUX.1-dev Text-to-Image • 12B • Updated Jun 27, 2025 • 668k • • 14.3k pyannote/speaker-diarization Automatic Speech Recognition • Updated May 10, 2024 • 339k • 1.31k
pyannote/speaker-diarization Automatic Speech Recognition • Updated May 10, 2024 • 339k • 1.31k
Text Collection by AlexanderGreat Aug 12, 2024 - meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
meta-llama/Llama-3.1-8B-Instruct Text Generation • 8B • Updated Sep 25, 2024 • 6.42M • • 6.67k
CodeMix Collection by Kartikeya Aug 12, 2024 - cjvt/roberta-en-hi-codemixed Fill-Mask • Updated Jan 8, 2023 • 9