Image-to-Text
Transformers
Safetensors
English
florence2
image-text-to-text
florence-2
image-captioning
anime
bf16
lora
fine-tuned
custom_code
Instructions to use Mitchins/Florence-Chan-2.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mitchins/Florence-Chan-2.1 with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "image-to-text" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("image-to-text", model="Mitchins/Florence-Chan-2.1", trust_remote_code=True)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Mitchins/Florence-Chan-2.1", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("Mitchins/Florence-Chan-2.1", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Florence-Chan-2.1
- What changed
- Training data
- Package
- Safe comparison set
- dataset/danbooru_images/danbooru_8428374.jpg
- dataset/danbooru_images/danbooru_5341516.jpg
- dataset/danbooru_images/danbooru_5294973.jpg
- dataset/danbooru_images/danbooru_7019581.jpg
- dataset/danbooru_images/danbooru_8654121.jpg
- dataset/danbooru_images/danbooru_5221484.jpg
- dataset/danbooru_images/danbooru_8643115.jpg
- dataset/danbooru_images/danbooru_6621403.jpg
- Notes
- Run with Transformers
- What changed
Florence-Chan-2.1
A Florence-2-large-derived anime caption model, baked into a standalone BF16 safetensors checkpoint. No adapter is required at inference time.
What changed
- Based on the v2 cumulative teacher line and its follow-up SFT continuation.
- Extended with the latest tranche plus an equal-sized v2 retention mix.
- Kept the standard Florence 2 large baseline behavior close while improving the validated follow-up row set.
- The training lineage used GLM-4.6V and Grok-4.3 teacher captions.
Training data
- Broader cumulative teacher lineage:
44,647rows. - Follow-up SFT used for this checkpoint:
980rows. - Follow-up SFT composition:
490latest-tranche rows +490retention rows.
Package
- model:
Florence-Chan-2.1 - parent model:
Florence 2 large - follow-up SFT size:
980 image-caption rows - cumulative teacher lineage size:
44,647 rows - base model:
microsoft/Florence-2-large - adapter source:
/data/anime-captioner/checkpoints/florence2_cumulative_teacher_v2_extension_late_stage_v1/best - merged checkpoint:
/data/anime-captioner/hf.co/florence2-cumulative-teacher-v2p1 - safetensors:
/data/anime-captioner/hf.co/florence2-cumulative-teacher-v2p1/model.safetensors
Safe comparison set
The table below compares the baked model against the standard Florence 2 large baseline and the raw WD14 tags on eight safe, diverse images.
dataset/danbooru_images/danbooru_8428374.jpg
| View | Content |
|---|---|
| WD14 tags | weapon (0.88) long hair (0.88) sword (0.85) animal ears (0.85) holding weapon (0.68) fox ears (0.67) holding (0.67) holding sword (0.65) katana (0.61) skirt (0.59) |
| Florence-2 large | Short: An illustration of a woman with a sword in a city. Detail: The image shows an animated person holding a sword in front of a cityscape, with buildings in the background and a watermark on the image. |
| Florence-Chan-2.1 | Short: Dark-haired horned warrior in dark armor wielding a large sword amid blue energy effects in a ruined city. Detail: A dark-haired figure with horns and a flowing black cloak stands in a dynamic pose, gripping a large sword with both hands, surrounded by swirling blue energy effects and debris in a ruined cityscape with broken buildings and a pink flower-like structure floating in the air. |
dataset/danbooru_images/danbooru_5341516.jpg
| View | Content |
|---|---|
| WD14 tags | mostima (arknights) (0.97) multiple girls (0.91) 2girls (0.91) doughnut (0.90) food (0.88) ? (0.87) tail (0.86) horns (0.86) halo (0.85) hug from behind (0.83) |
| Florence-2 large | Short: A couple of anime girls sitting next to each other. Detail: The image shows two anime girls sitting next to each other on a couch, with one of them holding a donut in her hand. The image is animated, giving it a lively and vibrant feel. |
| Florence-Chan-2.1 | Short: Blue-haired cat-eared girl holds a donut beside a red-haired horned girl in a hoodie. Detail: A blue-haired girl with cat ears and a halo holds a donut in her right hand while resting her left hand on the lap of a pink-haired horned girl in a black hoodie, their arms wrapped around each other in a close embrace. |
dataset/danbooru_images/danbooru_5294973.jpg
| View | Content |
|---|---|
| WD14 tags | guts (berserk) (0.97) armor (0.92) multiple boys (0.92) 2boys (0.92) weapon (0.88) long hair (0.83) sword (0.78) blue skin (0.74) colored skin (0.73) wavy hair (0.70) |
| Florence-2 large | Short: A group of anime characters standing next to each other. Detail: The image shows a group of four people standing next to each other in front of a white background. The person in the front is wearing a white dress and is holding a sword, while the person to the left is wearing an unknown outfit. The image is animated, giving it a dynamic feel. |
| Florence-Chan-2.1 | Short: White-haired woman in silver armor flanked by two dark-haired men in red robes. Detail: A woman with long white curly hair and blue eyes stands in the foreground wearing a white and silver armor with a glowing blue orb at her chest. To her left is a young man with short black hair and a red jacket, and to her right is a man with spiky dark hair wearing a red cape and holding a sword. |
dataset/danbooru_images/danbooru_7019581.jpg
| View | Content |
|---|---|
| WD14 tags | 1girl (0.94) scarf (0.91) solo (0.89) lying (0.85) animal ears (0.85) halo (0.85) one eye closed (0.82) on back (0.82) winter clothes (0.73) smile (0.68) |
| Florence-2 large | Short: A girl in a white coat laying on a pillow. Detail: The image shows a girl in a white coat laying on top of a bed next to a cat, with a pillow beneath her head. She appears to be sleeping peacefully, with her eyes closed and her arms tucked in close to her body. The background is a soft, muted color, giving the image a dreamy, ethereal feel. |
| Florence-Chan-2.1 | Short: Cat-eared girl in a white coat lies on a bed with a fluffy tail. Detail: A cat-eared girl with brown and white hair lies on her back across white sheets, eyes closed and mouth open, wearing a white puffy coat, pink scarf, black gloves, and brown boots. A large fluffy tail rests between her legs, and a pink circular emblem floats above her head. |
dataset/danbooru_images/danbooru_8654121.jpg
| View | Content |
|---|---|
| WD14 tags | 1girl (0.99) japanese clothes (0.96) kimono (0.95) umbrella (0.94) solo (0.89) animal ears (0.88) looking back (0.80) sword (0.79) weapon (0.79) from behind (0.78) |
| Florence-2 large | Short: A woman in a kimono holding an umbrella and a cat. Detail: The image shows a woman in a kimono sitting on a bench, holding an umbrella in one hand and a stick in the other. She is surrounded by trees and a shed in the background, creating a peaceful atmosphere. The image is animated, giving it a dynamic feel. |
| Florence-Chan-2.1 | Short: Cat-eared girl in a blue floral kimono holding a red umbrella. Detail: A young woman with black hair in a high bun and fox ears wears a teal kimono patterned with white and gold floral designs while seated on a wooden surface. She holds a small red object in one hand and rests the other on her lap, with a red umbrella visible behind her and a sword resting on the ground in front of her. |
dataset/danbooru_images/danbooru_5221484.jpg
| View | Content |
|---|---|
| WD14 tags | hina (blue archive) (0.99) 1girl (0.97) hina (swimsuit) (blue archive) (0.96) whistle (0.95) innertube (0.93) horns (0.92) purple eyes (0.85) ahoge (0.85) long hair (0.84) hair ornament (0.82) |
| Florence-2 large | Short: A cartoon of a girl floating in a pool with a duck. Detail: The image shows a cartoon of a girl floating in the water with a crown on her head, wearing a life jacket and holding an object in her hand. The background is white and there is some text at the bottom of the image. |
| Florence-Chan-2.1 | Short: White-haired horned girl in an inflatable ring with a white duck floating in blue water. Detail: A white-haired girl with spiky hair and a red bandage floats in blue water, her face flushed with eyes closed and mouth open, while a small white creature with a pink bow floats beside her and a purple crown floats above her head. |
dataset/danbooru_images/danbooru_8643115.jpg
| View | Content |
|---|---|
| WD14 tags | multiple girls (0.98) 2girls (0.96) kimono (0.96) japanese clothes (0.95) wings (0.93) seiza (0.91) hina (blue archive) (0.91) sash (0.89) halo (0.87) sitting (0.85) |
| Florence-2 large | Short: A couple of girls sitting on top of a tatami mat. Detail: The image shows three anime girls in kimono sitting on the floor in front of a window, with a houseplant on the left side and text on the image. Through the window, we can see trees and snow outside. |
| Florence-Chan-2.1 | Short: Pink-haired girl in a white kimono and white-haired dragon-eared girl in purple kimonos sit on a tatami mat. Detail: A pink-haired girl in a white kimono with a pink bow and floral headpiece sits on the left, holding a pink bowl, while a white-haired dragon-like girl with purple eyes kneels on the right in a purple kimonos with black wings and a white headpiece, both positioned on yellow tatami mats in front of large windows with snow-covered trees visible outside. |
dataset/danbooru_images/danbooru_6621403.jpg
| View | Content |
|---|---|
| WD14 tags | 1girl (0.99) solo (0.97) animal ears (0.95) sword (0.95) weapon (0.94) holding sword (0.93) holding (0.93) holding weapon (0.92) long hair (0.86) bear ears (0.86) |
| Florence-2 large | Short: A girl in a black dress holding a sword in a field. Detail: The image shows a girl in a black dress holding a sword in a field surrounded by plants with flowers, trees, hills, and a starry sky. |
| Florence-Chan-2.1 | Short: Purple-haired girl with bear ears holds a sword in a field at night. Detail: A young woman with long purple hair in twin pigtails tied with purple ribbons and teddy bear ears sits on a grassy field at night, wearing a black kimono-style outfit with purple trim and a plaid skirt, holding a tall silver sword in her right hand with her left arm raised. |
Notes
- Images are included at 512px for table display.
- Baseline is the standard Florence-2-large captioning path.
- WD14 tags are shown as the raw local tagger output for the same images.
- The checkpoint is merged and serialized as safetensors, so there is no adapter dependency.
Run with Transformers
Known-good runtime: transformers==4.46.3.
The standard Florence-2 processor/model API works directly against the baked checkpoint.
from PIL import Image
import torch
from transformers import AutoModelForCausalLM, AutoProcessor
model_id = "/data/anime-captioner/hf.co/florence2-cumulative-teacher-v2p1"
image_path = "/data/anime-captioner/dataset/danbooru_images/danbooru_8428374.jpg"
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
trust_remote_code=True,
).eval()
image = Image.open(image_path).convert("RGB")
for task in ["<CAPTION>", "<DETAILED_CAPTION>"]:
inputs = processor(text=task, images=image, return_tensors="pt")
inputs = {k: v.to(model.device) for k, v in inputs.items()}
with torch.no_grad():
generated_ids = model.generate(
input_ids=inputs["input_ids"],
pixel_values=inputs["pixel_values"],
max_new_tokens=128,
do_sample=False,
num_beams=1,
)
text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(task, text)
- Downloads last month
- 262
Model tree for Mitchins/Florence-Chan-2.1
Base model
microsoft/Florence-2-large






