This is an FP8 / INT8 quantized version of FLUX.1-schnell.

  • Optimized for efficient inference with reduced memory footprint. Same-seed LPIPS vs the bf16 model (lower is better): 0.151 INT8, 0.157 FP8.

Samples

Prompt: "cute sloth typing on a computer"

INT8 INT8
FP8 FP8

FLUX.1 [schnell] Grid

FLUX.1 [schnell] is a 12 billion parameter rectified flow transformer capable of generating images from text descriptions. For more information, please read our blog post.

Key Features

  1. Cutting-edge output quality and competitive prompt following, matching the performance of closed source alternatives.
  2. Trained using latent adversarial diffusion distillation, FLUX.1 [schnell] can generate high-quality images in only 1 to 4 steps.
  3. Released under the apache-2.0 licence, the model can be used for personal, scientific, and commercial purposes.

Usage

We provide a reference implementation of FLUX.1 [schnell], as well as sampling code, in a dedicated github repository. Developers and creatives looking to build on top of FLUX.1 [schnell] are encouraged to use this as a starting point.

API Endpoints

The FLUX.1 models are also available via API from the following sources

ComfyUI

FLUX.1 [schnell] is also available in Comfy UI for local inference with a node-based workflow.

Diffusers

To use FLUX.1 [schnell] with the 🧨 diffusers python library, first install or upgrade diffusers

pip install -U diffusers

Then you can use FluxPipeline to run the model

import torch
from diffusers import FluxPipeline

pipe = FluxPipeline.from_pretrained("black-forest-labs/FLUX.1-schnell", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload() #save some VRAM by offloading the model to CPU. Remove this if you have enough GPU power

prompt = "A cat holding a sign that says hello world"
image = pipe(
    prompt,
    guidance_scale=0.0,
    num_inference_steps=4,
    max_sequence_length=256,
    generator=torch.Generator("cpu").manual_seed(0)
).images[0]
image.save("flux-schnell.png")

To learn more check out the diffusers documentation


Limitations

  • This model is not intended or able to provide factual information.
  • As a statistical model this checkpoint might amplify existing societal biases.
  • The model may fail to generate output that matches the prompts.
  • Prompt following is heavily influenced by the prompting-style.

Out-of-Scope Use

The model and its derivatives may not be used

  • In any way that violates any applicable national, federal, state, local or international law or regulation.
  • For the purpose of exploiting, harming or attempting to exploit or harm minors in any way; including but not limited to the solicitation, creation, acquisition, or dissemination of child exploitative content.
  • To generate or disseminate verifiably false information and/or content with the purpose of harming others.
  • To generate or disseminate personal identifiable information that can be used to harm an individual.
  • To harass, abuse, threaten, stalk, or bully individuals or groups of individuals.
  • To create non-consensual nudity or illegal pornographic content.
  • For fully automated decision making that adversely impacts an individual's legal rights or otherwise creates or modifies a binding, enforceable obligation.
  • Generating or facilitating large-scale disinformation campaigns.

Quantized transformer checkpoints (this repo)

This repo adds pre-quantized diffusion transformer checkpoints for black-forest-labs/FLUX.1-schnell, built with torchao dynamic activation quantization from the dense bf16 transformer. The official model card above is unchanged from the source repo.

Files:

  • FLUX.1-schnell-INT8.pt (15.2 GB)
  • FLUX.1-schnell-FP8.pt (11.9 GB)

Details:

  • int8: Int8DynamicActivationInt8WeightConfig (per-token activation, per-channel weight, torch._int_mm).
  • fp8: Float8DynamicActivationFloat8WeightConfig with PerRow granularity (e4m3, torch._scaled_mm). The loader must floor the dynamic activation scale (activation_value_lb=1e-12 on torchao 0.13+) so all-zero activation token rows cannot produce a zero scale.
  • Loading a checkpoint is bit-identical to quantizing the dense bf16 transformer on the fly; the checkpoint skips the dense load and quantize step.
  • Validated against same-seed dense bf16 renders (SSIM, LPIPS-vgg, CLIP delta, non-finite and black-frame checks) on torch 2.12.1 and torchao 0.17.

Samples

Prompt: "cute sloth typing on a computer" (1024x1024, family default steps/guidance, seeds 0-2).

int8

fp8

Pre-cast fp8 text encoder (this repo)

FLUX.1-schnell-text_encoder_2-FP8.pt (5.9 GB) is the pipeline's text_encoder_2 (T5EncoderModel, T5-XXL, the text_encoder_2 subfolder of black-forest-labs/FLUX.1-schnell) with the layerwise fp8 storage cast Unsloth Studio applies at load time, saved pre-cast:

  • Loading it is bit-identical to downloading the dense encoder and casting on load (verified tensor for tensor: 220 tensors, 144 cast to fp8 storage; T5's fp32-kept wo projections stay dense per the runtime cast).
  • Cuts the T5 download from 9.5 GB to 5.9 GB. The T5 shards are byte-identical across FLUX.1-schnell, FLUX.1-dev and FLUX.1-Krea-dev (verified sha256), so this one artifact serves all three bases. The small CLIP-L text encoder stays dense.
  • Plain-tensor state dict: loads with torch.load(weights_only=True). Metadata records scheme fp8, component text_encoder_2, base black-forest-labs/FLUX.1-schnell.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Examples
Examples
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for unsloth/FLUX.1-schnell-FP8

Finetuned
(70)
this model