Qwen-Image-2.1-Turbo-Heretic-mlx-bf16

English | 中文说明


中文说明

本仓库提供专为 Apple Silicon (Mac) 优化的 mflux MLX BF16 原生格式权重,基于阿里巴巴官方开源的 Qwen/Qwen-Image-2.1-Turbo 加速模型,整合了无审核的 Heretic 文本编码器与原版高精度 FP32 VAE。

支持通过 mflux 在 Apple Silicon Mac 上直接运行 8 步极速文生图、单图精准编辑以及多图参考换装。


核心亮点

  1. **官方 8 步蒸馏 DiT (原生 BF16)**:直接采用官方最新微调的 Turbo Transformer 权重,无需额外挂载加速 LoRA,8 步即可达到高质感成像。
  2. 无审核 Heretic 文本编码器:使用无审核的 Heretic text_encoder,彻底解除安全词拦截与指令拒答,保持完整提示词遵循与图像编辑能力。
  3. 原生 qwen_turbo 采样调度器:完全对齐官方固化的 8 步 sample_sigmas 采样点,消除发丝泛白与噪点问题。
  4. 全功能支持:完整支持文生图 (Text-to-Image)、单图编辑 (Image Editing)、多参考图组合换装 (Multi-Reference Swap) 及额外效果 LoRA 串联。

调度器关键说明 (Scheduler Notice)

重要:运行官方 Turbo 模型时,务必指定 --scheduler qwen_turbo。

官方 Qwen-Image-2.1-Turbo 模型在 model_index.json 中固化了一组专为 8 步蒸馏训练定制的非线性采样节点:

sample_sigmas = [1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568, 0.0]
  • 为什么不能用通用 linear 调度器?
    通用的线性调度器会对步数重新均匀插值,并被基础模型的时间位移公式(time shift)改写。这会导致中间步的高频去噪轨迹发生偏移,实测表现为:发丝逆光边缘严重泛白起噪、眼神光涣散、皮肤微纹理干瘪塑料化。
  • qwen_turbo 调度器的作用:
    在 mflux 内部直接加载这组固定的 sample_sigmas,逐点注入降噪链。逆光发丝微高光、立体骨相与真实肤质即可完全还原官方 Turbo 模型的原生水准。

本地实测性能与画质对比 (M3 Max 128GB)

在 Apple M3 Max 128GB 设备上,统一固定随机种子实测数据如下:

测试场景与分辨率 官方 Turbo (8步 · bf16) Viggle r256 (6步 LoRA) Viggle r128 (6步 LoRA) 画面综合评价
写实肖像发丝光影 (832 × 1248) 54.27 秒 44.92 秒 43.35 秒 官方 Turbo 发丝逆光层次最细腻,立体骨相自然;r256 胶片感强;r128 最快
密集静物与黑板文字 (1024 × 1024) 56.75 秒 51.54 秒 43.32 秒 官方 Turbo 的 DAILY SPECIAL 粉笔字完美排版,金属反射最通透
图像编辑换装重光照 (832 × 1248) 82.96 秒 60.87 秒 59.90 秒 官方 Turbo 的五官一致性与古典石阶夕阳光照最自然,刺绣立体感优秀

1. 图像编辑:人脸一致性、换装与背景重光照 (832 × 1248)

输入原图 vs 官方 Turbo (8步) vs Viggle r256 (6步) vs Viggle r128 (6步)
提示词:保持肖像五官发型一致,换墨绿刺绣丝绒礼服,背景替换为夕阳古典石阶花园

图像编辑横向对比

2. 文生图:写实肖像骨相与逆光发丝微高光 (832 × 1248)

官方 Turbo 使用原生 sample_sigmas 消除边缘白噪,发丝立体有呼吸感。

写实肖像对比

3. 文生图:密集静物、材质反射与黑板文字排版 (1024 × 1024)

考察咖啡机铜质微反射与 DAILY SPECIAL 英文粉笔字排版能力。

静物与文字对比


使用方法

1. 命令行 (CLI) 运行文生图 (8 步)

mflux-generate-qwen-2.1-edit   --model vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16   --prompt "写实摄影风格,一名年轻东亚女性站在窗边,逆光发丝微高光,眼神清澈,真实皮肤细腻微纹理,低饱和柔光,8k清晰度"   --width 832   --height 1248   --steps 8   --guidance 1.0   --scheduler qwen_turbo   --use-kv-cache   --output portrait_turbo.png

2. 命令行 (CLI) 单图编辑 (8 步)

mflux-generate-qwen-2.1-edit   --model vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16   --image-paths input.png   --prompt "保持人物五官骨相与发型完全一致,将背景替换为夕阳下的古典欧洲石阶花园,换穿墨绿色刺绣丝绒礼服"   --steps 8   --guidance 1.0   --scheduler qwen_turbo   --use-kv-cache   --output edited_turbo.png

3. 命令行 (CLI) 多图参考换装 (Multi-Reference)

mflux-generate-qwen-2.1-edit   --model vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16   --image-paths person_face.png dress_reference.png   --prompt "图一为人物面部与体态基准,图二为服装款式参考。保持图一人物容貌完全不变,换上图二的服装,自然站姿摄影"   --steps 8   --guidance 1.0   --scheduler qwen_turbo   --output outfit_swap.png

4. 叠加效果 LoRA 使用

支持在 8 步 Turbo 基础上继续串联各类风格或动作 LoRA:

mflux-generate-qwen-2.1-edit   --model vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16   --lora-paths /path/to/style_lora.safetensors   --lora-scales 0.85   --prompt "..."   --steps 8   --scheduler qwen_turbo   --output lora_turbo.png

English

This repository provides native MLX BF16 weights for Qwen/Qwen-Image-2.1-Turbo, combined with the uncensored Heretic vision-language text encoder and the original full-precision FP32 VAE.

It is designed specifically for running on Apple Silicon (M-series Macs) via mflux.

Key Features

  • 8-Step Turbo Inference: Employs the official Qwen-Image-2.1-Turbo DiT weights for fast, high-quality generation in just 8 steps without third-party acceleration LoRAs.
  • Uncensored Heretic Text Encoder: Full prompt fidelity without arbitrary safety refusals or filter interference.
  • Native qwen_turbo Scheduler: Aligned with the official 8-step fixed sample_sigmas grid ([1.0, 0.978453, 0.95418, 0.926626, 0.89508, 0.845148, 0.704534, 0.414568, 0.0]).
  • Complete Feature Set: Supports Text-to-Image, single-image editing, and multi-reference image composition out of the box.

Visual Comparison Gallery

Image Editing (Face Consistency & Relighting)

Image Editing Comparison

Text-to-Image (Hair Strands & Realistic Skin Details)

Portrait Comparison

Still Life & Text Rendering (Coffee Machine & Daily Special Chalkboard)

Still Life Comparison

Quick Start (CLI)

Text-to-Image

mflux-generate-qwen-2.1-edit   --model vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16   --prompt "A photorealistic portrait of an elegant East Asian woman by a window, natural rim lighting, sharp focus, fine skin textures, 8k resolution"   --width 832   --height 1248   --steps 8   --guidance 1.0   --scheduler qwen_turbo   --use-kv-cache   --output portrait.png

Image Editing

mflux-generate-qwen-2.1-edit   --model vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16   --image-paths input.png   --prompt "Keep facial features identical, change background to an autumn park with golden sunset glow, wear dark green velvet evening dress"   --steps 8   --guidance 1.0   --scheduler qwen_turbo   --use-kv-cache   --output edited.png

Acknowledgments & Citation

Downloads last month

-

Downloads are not tracked for this model. How to track
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vanch007/Qwen-Image-2.1-Turbo-Heretic-mlx-bf16

Finetuned
(7)
this model