Qwen-Image-2.1 Turbo โ experimental browser ONNX pack
Built with Qwen. For research and evaluation under the included Qwen Research License. Inference runs on the visitor's hardware using WebGPU. No hosted inference is included.
The Viggle rank-64 four-step Turbo LoRA is merged into the base denoiser, then quantized to Q4 MatMulNBits with block size 128. Text and vision encoders retain the qualified mixed Q8/fp16 exports; the VAE is unchanged in precision. External weights are split into small shards. The base model revision is 790c92633540aa0cb11d9abf19eb46d861714758; the adapter revision is bafc91e4cc934f5fb1406b22496a0bed9b99c548.
User testing completed 256px Q4 Turbo browser generation. Larger resolutions and reference editing remain experimental; automated tensor and lifecycle checks do not establish image quality. Four steps is the trained configuration; two steps is an unqualified speed comparison. Q2 is excluded because user testing produced noise.
Requires the FreeGen browser pipeline at https://github.com/CatsWithKeyboards/sharemyai (src/freegen/qwen-image21.js), a desktop WebGPU adapter with shader-f16, and Float16Array browser support. Up to 17.2 GB of assets are fetched as needed and cached locally; generation uses a subset of reference-only assets. GPU and system memory requirements vary with size and loading settings.
This is a modified derivative: ONNX export, external-data splitting, mixed precision encoder conversion, LoRA fusion, and four-bit denoiser quantization. See LICENSE and NOTICE for upstream terms and attribution. No commercial license is granted by this repository.
Model tree for cgb/Qwen-Image-2.1-Turbo-ONNX
Base model
Qwen/Qwen-Image-2.1