Qwen-Image-2.1 Turbo โ€” experimental browser ONNX pack

Built with Qwen. For research and evaluation under the included Qwen Research License. Inference runs on the visitor's hardware using WebGPU. No hosted inference is included.

The Viggle rank-64 four-step Turbo LoRA is merged into the base denoiser, then quantized to Q4 MatMulNBits with block size 128. Text and vision encoders retain the qualified mixed Q8/fp16 exports; the VAE is unchanged in precision. External weights are split into small shards. The base model revision is 790c92633540aa0cb11d9abf19eb46d861714758; the adapter revision is bafc91e4cc934f5fb1406b22496a0bed9b99c548.

User testing completed 256px Q4 Turbo browser generation. Larger resolutions and reference editing remain experimental; automated tensor and lifecycle checks do not establish image quality. Four steps is the trained configuration; two steps is an unqualified speed comparison. Q2 is excluded because user testing produced noise.

Requires the FreeGen browser pipeline at https://github.com/CatsWithKeyboards/sharemyai (src/freegen/qwen-image21.js), a desktop WebGPU adapter with shader-f16, and Float16Array browser support. Up to 17.2 GB of assets are fetched as needed and cached locally; generation uses a subset of reference-only assets. GPU and system memory requirements vary with size and loading settings.

This is a modified derivative: ONNX export, external-data splitting, mixed precision encoder conversion, LoRA fusion, and four-bit denoiser quantization. See LICENSE and NOTICE for upstream terms and attribution. No commercial license is granted by this repository.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for cgb/Qwen-Image-2.1-Turbo-ONNX

Quantized
(125)
this model