Qwen3-1.7B LiteRT-LM Model

This repository contains LiteRT-LM variants of Qwen/Qwen3-1.7B optimized for on-device text generation.

Available Artifact

File Quantization Recipe Context Size
Qwen3_1.7B.litertlm dynamic_wi8_afp32 - 2.1 GB
Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm dynamic_wi4b32_afp32 4096 932 MB

Conversion Notes

Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm is a dynamic INT4 variant (block-32 weights, FP32 activations). It was converted through the LiteRT Torch (litert-torch) path and quantized with AI Edge Quantizer. This artifact incorporates LiteRT-LM GPU graph optimizations, including composite ops for RoPE, fused QKV, and fused Gate/Up projections, and is configured with static prefill memory allocation.

How to Use

Command-Line Interface

  1. Install the prerequisites:
   pip install litert-lm
  1. Run the command in CLI:
   litert-lm run --from-huggingface-repo=litert-community/Qwen3-1.7B Qwen3_1.7B.litertlm --prompt="Write me a poem on nature"

Edge Gallery App

  1. Get the App: Install the app from Google Play or download the latest APK from the GitHub releases page.
  2. Importing the Model: Navigate to the Model manager within the app and click the "+" (plus) icon in the bottom-right corner. Two options will appear:
    • Import from HF (Recommended): Select this option, and a dialog box will appear showing an example Hugging Face model URL. Enter the HF link for the desired .litertlm model and click submit. The model will then appear in your list, and you can proceed to download it (a Hugging Face account login is required).
    • From local model file: First, download the .litertlm model directly to your Android device, OR download it to your computer and push it via ADB (e.g., adb push Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm /sdcard/Download/). Then, select this option, choose the downloaded file from your storage, configure your preferred parameters, and tap "Import".

For full details on importing models and other features, see the Edge Gallery App Wiki.

To build the demo app from source, please follow the instructions from the GitHub repository.

Python

  1. Install the prerequisites:
   pip install litert-lm huggingface_hub
  1. Download the model file:
   from huggingface_hub import hf_hub_download

   model_path = hf_hub_download(
       repo_id="litert-community/Qwen3-1.7B",
       filename="Qwen3_1.7B.litertlm"
   )
  1. Run inference:
   import litert_lm

   litert_lm.set_min_log_severity(litert_lm.LogSeverity.ERROR)  # Hide log for TUI app

   with litert_lm.Engine(model_path) as engine:
     with engine.create_conversation() as conversation:
       while True:
         user_input = input("\n>>> ")
         for chunk in conversation.send_message_async(user_input):
           print(chunk["content"][0]["text"], end="", flush=True)

Performance (on device, measured)

Community measurement on a physical Samsung Galaxy S26 (SM-S942Q, Snapdragon 8 Elite Gen 5 / SM8850, Android 16): litert_lm_advanced_main from the litert-lm v0.16.0 release, GPU backend OpenCL (LITERT_CL), one fixed 205-token prompt text (202 tokens under this tokenizer), --benchmark. Two runs per backend taken back-to-back β€” cells show the range (the phone warms across a series; the CPU's first run also pays cold file I/O). Peak RSS is the process VmHWM. Before quoting, every file/backend pair was run with a real prompt and produced a correct text answer; the GPU rows are full delegation (e.g. Qwen3-1.7B_dynamic_wi4b32_afp32: decode 994/994 ops on LITERT_CL, nothing on the CPU fallback).

File Backend Prefill (202 tok) Decode Time-to-first-token Peak RSS
Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm GPU (OpenCL) 532–568 tok/s 40.5–40.7 tok/s 0.38–0.40 s 1036 MB
Qwen3-1.7B_dynamic_wi4b32_afp32.litertlm CPU (XNNPACK) 48–76 tok/s 8.0–8.3 tok/s 2.8–4.3 s 2523 MB
Qwen3_1.7B.litertlm GPU (OpenCL) 312–504 tok/s 14.9–18.1 tok/s 0.46–0.71 s 854 MB
Qwen3_1.7B.litertlm CPU (XNNPACK) 238–245 tok/s 7.7–7.8 tok/s 0.95–0.98 s 3185 MB

What the table says:

  • On the S26 GPU the wi4b32 file is the one to run: 40.7 tok/s decode against 18.1 for the int8 file β€” the GPU graph optimizations described in Conversion Notes above, visible in measurement. It is also the fastest prefill and the lowest time-to-first-token on the card.
  • The wi4b32 file is a GPU specialist. On CPU it decodes at 8 tok/s and prefills at 48–76 tok/s β€” int4 block-32 weights suit the GPU path, not XNNPACK. If you need CPU, the int8 Qwen3_1.7B.litertlm prefills 3–5Γ— faster than wi4b32 on the same CPU.
  • Peak RSS is 2.4–3.7Γ— lower on the GPU for both files (854–1036 MB against 2.5–3.2 GB).
  • GPU engine init is the one cost the table hides: 4.4–8.3 s per process against 2.2–3.5 s on CPU.

Integration

Ready to integrate this into your product? Get started in the LiteRT-LM documentation.

Downloads last month
3,513
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ 2 Ask for provider support

Model tree for litert-community/Qwen3-1.7B

Finetuned
Qwen/Qwen3-1.7B
Quantized
(353)
this model

Collections including litert-community/Qwen3-1.7B