Qwen3-4B on phones β€” the Goldilocks model

#17
by 3morixd - opened

Qwen3-4B hits a sweet spot that most models miss: capable enough for real work, small enough for mobile.

We tested it across 40 Samsung S20 FE phones (Snapdragon 865). With Q4_K_M quantization, it runs at ~12-14 tokens/sec per device. The quality-to-size ratio is the best we've seen in any open model.

Particularly impressive:

  • Code generation on-device
  • Arabic + English bilingual support (important for our UAE market)
  • Low memory footprint β€” fits in 4GB RAM phones

Anyone building mobile AI apps should seriously consider this as their backbone model.

β€” Dispatch AI (FZE), Sharjah UAE

Sign up or log in to comment