Qwenity 3.6 27B (msv2) — GGUF

GGUF quantizations of wrayy/Qwenity3.6-27B-msv2 — a specialized version of Qwen3.6-27B, fine-tuned to act as an expert assistant for Unity game development. Ready to run locally with llama.cpp and compatible runtimes.

Model summary

Quantized, GGUF builds of the merged Qwenity 3.6 27B model (text / language model). It specializes the base model for the Unity engine and C# scripting while retaining its general capabilities.

  • Developed by: wrayy
  • Source model: wrayy/Qwenity3.6-27B-msv2 (16-bit merged)
  • Original base: Qwen/Qwen3.6-27B
  • Format: GGUF (text generation), qwen3_5 architecture
  • Language: English
  • License: Apache-2.0 (inherited from the base model)

Files

Two quantizations are provided. Each is split into <5 GB GGUF shards; download all shards of the quant you want into one folder.

Quant Shards Total size Notes
Q4_K_M …-Q4_K_M-00001-of-00005.gguf → …-00005-of-00005.gguf (5 files) ~16.5 GB Recommended — best size/quality balance; runs on a 24 GB GPU or CPU.
Q8_0 …-Q8_0-00001-of-00008.gguf → …-00008-of-00008.gguf (8 files) ~28.6 GB Near-lossless; highest quality.

These are native GGUF splits: point your runtime at the first shard (-00001-of-…) and it loads the rest automatically. (You can also merge them back with llama-gguf-split --merge.)

Intended use

Purpose-built to help developers build with Unity: C# gameplay scripting and the Unity scripting API, the Editor and project workflow, rendering, physics, animation, UI, and Unity-oriented debugging and tool use. It behaves as a chat / instruction-following assistant.

How to use

Runtime requirement. This is the new qwen3_5 (Qwen3.6) architecture. Use a llama.cpp build (and Ollama / LM Studio versions) that include Qwen3.5 / 3.6 support; older builds will refuse to load it.

llama.cpp

# downloads all 5 shards into the current folder, then point at shard 1:
llama-server -m Qwenity3.6-27B-msv2-Q4_K_M-00001-of-00005.gguf -ngl 99 -c 4096 --host 127.0.0.1 --port 8080

Ollama

printf 'FROM ./Qwenity3.6-27B-msv2-Q4_K_M-00001-of-00005.gguf\n' > Modelfile
ollama create qwenity3.6-27b -f Modelfile
ollama run qwenity3.6-27b "How do I raycast from the camera in Unity C#?"

LM Studio and other GGUF apps also load split files (point them at the first shard) with a Qwen3.5/3.6-capable runtime.

Training

  • Base model: Qwen/Qwen3.6-27B
  • Method: LoRA fine-tuning at 16-bit via Unsloth + TRL (SFT), merged to full bf16 (checkpoint 1200), then converted to GGUF and quantized with llama.cpp.
  • Training data: wrayy/unity.masterset.sbv2 — a privately compiled and synthesized instruction dataset for Unity development (English; conversational instruction / Q&A spanning C# scripting, the Editor, rendering, physics, UI, and tool use). The dataset is private / research-only, and its contents are not publicly disclosed.

Limitations

  • Domain-focused: optimized for Unity / C# / game development; less reliable elsewhere.
  • Verify generated code: like all LLMs it can produce incorrect, outdated, or insecure code — review and test before use.
  • Runtime: requires a Qwen3.5/3.6-capable llama.cpp / Ollama / LM Studio build.

License

Released under Apache-2.0, inherited from the base model Qwen/Qwen3.6-27B. The training dataset is private and research-only.

Downloads last month
76
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for wrayy/Qwenity3.6-27B-msv2-GGUF

Base model

Qwen/Qwen3.6-27B
Quantized
(1)
this model

Collection including wrayy/Qwenity3.6-27B-msv2-GGUF