Instructions to use wrayy/Qwenity3.6-27B-msv2-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Use Docker
docker model run hf.co/wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "wrayy/Qwenity3.6-27B-msv2-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wrayy/Qwenity3.6-27B-msv2-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
- Ollama
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with Ollama:
ollama run hf.co/wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with Docker Model Runner:
docker model run hf.co/wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
- Lemonade
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Qwenity3.6-27B-msv2-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use wrayy/Qwenity3.6-27B-msv2-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "wrayy/Qwenity3.6-27B-msv2-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwenity 3.6 27B (msv2) — GGUF
GGUF quantizations of wrayy/Qwenity3.6-27B-msv2 —
a specialized version of Qwen3.6-27B, fine-tuned to act as an expert assistant for Unity game development.
Ready to run locally with llama.cpp and compatible runtimes.
Model summary
Quantized, GGUF builds of the merged Qwenity 3.6 27B model (text / language model). It specializes the base model for the Unity engine and C# scripting while retaining its general capabilities.
- Developed by: wrayy
- Source model:
wrayy/Qwenity3.6-27B-msv2(16-bit merged) - Original base:
Qwen/Qwen3.6-27B - Format: GGUF (text generation),
qwen3_5architecture - Language: English
- License: Apache-2.0 (inherited from the base model)
Files
Two quantizations are provided. Each is split into <5 GB GGUF shards; download all shards of the quant you want into one folder.
| Quant | Shards | Total size | Notes |
|---|---|---|---|
| Q4_K_M | …-Q4_K_M-00001-of-00005.gguf → …-00005-of-00005.gguf (5 files) |
~16.5 GB | Recommended — best size/quality balance; runs on a 24 GB GPU or CPU. |
| Q8_0 | …-Q8_0-00001-of-00008.gguf → …-00008-of-00008.gguf (8 files) |
~28.6 GB | Near-lossless; highest quality. |
These are native GGUF splits: point your runtime at the first shard (
-00001-of-…) and it loads the rest automatically. (You can also merge them back withllama-gguf-split --merge.)
Intended use
Purpose-built to help developers build with Unity: C# gameplay scripting and the Unity scripting API, the Editor and project workflow, rendering, physics, animation, UI, and Unity-oriented debugging and tool use. It behaves as a chat / instruction-following assistant.
How to use
Runtime requirement. This is the new
qwen3_5(Qwen3.6) architecture. Use allama.cppbuild (and Ollama / LM Studio versions) that include Qwen3.5 / 3.6 support; older builds will refuse to load it.
llama.cpp
# downloads all 5 shards into the current folder, then point at shard 1:
llama-server -m Qwenity3.6-27B-msv2-Q4_K_M-00001-of-00005.gguf -ngl 99 -c 4096 --host 127.0.0.1 --port 8080
Ollama
printf 'FROM ./Qwenity3.6-27B-msv2-Q4_K_M-00001-of-00005.gguf\n' > Modelfile
ollama create qwenity3.6-27b -f Modelfile
ollama run qwenity3.6-27b "How do I raycast from the camera in Unity C#?"
LM Studio and other GGUF apps also load split files (point them at the first shard) with a Qwen3.5/3.6-capable runtime.
Training
- Base model:
Qwen/Qwen3.6-27B - Method: LoRA fine-tuning at 16-bit via Unsloth + TRL (SFT), merged to full bf16 (checkpoint 1200), then converted to GGUF and quantized with
llama.cpp. - Training data:
wrayy/unity.masterset.sbv2— a privately compiled and synthesized instruction dataset for Unity development (English; conversational instruction / Q&A spanning C# scripting, the Editor, rendering, physics, UI, and tool use). The dataset is private / research-only, and its contents are not publicly disclosed.
Limitations
- Domain-focused: optimized for Unity / C# / game development; less reliable elsewhere.
- Verify generated code: like all LLMs it can produce incorrect, outdated, or insecure code — review and test before use.
- Runtime: requires a Qwen3.5/3.6-capable
llama.cpp/ Ollama / LM Studio build.
License
Released under Apache-2.0, inherited from the base model
Qwen/Qwen3.6-27B.
The training dataset is private and research-only.
- Downloads last month
- 76
4-bit
8-bit