Text Generation
Transformers
Safetensors
English
qwen3_5
image-text-to-text
grug
coding
tool-use
agentic
mtp
conversational
Instructions to use ProCreations/grug-27b-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ProCreations/grug-27b-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="ProCreations/grug-27b-v2") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ProCreations/grug-27b-v2") model = AutoModelForMultimodalLM.from_pretrained("ProCreations/grug-27b-v2", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ProCreations/grug-27b-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ProCreations/grug-27b-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ProCreations/grug-27b-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ProCreations/grug-27b-v2
- SGLang
How to use ProCreations/grug-27b-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ProCreations/grug-27b-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ProCreations/grug-27b-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ProCreations/grug-27b-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ProCreations/grug-27b-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use ProCreations/grug-27b-v2 with Docker Model Runner:
docker model run hf.co/ProCreations/grug-27b-v2
Download gguf_conversion.json from ProCreations/grug-27b-v2: direct link, hf CLI and curl.
- Browser
- Download file 5.22 kB
-
https://huggingface.co/ProCreations/grug-27b-v2/resolve/main/gguf_conversion.json
- Command line
-
hf download hf://ProCreations/grug-27b-v2/gguf_conversion.json
-
curl -L -o gguf_conversion.json https://huggingface.co/ProCreations/grug-27b-v2/resolve/main/gguf_conversion.json
5.22 kB
| { | |
| "conversion_transformers": "5.16.1", | |
| "source_model": "ProCreations/grug-27b-v2-candidate-20260912-qwen", | |
| "source_revision": "79a2b2b81bf6ee0123da02467cf187db04e03708", | |
| "llama_cpp_commit": "2a3005c23f60cb38dab70b8ea2ddbd969bcf3e87", | |
| "mtp_embedded_in_every_text_gguf": true, | |
| "artifacts": [ | |
| { | |
| "file": "grug-27b-v2-Q4_K_M.gguf", | |
| "bytes": 16998721440, | |
| "sha256": "5db0549d718c0accbe5979af0a959099f7165e3b0bea7a8604c7a1bd6872363f", | |
| "mtp_tensor_count": 15, | |
| "mtp_matrix_quantization": "Q8_0", | |
| "smoke_stdout": "| model | size | params | backend | threads | test | t/s |\n| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |\n| qwen35 27B Q4_K - Medium | 15.82 GiB | 27.32 B | CPU | 24 | pp16 | 21.50 ± 0.00 |\n| qwen35 27B Q4_K - Medium | 15.82 GiB | 27.32 B | CPU | 24 | tg8 | 3.16 ± 0.00 |\n\nbuild: 2a3005c (1)\n", | |
| "smoke_stderr": "" | |
| }, | |
| { | |
| "file": "grug-27b-v2-Q5_K_M.gguf", | |
| "bytes": 19682420640, | |
| "sha256": "bf02d822be68618f9dca224065aa47e0cbb478de8b432a0ccc672a28d06438f8", | |
| "mtp_tensor_count": 15, | |
| "mtp_matrix_quantization": "Q8_0", | |
| "smoke_stdout": "| model | size | params | backend | threads | test | t/s |\n| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |\n| qwen35 27B Q5_K - Medium | 18.32 GiB | 27.32 B | CPU | 24 | pp16 | 13.27 ± 0.00 |\n| qwen35 27B Q5_K - Medium | 18.32 GiB | 27.32 B | CPU | 24 | tg8 | 2.81 ± 0.00 |\n\nbuild: 2a3005c (1)\n", | |
| "smoke_stderr": "" | |
| }, | |
| { | |
| "file": "grug-27b-v2-Q6_K.gguf", | |
| "bytes": 22533851040, | |
| "sha256": "c110ad0b8863c06dacd753c41c301ba3d7c4901e5776d0d80e2396b6c184fd0c", | |
| "mtp_tensor_count": 15, | |
| "mtp_matrix_quantization": "Q8_0", | |
| "smoke_stdout": "| model | size | params | backend | threads | test | t/s |\n| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |\n| qwen35 27B Q6_K | 20.98 GiB | 27.32 B | CPU | 24 | pp16 | 15.91 ± 0.00 |\n| qwen35 27B Q6_K | 20.98 GiB | 27.32 B | CPU | 24 | tg8 | 2.53 ± 0.00 |\n\nbuild: 2a3005c (1)\n", | |
| "smoke_stderr": "" | |
| }, | |
| { | |
| "file": "grug-27b-v2-Q8_0.gguf", | |
| "bytes": 29047084960, | |
| "sha256": "01431cb3864ee4cd7ef5a9e5d73079194f6c3bca66d0cac5ce19502e2e948f44", | |
| "mtp_tensor_count": 15, | |
| "mtp_matrix_quantization": "Q8_0", | |
| "smoke_stdout": "| model | size | params | backend | threads | test | t/s |\n| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |\n| qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CPU | 24 | pp16 | 17.03 ± 0.00 |\n| qwen35 27B Q8_0 | 27.04 GiB | 27.32 B | CPU | 24 | tg8 | 1.92 ± 0.00 |\n\nbuild: 2a3005c (1)\n", | |
| "smoke_stderr": "" | |
| }, | |
| { | |
| "file": "grug-27b-v2-Q3_K_M.gguf", | |
| "bytes": 13752764320, | |
| "sha256": "3b36b201c4c7053bea301533d2804aa41912f701c19e320bffe0cfa5e72e07c4", | |
| "mtp_tensor_count": 15, | |
| "mtp_matrix_quantization": "Q8_0", | |
| "smoke_stdout": "| model | size | params | backend | threads | test | t/s |\n| ------------------------------ | ---------: | ---------: | ---------- | ------: | --------------: | -------------------: |\n| qwen35 27B Q3_K - Medium | 12.80 GiB | 27.32 B | CPU | 24 | pp16 | 17.12 ± 0.00 |\n| qwen35 27B Q3_K - Medium | 12.80 GiB | 27.32 B | CPU | 24 | tg8 | 4.01 ± 0.00 |\n\nbuild: 2a3005c (1)\n", | |
| "smoke_stderr": "" | |
| }, | |
| { | |
| "file": "mmproj-grug-27b-v2-F16.gguf", | |
| "bytes": 927606976, | |
| "sha256": "1335119dd8d9f9e8c4b0a65de12b7d218cfc47282c706a0c6be0dd85be209da6" | |
| } | |
| ], | |
| "elapsed_seconds": 1267.8098618984222, | |
| "conversion_time_remaining_validation": "GPU generation and embedded MTP parity must pass before public release", | |
| "post_conversion_validation": { | |
| "completed": true, | |
| "record": "results/runtime_validation.json", | |
| "basic_functional_checks": "12/12 for every quantization with MTP off and on", | |
| "native_api_checks": "14/15 for every mode; reasoning alias is unsupported by the pinned llama-server parser", | |
| "token_parity": "Measured and reported per quantization; not a bit-identical decoding guarantee", | |
| "coding_grader": "Corrected final-answer replay on HF CPU Jobs; see results/coding_rescore.json" | |
| } | |
| } | |