Instructions to use GadflyII/Qwen3-Coder-Next-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

Libraries

How to use GadflyII/Qwen3-Coder-Next-NVFP4 with Transformers:

# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="GadflyII/Qwen3-Coder-Next-NVFP4")
messages = [
    {"role": "user", "content": "Who are you?"},
]
pipe(messages)

# Load model directly
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("GadflyII/Qwen3-Coder-Next-NVFP4")
model = AutoModelForCausalLM.from_pretrained("GadflyII/Qwen3-Coder-Next-NVFP4")
messages = [
    {"role": "user", "content": "Who are you?"},
]
inputs = tokenizer.apply_chat_template(
	messages,
	add_generation_prompt=True,
	tokenize=True,
	return_dict=True,
	return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=40)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))

Notebooks
Google Colab
Kaggle
Local Apps Settings

vLLM

How to use GadflyII/Qwen3-Coder-Next-NVFP4 with vLLM:

Install from pip and serve model

# Install vLLM from pip:
pip install vllm
# Start the vLLM server:
vllm serve "GadflyII/Qwen3-Coder-Next-NVFP4"
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:8000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "GadflyII/Qwen3-Coder-Next-NVFP4",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker

docker model run hf.co/GadflyII/Qwen3-Coder-Next-NVFP4

SGLang

How to use GadflyII/Qwen3-Coder-Next-NVFP4 with SGLang:

Install from pip and serve model

# Install SGLang from pip:
pip install sglang
# Start the SGLang server:
python3 -m sglang.launch_server \
    --model-path "GadflyII/Qwen3-Coder-Next-NVFP4" \
    --host 0.0.0.0 \
    --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "GadflyII/Qwen3-Coder-Next-NVFP4",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Use Docker images

docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server \
        --model-path "GadflyII/Qwen3-Coder-Next-NVFP4" \
        --host 0.0.0.0 \
        --port 30000
# Call the server using curl (OpenAI-compatible API):
curl -X POST "http://localhost:30000/v1/chat/completions" \
	-H "Content-Type: application/json" \
	--data '{
		"model": "GadflyII/Qwen3-Coder-Next-NVFP4",
		"messages": [
			{
				"role": "user",
				"content": "What is the capital of France?"
			}
		]
	}'

Docker Model Runner
How to use GadflyII/Qwen3-Coder-Next-NVFP4 with Docker Model Runner:
```
docker model run hf.co/GadflyII/Qwen3-Coder-Next-NVFP4
```

Is it possible for vllm (or other) to run this NVFP4 on RTX 5090 with 128GB RAM?

by vladulidlo - opened Feb 27

Discussion

vladulidlo

Feb 27

I tried multiple vllm installs including 0.16.0 final and all I get are errors with many different combinations of parameters including cpu offloading. Consulted with ChatGPT, tried over 100 attempts but it never worked for me. Always crashing:

(EngineCore_DP0 pid=84120) WARNING 02-27 14:11:29 [decorators.py:590] unable to save AOT compiled function to /home/vlady/.cache/vllm/torch_compile_cache/torch_aot_compile/60dd78938ac85ab23fa92572573f3271e84bc07d0b4c16e1e978a186b85cc995/rank_0_0/model:
(EngineCore_DP0 pid=84120) INFO 02-27 14:11:30 [gpu_worker.py:405] Available KV cache memory: 0.94 GiB
(EngineCore_DP0 pid=84120) INFO 02-27 14:11:30 [kv_cache_utils.py:1314] GPU KV cache size: 9,792 tokens
(EngineCore_DP0 pid=84120) INFO 02-27 14:11:30 [kv_cache_utils.py:1319] Maximum concurrency for 16,384 tokens per request: 2.21x
(EngineCore_DP0 pid=84120) ERROR 02-27 14:11:30 [core.py:1078] EngineCore failed to start.

torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB. GPU 0 has a total capacity of 31.36 GiB of which 357.62 MiB is free. Process 3134 has 703.11 MiB memory in use. Including non-PyTorch memory, this process has 29.33 GiB memory in use. Of the allocated memory 28.54 GiB is allocated by PyTorch, and 143.66 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[rank0]:[W227 14:23:16.541489619 ProcessGroupNCCL.cpp:1524] Warning: WARNING: destroy_process_group() was not called before program exit, which can leak resources. For more info, please see https://pytorch.org/docs/stable/distributed.html#shutdown (function operator())

vladulidlo changed discussion title from Is it possible to run this on RTX 5090 with 128GB RAM? to Is it possible to run this NVFP4 on RTX 5090 with 128GB RAM? Feb 27

vladulidlo changed discussion title from Is it possible to run this NVFP4 on RTX 5090 with 128GB RAM? to Is it possible fro vllm (or other) to run this NVFP4 on RTX 5090 with 128GB RAM? Feb 27

vladulidlo changed discussion title from Is it possible fro vllm (or other) to run this NVFP4 on RTX 5090 with 128GB RAM? to Is it possible for vllm (or other) to run this NVFP4 on RTX 5090 with 128GB RAM? Feb 27

Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.

Tap or paste here to upload images

· Sign up or log in to comment