Instructions to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Use Docker
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "dolev31/ProactiveInquirer-Qwen3-8B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "dolev31/ProactiveInquirer-Qwen3-8B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
- Ollama
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with Ollama:
ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with Docker Model Runner:
docker model run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
- Lemonade
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.ProactiveInquirer-Qwen3-8B-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use dolev31/ProactiveInquirer-Qwen3-8B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
ProactiveInquirer-Qwen3-8B-GGUF
Ido Levy1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM 2Weizmann Institute of Science
▶ The paper's example, step by step (22 seconds): the questioner trained with Q&D finds the account, the order with the boots and the size-8 boots, and the task is completed.
GGUF quantizations of the trained questioner from Asking for What Was Never Requested: Horizontal
and Vertical Proactivity in Agents, for llama.cpp, Ollama, LM Studio and Jan. They were made from the
merged model with llama.cpp
(commit 9adc7f4).
| File | Quantization | Size | Notes |
|---|---|---|---|
ProactiveInquirer-Qwen3-8B-Q4_K_M.gguf |
Q4_K_M | 5.0 GB | the usual choice, runs on a laptop |
ProactiveInquirer-Qwen3-8B-Q5_K_M.gguf |
Q5_K_M | 5.9 GB | a step closer to the full model |
ProactiveInquirer-Qwen3-8B-Q8_0.gguf |
Q8_0 | 8.7 GB | closest to the full model |
Before upload, each file ran the adapter card's two-turn example with greedy decoding. Every file asked the same two questions as the full-precision model: "Who directed the film The Great Flamarion?" and, once the evidence named the director, "Who was the spouse of film director Anthony Mann?".
Results
The results are the trained questioner's, as the paper reports them: see the adapter card's Results. The paper evaluated the unquantized model, not these files.
Run it
Ollama
ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M
llama.cpp
llama-server -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M --jinja
LM Studio: search for ProactiveInquirer in the model browser.
The questioner reads the prompt template it was trained on, in the adapter repository's
prompts/, and replies
with one JSON action per step: {"action": "ask", "question": ...} or {"action": "stop", ...}. It
was trained with Qwen3's thinking off, so keep it off: in Ollama run it with --think=false (or send
"think": false to its API), and with llama.cpp's server send "chat_template_kwargs": {"enable_thinking": false}.
curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
"messages": [{"role": "user", "content": "<the filled template>"}],
"chat_template_kwargs": {"enable_thinking": false},
"temperature": 0
}'
Limitations
- The questioner's own limitations, from the paper: it has learned what to ask more readily than when to stop, the extra evidence it finds does not yet translate into better final answers, and its user-facing results come from a simulated customer, not from real people.
- It is a component inside an agent, meant to be called with its prompt template. It is not a chat assistant, and it was trained and evaluated in English.
- Quantization can change the model's choices. Each file was checked on one example, as above: a check, not an evaluation.
Citation
@article{levy2026asking,
title = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
author = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
journal = {arXiv preprint arXiv:2609.37236},
url = {https://arxiv.org/abs/2609.37236},
year = {2026}
}
License
Apache-2.0, like the base model Qwen3-8B.
- Downloads last month
- 509
4-bit
5-bit
8-bit
Model tree for dolev31/ProactiveInquirer-Qwen3-8B-GGUF
Base model
Qwen/Qwen3-8B-Base