ProactiveInquirer-Qwen3-8B-GGUF

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

Ido Levy1,2 · Asaf Yehudai1 · Segev Shlomov1 · Asaf Adi1 · Leshem Choshen1,2
1IBM   2Weizmann Institute of Science

Project page Paper Code License

▶ The paper's example, step by step (22 seconds): the questioner trained with Q&D finds the account, the order with the boots and the size-8 boots, and the task is completed.

GGUF quantizations of the trained questioner from Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents, for llama.cpp, Ollama, LM Studio and Jan. They were made from the merged model with llama.cpp (commit 9adc7f4).

File Quantization Size Notes
ProactiveInquirer-Qwen3-8B-Q4_K_M.gguf Q4_K_M 5.0 GB the usual choice, runs on a laptop
ProactiveInquirer-Qwen3-8B-Q5_K_M.gguf Q5_K_M 5.9 GB a step closer to the full model
ProactiveInquirer-Qwen3-8B-Q8_0.gguf Q8_0 8.7 GB closest to the full model

Before upload, each file ran the adapter card's two-turn example with greedy decoding. Every file asked the same two questions as the full-precision model: "Who directed the film The Great Flamarion?" and, once the evidence named the director, "Who was the spouse of film director Anthony Mann?".

Results

The results are the trained questioner's, as the paper reports them: see the adapter card's Results. The paper evaluated the unquantized model, not these files.

Run it

Ollama

ollama run hf.co/dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M

llama.cpp

llama-server -hf dolev31/ProactiveInquirer-Qwen3-8B-GGUF:Q4_K_M --jinja

LM Studio: search for ProactiveInquirer in the model browser.

The questioner reads the prompt template it was trained on, in the adapter repository's prompts/, and replies with one JSON action per step: {"action": "ask", "question": ...} or {"action": "stop", ...}. It was trained with Qwen3's thinking off, so keep it off: in Ollama run it with --think=false (or send "think": false to its API), and with llama.cpp's server send "chat_template_kwargs": {"enable_thinking": false}.

curl http://localhost:8080/v1/chat/completions -H "Content-Type: application/json" -d '{
  "messages": [{"role": "user", "content": "<the filled template>"}],
  "chat_template_kwargs": {"enable_thinking": false},
  "temperature": 0
}'

Limitations

  • The questioner's own limitations, from the paper: it has learned what to ask more readily than when to stop, the extra evidence it finds does not yet translate into better final answers, and its user-facing results come from a simulated customer, not from real people.
  • It is a component inside an agent, meant to be called with its prompt template. It is not a chat assistant, and it was trained and evaluated in English.
  • Quantization can change the model's choices. Each file was checked on one example, as above: a check, not an evaluation.

Citation

@article{levy2026asking,
  title   = {Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents},
  author  = {Levy, Ido and Yehudai, Asaf and Shlomov, Segev and Adi, Asaf and Choshen, Leshem},
  journal = {arXiv preprint arXiv:2609.37236},
  url     = {https://arxiv.org/abs/2609.37236},
  year    = {2026}
}

License

Apache-2.0, like the base model Qwen3-8B.

Downloads last month
509
GGUF
Model size
8B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dolev31/ProactiveInquirer-Qwen3-8B-GGUF

Finetuned
Qwen/Qwen3-8B
Quantized
(1)
this model

Collection including dolev31/ProactiveInquirer-Qwen3-8B-GGUF

Paper for dolev31/ProactiveInquirer-Qwen3-8B-GGUF