GrepSeek-Qwen3.5-9B-SFT

The cold-start SFT policy for GrepSeek — a Direct Corpus Interaction (DCI) search agent that answers questions by issuing Unix shell commands (rg, grep, head, …) directly over a 21M-passage Wikipedia corpus, instead of retrieving from a dense/sparse index. This checkpoint is Qwen/Qwen3.5-9B supervised-fine-tuned on the cold-start trajectories; it is the initialization for the RL stage and corresponds to the "w/o GRPO" ablation in the paper.

📄 GrepSeek: Training Search Agents for Direct Corpus Interaction · 💻 https://github.com/alirezasalemi7/grepseek

  • Base: Qwen/Qwen3.5-9B
  • Training data: alireza7/GrepSeek-ColdStart-SFT-10k (5k NQ + 5k HotpotQA, Answer-Aware-Tutor / Answer-Blind-Planner trajectories)
  • Recipe: verl FSDP, 1 epoch, AdamW, peak LR 5e-6 (constant + 5% warmup), 16,384-token sequences, bf16, Ulysses SP=4, on 4×A100-80GB.
  • RL-optimized successor: alireza7/GrepSeek-Qwen3.5-9B-GRPO

Why SFT first: RL directly from the base model over millions of documents is unstable (overly broad queries, context blow-ups, OOM). Cold-start SFT instills concise, causally-grounded shell-search behavior — establishing the low-level retrieval "primitives" (fixed-string -F matching, | head truncation, cascaded rg ... | rg ... filtering) that RL later refines.

⚠️ A tool-using agent, not a standalone chatbot

The model emits <tool_call> shell commands that must be executed against the Wikipedia corpus and fed back as <tool_response> turns. To use it you need: (1) the corpus PeterJinGo/wiki-18-corpus, (2) a tool-calling vLLM server, and (3) the GrepSeek inference harness (grep tool

Usage

git clone https://github.com/alirezasalemi7/grepseek && cd grepseek
# env: TRAINING_ENV.md  ·  corpus: cold_start_sft/download_corpus.py

# 1. serve this checkpoint
MODEL_PATH=alireza7/GrepSeek-Qwen3.5-9B-SFT bash rl/serve_rl.sh         # -> http://localhost:10730/v1

# 2. run the agent (paper inference: temperature 0.6, <=6 turns, 16k context)
GREPSEEK_CORPUS_ROOT=/path/to/wiki_18_corpus \
  bash inference/run_inference.sh --base_url http://localhost:10730/v1 \
    --model grepseek --temperature 0.6 --input my_questions.jsonl --out_dir out

Evaluation (token-F1 / EM, micro-average over 7 QA benchmarks)

This SFT-only policy already substantially beats the untuned base model, but RL adds large gains on multi-hop reasoning:

variant micro-avg F1 micro-avg EM
base (no SFT, no RL) 0.3314 0.2836
this model (SFT only) 0.4249 0.3569
+ GRPO → GrepSeek-Qwen3.5-9B-GRPO 0.5691 0.4948

(7 benchmarks: NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue, Bamboogle; trained only on NQ + HotpotQA, the rest are out-of-distribution.)

License

Inherits the license of the base model Qwen/Qwen3.5-9B — confirm and update the license field above if needed.

Citation

@misc{salemi2026grepseektrainingsearchagents,
      title={GrepSeek: Training Search Agents for Direct Corpus Interaction},
      author={Alireza Salemi and Chang Zeng and Atharva Nijasure and Jui-Hui Chung and Razieh Rahimi and Fernando Diaz and Hamed Zamani},
      year={2026},
      eprint={2605.29307},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2605.29307},
}
Downloads last month
45
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for alireza7/GrepSeek-Qwen3.5-9B-SFT

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(995)
this model
Finetunes
1 model

Dataset used to train alireza7/GrepSeek-Qwen3.5-9B-SFT

Collection including alireza7/GrepSeek-Qwen3.5-9B-SFT

Paper for alireza7/GrepSeek-Qwen3.5-9B-SFT