Instructions to use alireza7/GrepSeek-Qwen3.5-9B-SFT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use alireza7/GrepSeek-Qwen3.5-9B-SFT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="alireza7/GrepSeek-Qwen3.5-9B-SFT") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("alireza7/GrepSeek-Qwen3.5-9B-SFT") model = AutoModelForMultimodalLM.from_pretrained("alireza7/GrepSeek-Qwen3.5-9B-SFT", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use alireza7/GrepSeek-Qwen3.5-9B-SFT with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "alireza7/GrepSeek-Qwen3.5-9B-SFT" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alireza7/GrepSeek-Qwen3.5-9B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/alireza7/GrepSeek-Qwen3.5-9B-SFT
- SGLang
How to use alireza7/GrepSeek-Qwen3.5-9B-SFT with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "alireza7/GrepSeek-Qwen3.5-9B-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alireza7/GrepSeek-Qwen3.5-9B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "alireza7/GrepSeek-Qwen3.5-9B-SFT" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "alireza7/GrepSeek-Qwen3.5-9B-SFT", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use alireza7/GrepSeek-Qwen3.5-9B-SFT with Docker Model Runner:
docker model run hf.co/alireza7/GrepSeek-Qwen3.5-9B-SFT
GrepSeek-Qwen3.5-9B-SFT
The cold-start SFT policy for GrepSeek — a Direct Corpus Interaction (DCI)
search agent that answers questions by issuing Unix shell commands (rg, grep,
head, …) directly over a 21M-passage Wikipedia corpus, instead of retrieving
from a dense/sparse index. This checkpoint is Qwen/Qwen3.5-9B supervised-fine-tuned
on the cold-start trajectories; it is the initialization for the RL stage and
corresponds to the "w/o GRPO" ablation in the paper.
📄 GrepSeek: Training Search Agents for Direct Corpus Interaction · 💻 https://github.com/alirezasalemi7/grepseek
- Base:
Qwen/Qwen3.5-9B - Training data:
alireza7/GrepSeek-ColdStart-SFT-10k(5k NQ + 5k HotpotQA, Answer-Aware-Tutor / Answer-Blind-Planner trajectories) - Recipe: verl FSDP, 1 epoch, AdamW, peak LR 5e-6 (constant + 5% warmup), 16,384-token sequences, bf16, Ulysses SP=4, on 4×A100-80GB.
- RL-optimized successor:
alireza7/GrepSeek-Qwen3.5-9B-GRPO
Why SFT first: RL directly from the base model over millions of documents is
unstable (overly broad queries, context blow-ups, OOM). Cold-start SFT instills
concise, causally-grounded shell-search behavior — establishing the low-level
retrieval "primitives" (fixed-string -F matching, | head truncation, cascaded
rg ... | rg ... filtering) that RL later refines.
⚠️ A tool-using agent, not a standalone chatbot
The model emits <tool_call> shell commands that must be executed against the
Wikipedia corpus and fed back as <tool_response> turns. To use it you need:
(1) the corpus PeterJinGo/wiki-18-corpus,
(2) a tool-calling vLLM server, and (3) the GrepSeek inference harness (grep tool
- agent loop), all in the code repo.
Usage
git clone https://github.com/alirezasalemi7/grepseek && cd grepseek
# env: TRAINING_ENV.md · corpus: cold_start_sft/download_corpus.py
# 1. serve this checkpoint
MODEL_PATH=alireza7/GrepSeek-Qwen3.5-9B-SFT bash rl/serve_rl.sh # -> http://localhost:10730/v1
# 2. run the agent (paper inference: temperature 0.6, <=6 turns, 16k context)
GREPSEEK_CORPUS_ROOT=/path/to/wiki_18_corpus \
bash inference/run_inference.sh --base_url http://localhost:10730/v1 \
--model grepseek --temperature 0.6 --input my_questions.jsonl --out_dir out
Evaluation (token-F1 / EM, micro-average over 7 QA benchmarks)
This SFT-only policy already substantially beats the untuned base model, but RL adds large gains on multi-hop reasoning:
| variant | micro-avg F1 | micro-avg EM |
|---|---|---|
| base (no SFT, no RL) | 0.3314 | 0.2836 |
| this model (SFT only) | 0.4249 | 0.3569 |
+ GRPO → GrepSeek-Qwen3.5-9B-GRPO |
0.5691 | 0.4948 |
(7 benchmarks: NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue, Bamboogle; trained only on NQ + HotpotQA, the rest are out-of-distribution.)
License
Inherits the license of the base model Qwen/Qwen3.5-9B — confirm and update the
license field above if needed.
Citation
@misc{salemi2026grepseektrainingsearchagents,
title={GrepSeek: Training Search Agents for Direct Corpus Interaction},
author={Alireza Salemi and Chang Zeng and Atharva Nijasure and Jui-Hui Chung and Razieh Rahimi and Fernando Diaz and Hamed Zamani},
year={2026},
eprint={2605.29307},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2605.29307},
}
- Downloads last month
- 45