Instructions to use srpone/zoowork-shopranker-8b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use srpone/zoowork-shopranker-8b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("srpone/zoowork-shopranker-8b") model = PeftModel.from_pretrained(base_model, "srpone/zoowork-shopranker-8b") - Notebooks
- Google Colab
- Kaggle
ZooWork-ShopRanker-8B
An e-commerce reranker aligned to judged shopper preference rather than topical
relevance, from the paper ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce
Reranker. This repository is self-contained: it holds the
unmodified Qwen/Qwen3-Reranker-8B weights together with our LoRA
adapter, so everything loads from this one repo id -- no separate base download. Same scoring
interface as the base: one forward pass per (query, product), works at any k, and scores
are cacheable per pair.
🚀 An optimized commercial version of the ZooWork-ShopRanker rerankers is available through the zoodata.ai API.
Results on ShopRank-Bench
Pairwise accuracy in percent on ShopRank-Bench (10,511 judged preference pairs from 2,991 queries); intervals are 95% query-clustered bootstrap CIs. Tiers are the number of judge families (of three) that committed to the label.
| Model | Product text | Overall [95% CI] | Gold (3/3) | Silver (2/3) | Bronze (1/3) |
|---|---|---|---|---|---|
| ZooWork-ShopRanker-8B | Structured | 83.4 [82.5, 84.2] | 96.5 | 86.1 | 74.8 |
Qwen3-Reranker-8B (base) |
Structured | 79.2 [78.2, 80.1] | 93.2 | 82.6 | 69.4 |
| ZooWork-ShopRanker-8B | Natural language | 81.8 [80.9, 82.7] | 95.0 | 85.4 | 72.2 |
Qwen3-Reranker-8B (base) |
Natural language | 78.4 [77.5, 79.4] | 91.9 | 82.3 | 68.4 |
The flagship and the distillation teacher for the 4B and 0.6B. Best reranker overall in both formats and on every tier of the structured format; +4.2 over its own base on structured text.
Report per tier: every reranker loses 20–25 points from gold to bronze, so an aggregate score is dominated by pairs nothing gets wrong.
Serving cost
One NVIDIA H200, fp16, 1024-token limit, 256 real (query, product) pairs, adapter merged into the weights (LoRA adds no inference cost once merged).
| Params | Batch-1 p50 | Batch-1 p95 | Docs/s (batch 32) | Peak memory |
|---|---|---|---|---|
| 8.19B | 25.8 ms | 30.4 ms | 102.7 | 18.9 GB |
Usage
Requires pip install transformers peft. Load the base weights and the adapter from this
repo as shown: wrapping with PeftModel keeps the adapter in fp32, which reproduces the paper
numbers exactly (loading the repo id with from_pretrained alone also attaches the adapter,
but in bf16, and drifts by up to ~0.3 points).
The score is P(yes) = softmax over the bare no / yes logits at the last position of
the native Qwen3-Reranker prompt -- exactly the scorer behind the numbers above.
import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "srpone/zoowork-shopranker-8b" # base weights + LoRA adapter, both in this repo
tok = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16)
# Re-wrapping keeps the adapter in fp32, as in the paper. peft may warn "Already found a
# `peft_config` attribute"; that is expected here -- exactly one adapter ends up loaded.
model = PeftModel.from_pretrained(model, model_id).cuda().eval()
PREFIX = ("<|im_start|>system\nJudge whether the Document meets the requirements based on "
"the Query and the Instruct provided. Note that the answer can only be \"yes\" "
"or \"no\".<|im_end|>\n<|im_start|>user\n")
SUFFIX = "<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n"
INSTRUCTION = "Given a web search query, retrieve relevant passages that answer the query"
YES, NO = tok.convert_tokens_to_ids("yes"), tok.convert_tokens_to_ids("no")
@torch.no_grad()
def score(query: str, products: list[str]) -> list[float]:
texts = [PREFIX + f"<Instruct>: {INSTRUCTION}\n<Query>: {query}\n<Document>: {p}" + SUFFIX
for p in products]
enc = tok(texts, return_tensors="pt", padding=True, truncation=True,
max_length=4096, add_special_tokens=False).to(model.device)
logits = model(**enc).logits[:, -1, :]
return torch.softmax(logits[:, [NO, YES]].float(), dim=-1)[:, 1].tolist()
print(score("jumpsuit under $100 boho", [
"title: Boho Wide-Leg Jumpsuit | price: $44 | color: rust",
"title: Bohemian Silk Jumpsuit | price: $239 | color: ivory",
]))
Product text may be the structured attribute schema or natural-language prose. The model was trained on the structured schema and carries its gain to prose (see the table).
Repository contents. config.json, model*.safetensors and the tokenizer files are
the unmodified Qwen3-Reranker-8B release; adapter_config.json and
adapter_model.safetensors are the ZooWork-ShopRanker LoRA adapter. The adapter is kept
separate rather than merged so the published model reproduces the paper numbers exactly.
If your serving stack needs a single set of weights, merging the adapter into the base in
bf16 stays within 0.4 points of the reported accuracy.
Training
LoRA (rank 16, α=32, dropout 0.05, on q/k/v/o projections) over judge-labelled preference pairs from private search traffic, labelled by a three-family LLM judge panel in both presentation orders, mixed with general retrieval replay data. This model was aligned directly on judged preference pairs. The benchmark comes from a disjoint slice of traffic (zero query and zero exact-pair overlap).
Limitations
- Labels come from an LLM judge panel and are not human-verified. No click, purchase or interleaving signal validates that the panel tracks what shoppers actually choose.
- Attribute hierarchy is not installed. On the AHP diagnostic, alignment lifts the
aggregate by transforming
priceandcolorwhile degradingstyleandproduct type; treat the gain as a side effect, not hierarchy-following. - Explicit price constraints are largely ignored by open rerankers, this one included.
- Zero-shot reasoning LLMs remain 8+ points more accurate overall, at far higher latency.
- The category mix leans toward apparel; evaluated in English only.
License and attribution
Apache License 2.0 (see LICENSE), inherited from Qwen/Qwen3-Reranker-8B
(Apache-2.0, by the Qwen team). This repository redistributes the base weights
and tokenizer unmodified; the ZooWork-ShopRanker modification is the LoRA adapter,
trained on judged preference data.
Citation
@article{xue2026zooworkshopranker,
title = {ZooWork-ShopRanker: An Open, Preference-Aligned E-Commerce Reranker},
author = {Xue, Siqiao and Liu, Shuxuan and Hu, Ning},
journal = {arXiv preprint arXiv:2609.31002},
year = {2026},
url = {https://arxiv.org/abs/2609.31002}
}
- Downloads last month
- 54