Qwen3-1.7B — Direct-OPD (SFT-induced teacher shift), 100 steps

Pilot artifact of the Direct-OPD SFT-vs-RL policy-shift experiment (condition sft). Student Qwen/Qwen3-1.7B trained for 100 Direct-OPD steps against the token-level policy-shift signal between:

role model
pi_T (post-shift teacher) cmpatino/DeepSeek-R1-Distill-Qwen-1.5B-DeepMath-SFT100 @ baee02cc3858ddb200601f247b6a22f072b599d1
pi_Tref (pre-shift teacher) deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B @ ad9f0ae0864d7fbcd1cd905e3c6c5b069cc8b562

Root = step 100. checkpoint-{20,40,60,80,100}/ = intermediate merged checkpoints. Weights are bf16 (verl's FSDP->HF merge downcasts the fp32 masters).

Reproduction

  • code: https://github.com/BytedTsinghua-SIA/Direct-OPD @ 3a9d6bd37b00a38e7a9b2959239e4631e5324aea + logs/phase4_seed.patch (seed 42 shim)
  • data: cmpatino/direct-opd-sft-deepmath-pilot-data @ 22625ae5db434947195bf862c429cd94504a4809 :: opd_train.parquet (6,400 prompts, one pass)
  • 100 steps x 64 prompts x 4 rollouts, lr 1e-6, adaptive KL (init/max 2.5), max prompt 1024 / response 2048 tokens, seed 42
  • driver + full env block: logs/run_manifest.json, console log logs/train.log.gz

Caveats: no in-training validation (test_freq=-1); bit-exact reproducibility is not attainable (vLLM continuous batching, dynamic micro-batching, FSDP reductions); teacher scoring reuses the student's Qwen3 token ids verbatim (shared-vocab assumption, unasserted upstream).

Downloads last month
1
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cmpatino/Qwen3-1.7B-DirectOPD-SFTShift-100

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(975)
this model