community-airline-voice-outcome-filter-1.7b

Recipe: recipes/community/airline-voice-concise-under-probe-outcome-filter · Collection: Course and community runs

The filter-metric paper's change, moved from GSM8K to an airline agent's register. GRPO on the airline-voice-concise rows with a shaped reward; the baseline drops all-equal groups by shaped score, the method by binary outcome. The eval/ folder holds the eval rows.

Result, from the recipe README

comparison target delta verdict
baseline vs base +0.029 [-0.004, +0.063] flat (noise < 0.142)
method vs base +0.023 [-0.011, +0.058] flat, and over-optimized
method vs baseline -0.005 [-0.041, +0.029] flat, no difference

The method dropped 77.5% of groups against the baseline's 6.2%, moved covered_all +0.090 and the training reward +0.064, and did not move the target. On the probed rows the target went -0.058 [-0.108, -0.017]. About 73 GPU minutes on H100s.

Arms in this repo

The root holds the arm the recipe README's headline number reports. Every other arm is a subfolder named after it. checkpoints/ never ships.

folder arm
. method: flat-group test on the binary outcome
baseline baseline: flat-group test on the shaped score
eval eval rows for base, baseline and method

Load

from peft import PeftModel
from transformers import AutoModelForCausalLM

base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-1.7B")
model = PeftModel.from_pretrained(base, "while-ai/community-airline-voice-outcome-filter-1.7b")  # the headline arm
model = PeftModel.from_pretrained(base, "while-ai/community-airline-voice-outcome-filter-1.7b", subfolder="baseline")  # another arm

Reproduce

git clone https://github.com/whilehq/whileai-sdk && cd whileai-sdk/recipes/community/airline-voice-concise-under-probe-outcome-filter
python run.py

The recipe README pins the seed, the library versions and the GPU, and its Checks table says what the eval verified. Read the Learned section before quoting a number from this card.

Downloads last month
32
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for while-ai/community-airline-voice-outcome-filter-1.7b

Finetuned
Qwen/Qwen3-1.7B
Adapter
(763)
this model

Dataset used to train while-ai/community-airline-voice-outcome-filter-1.7b

Collections including while-ai/community-airline-voice-outcome-filter-1.7b