SETA-SFT: Qwen3-8B Fine-tuned with SFT on Kimi-K2.5 Trajectories

This model is Qwen3-8B fine-tuned with supervised fine-tuning (SFT) on trajectories collected from a strong teacher model (Kimi-K2.5) on SETA terminal-agent tasks.

It is a checkpoint released alongside the paper SETA: Scaling Environments for Terminal Agents (anonymous submission under double-blind review).

Model Details

Base model Qwen/Qwen3-8B
Training method Supervised Fine-Tuning (SFT)
Teacher model Kimi-K2.5
Training data 1 488 terminal-agent trajectories collected on SETA tasks
Trainable tokens ~7.79 M (thinking variant)
Context length 32 768 tokens
Thinking Enabled (trajectories preserve <think>…</think> reasoning blocks)

Intended Use

This model is designed for terminal agent tasks: completing multi-step shell-based tasks inside a Docker container environment using tools such as shell_exec, shell_view, and shell_write_content_to_file.

It serves as an SFT warm-start baseline and can be further improved with reinforcement learning.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "AnonymousSubmissionUnderDouble-BlindRevi/seta-sft-kimi-qwen3"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

For evaluation with the SETA framework, serve via SGLang and run:

python -m sglang.launch_server --model AnonymousSubmissionUnderDouble-BlindRevi/seta-sft-kimi-qwen3 --port 30000

python scripts/evaluation/eval.py \
    --config scripts/evaluation/configs/eval_default_qwen3_8b.yaml \
    terminal_env.model.model_type=AnonymousSubmissionUnderDouble-BlindRevi/seta-sft-kimi-qwen3 \
    terminal_env.model.url=http://localhost:30000/v1 \
    dataset=seta-env

Training Configuration

Key hyperparameters (see scripts/areal_sft/configs/seta_kimi_qwen3_sft_thinking.yaml in the companion code repository):

Hyperparameter Value
Learning rate 2e-5
LR scheduler cosine
Optimizer AdamW (β₁=0.9, β₂=0.95)
Weight decay 0.05
Total epochs 3
Batch size 8
Max sequence length 16 384 tokens
GPUs 2 × (d2p1t1)

SFT Data Summary

Trajectories were collected by running Kimi-K2.5 on SETA tasks and filtering for verified reward:

Thinking variant
Total trajectories 1 488
Fully-passing trajectories 1 112
Mean reward 0.932
Total tokens 15.69 M
Trainable tokens 7.79 M (49.6%)

Limitations

  • Evaluated on terminal-agent benchmarks; performance on general language tasks is not characterized.
  • SFT trajectories were collected from a single teacher model; diversity may be limited.

License

Apache 2.0 (inherited from Qwen3-8B base).

Downloads last month
15
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AnonymousSubmissionUnderDouble-BlindRevi/seta-sft-kimi-qwen3

Finetuned
Qwen/Qwen3-8B
Finetuned
(2158)
this model