Pradheep1647/kibitzer-sft-elo-0808-step-002000

Kibitzer checkpoint for a chess policy/value model trained on game sequences. Each board is encoded as 64 square tokens plus auxiliary state, summarized by a square-level encoder, then processed by a causal transformer over the game timeline.

Status

  • checkpoint: step_002000.pt
  • step: 002000
  • elo_rating: 808 estimated vs stockfish-elo-1320 (-511.5 Elo diff)
  • github_repo: https://github.com/Mantissagithub/kibitzer
  • eval_opponent: stockfish-elo-1320
  • elo_diff: -511.5
  • eval_games: 20
  • eval_score: 1.00
  • elo_error: n/a

Architecture

  • Input: 64 chess-square piece tokens plus 7 auxiliary scalars covering side to move, castling rights, en-passant file, and halfmove clock.
  • Position encoder: bidirectional square-level transformer that compresses each board into one timeline token.
  • Timeline trunk: causal transformer with d_model=384, sequence length 256, 12 layers, and 8 attention heads.
  • Transformer blocks: pre-norm RMSNorm, RoPE on causal Q/K attention, PyTorch scaled dot-product attention, and SwiGLU feed-forward layers.
  • Outputs: policy head over 4,672 AlphaZero-style moves (64 * 73) plus a tanh value head for bounded outcome prediction.

Files

  • step_002000.pt: PyTorch checkpoint.
  • training_metadata.yaml: training config and checkpoint metrics.
  • post_eval.yaml: uploaded after local Stockfish/cutechess eval.

Usage

git clone https://github.com/Mantissagithub/kibitzer.git
cd kibitzer
uv run python scripts/uci.py --checkpoint <path-to>/step_002000.pt

This checkpoint is one artifact in a sequence. Repos with elo-pending in the name have not been evaluated yet; rated repos are renamed after scripts/eval_and_rename_hf.py --from-hf completes.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Pradheep1647/kibitzer-sft-elo-0808-step-002000