LFM2.5-2.6B-GGUF / qad /README.md
adityatadimeti's picture
Add full-precision QAD safetensors source checkpoint
e7caca5 verified
|
Raw History Blame Contribute Delete
3.1 kB
metadata
library_name: transformers
pipeline_tag: text-generation
license: other
license_name: lfm1.0
license_link: LICENSE
base_model: LiquidAI/LFM2.5-2.6B
tags:
  - liquid
  - lfm2.5
  - qad
  - safetensors

LFM2.5-2.6B QAD — FP32 source checkpoint

This directory contains the unquantized FP32 safetensors source weights for LFM2.5-2.6B-QAD-Q4_0.gguf. These are the weights after Quantization-Aware Distillation (QAD), before GGUF conversion and quantization. The architecture and tokenizer are those of LiquidAI/LFM2.5-2.6B.

The checkpoint is provided for fine-tuning, inspecting weights, and experimenting with deployment formats. It was optimized for Q4_0 deployment; published QAD Q4_0 benchmark results do not describe direct FP32/BF16 inference or other quantization formats. See the QAD release.

Usage

Install current torch, transformers, and accelerate. This package was checked with Transformers 5.9.0. dtype="auto" preserves the stored FP32 weights.

from transformers import AutoModelForCausalLM, AutoTokenizer

repo_id = "LiquidAI/LFM2.5-2.6B-GGUF"
tokenizer = AutoTokenizer.from_pretrained(repo_id, subfolder="qad")
model = AutoModelForCausalLM.from_pretrained(
    repo_id, subfolder="qad", dtype="auto", device_map="auto"
)

inputs = tokenizer.apply_chat_template(
    [{"role": "user", "content": "What is 2 + 2?"}],
    tokenize=True, add_generation_prompt=True, return_dict=True, return_tensors="pt",
).to(model.device)
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0, inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

Files and validation

  • model.safetensors: original FP32 QAD weights.
  • config.json: the validated source architecture/configuration, including its FP32 dtype.
  • tokenizer.json, tokenizer_config.json: tokenizer files.
  • chat_template.jinja: exactly the template embedded in the public QAD GGUF.
  • generation_config.json: defaults from the corresponding public HF model.
  • LICENSE: the repository's LFM license.
  • provenance.json: source hash and the validated GGUF reference.

The FP32 weights passed finite-value checks and Transformers loading, generation, and cache-consistency smoke tests. Converting these weights to F16 GGUF and then Q4_0 with llama.cpp revision 74ade52741203e5c8f81eaf06a96cb1cfe15f2a3 reproduced every tensor byte of the released QAD artifact. Applying the same release metadata also reproduced its complete SHA-256. File names and descriptive GGUF metadata can otherwise cause different whole-file hashes despite identical tensors.

Preserve FP32 weights when reproducing that conversion: saving a BF16 copy first changes the source precision. Generation defaults follow the public model; use do_sample=False for deterministic greedy decoding.

License

This checkpoint is distributed under the same LFM license as the corresponding model release.