granite-embedding-278m-multilingual-burnpack
Original model: https://huggingface.co/ibm-granite/granite-embedding-278m-multilingual
Original authors: the Granite Embedding Team, IBM Research
Converted by: Lucie666, using burn-onnx โ format only
This is not an original model, and no part of it is my work. It is a mechanical format conversion of ibm-granite/granite-embedding-278m-multilingual โ nothing was trained, fine-tuned, distilled, quantised or modified. No new weights were produced. All credit belongs to IBM Research; the model is released under Apache-2.0, and so is this conversion. For the paper, the benchmarks and the citation, see the upstream card.
If you are looking for the model itself, go to the upstream repository. This one only exists so people running Burn don't each redo the conversion.
The file model.bpk holds the same weights as the upstream model.onnx, re-serialised
into Burn's burnpack format so they can be loaded by a pure-Rust inference stack โ no
Python, no PyTorch, no ONNX Runtime at inference time.
The model, in one line
XLM-RoBERTa encoder, 12 layers, hidden size 768, 768-dimensional embeddings, 512 tokens, trained on text and code in twelve languages.
The embedding is last_hidden_state[:, 0] (the CLS token) followed by an L2 normalisation,
as on the upstream card. There are no token_type_ids.
Provenance
ibm-granite/granite-embedding-278m-multilingual model.onnx
(1 112 413 925 bytes, sha256 aefac97b384f92932a61a19900d41c870679d5b8e6ceb682768eb153d0e31c7d)
โ
โ burn-onnx 0.22.0-pre.3 (mechanical ONNX โ Burn conversion, LoadStrategy::Bytes)
โผ
model.bpk weights, burnpack format
model.rs model graph, generated Rust source (not distributed here)
Nothing in this pipeline is hand-written. The generated graph lives in
rag3weaver as generated/granite_278m_onnx.rs, where it is
used as the default code-and-text embedder.
Files
model.bpk 1 112 227 840 bytes
sha256 a54628b51156caa158a4ede09f1c4d1577e4f94be9229e8455bef4787fa342d3
tokenizer.json 9 081 351 bytes (unchanged from upstream)
sha256 2a0d7366dd7780ea36cc42431dd74cd79289b783ab01acd33013fcc96865a8e9
A caveat before you regenerate. Burnpack serialisation is not byte-deterministic: two builds from the same ONNX produce files of identical size but different bytes. The tensor values are unaffected. The checksum above verifies this download, not a rebuild.
Verification
Converted on 6 September 2026. Checked on 3 October 2026 against the upstream ONNX on a fixed set of 24 sentences (French, English, source code, four other languages; from 4 tokens to texts truncated at 512), one sentence per forward pass, comparing the L2-normalised CLS vectors element by element:
| reference | this file, loaded by Burn | max absolute difference |
|---|---|---|
| onnxruntime 1.30.0, CPU, f32 | Burn 0.22.0-pre.3 on Vulkan (AMD Radeon 8060S, radv), Flex32 (f32 storage, f16 matmul, f32 accumulation) | 4.9e-5 |
The vectors are unit-norm, so their components are of the order of a few hundredths. A rebuild from the same ONNX on 2 October 2026 measured 3e-7 in plain f32 on the same device, which is rounding noise; the Flex32 figure is the cost of the half-precision matmul, not of the conversion.
Loading
Generate model.rs from the upstream model.onnx with burn-onnx 0.22.0-pre.3 (ModelGen
with LoadStrategy::Bytes), load model.bpk into the generated model, and call
forward(input_ids, attention_mask). It returns (last_hidden_state, pooler_output): the
embedding is the L2-normalised last_hidden_state[:, 0]; the second output is Hugging
Face's pooler_output, which the upstream card does not use.
Tokenise with tokenizer.json (XLM-RoBERTa, pad id 1) and truncate at 512 tokens.