granite-embedding-278m-multilingual-burnpack

Original model: https://huggingface.co/ibm-granite/granite-embedding-278m-multilingual Original authors: the Granite Embedding Team, IBM Research Converted by: Lucie666, using burn-onnx โ€” format only


This is not an original model, and no part of it is my work. It is a mechanical format conversion of ibm-granite/granite-embedding-278m-multilingual โ€” nothing was trained, fine-tuned, distilled, quantised or modified. No new weights were produced. All credit belongs to IBM Research; the model is released under Apache-2.0, and so is this conversion. For the paper, the benchmarks and the citation, see the upstream card.

If you are looking for the model itself, go to the upstream repository. This one only exists so people running Burn don't each redo the conversion.

The file model.bpk holds the same weights as the upstream model.onnx, re-serialised into Burn's burnpack format so they can be loaded by a pure-Rust inference stack โ€” no Python, no PyTorch, no ONNX Runtime at inference time.

The model, in one line

XLM-RoBERTa encoder, 12 layers, hidden size 768, 768-dimensional embeddings, 512 tokens, trained on text and code in twelve languages. The embedding is last_hidden_state[:, 0] (the CLS token) followed by an L2 normalisation, as on the upstream card. There are no token_type_ids.

Provenance

ibm-granite/granite-embedding-278m-multilingual   model.onnx
        (1 112 413 925 bytes, sha256 aefac97b384f92932a61a19900d41c870679d5b8e6ceb682768eb153d0e31c7d)
        โ”‚
        โ”‚  burn-onnx 0.22.0-pre.3   (mechanical ONNX โ†’ Burn conversion, LoadStrategy::Bytes)
        โ–ผ
model.bpk        weights, burnpack format
model.rs         model graph, generated Rust source (not distributed here)

Nothing in this pipeline is hand-written. The generated graph lives in rag3weaver as generated/granite_278m_onnx.rs, where it is used as the default code-and-text embedder.

Files

model.bpk        1 112 227 840 bytes
sha256           a54628b51156caa158a4ede09f1c4d1577e4f94be9229e8455bef4787fa342d3
tokenizer.json   9 081 351 bytes   (unchanged from upstream)
sha256           2a0d7366dd7780ea36cc42431dd74cd79289b783ab01acd33013fcc96865a8e9

A caveat before you regenerate. Burnpack serialisation is not byte-deterministic: two builds from the same ONNX produce files of identical size but different bytes. The tensor values are unaffected. The checksum above verifies this download, not a rebuild.

Verification

Converted on 6 September 2026. Checked on 3 October 2026 against the upstream ONNX on a fixed set of 24 sentences (French, English, source code, four other languages; from 4 tokens to texts truncated at 512), one sentence per forward pass, comparing the L2-normalised CLS vectors element by element:

reference this file, loaded by Burn max absolute difference
onnxruntime 1.30.0, CPU, f32 Burn 0.22.0-pre.3 on Vulkan (AMD Radeon 8060S, radv), Flex32 (f32 storage, f16 matmul, f32 accumulation) 4.9e-5

The vectors are unit-norm, so their components are of the order of a few hundredths. A rebuild from the same ONNX on 2 October 2026 measured 3e-7 in plain f32 on the same device, which is rounding noise; the Flex32 figure is the cost of the half-precision matmul, not of the conversion.

Loading

Generate model.rs from the upstream model.onnx with burn-onnx 0.22.0-pre.3 (ModelGen with LoadStrategy::Bytes), load model.bpk into the generated model, and call forward(input_ids, attention_mask). It returns (last_hidden_state, pooler_output): the embedding is the L2-normalised last_hidden_state[:, 0]; the second output is Hugging Face's pooler_output, which the upstream card does not use.

Tokenise with tokenizer.json (XLM-RoBERTa, pad id 1) and truncate at 512 tokens.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Lucie666/granite-embedding-278m-multilingual-burnpack

Finetuned
(15)
this model