bge-m3 ONNX (opset 18, fused LayerNormalization, last_hidden_state output)
ONNX export of the dense encoder of BAAI/bge-m3 that
produces the same embeddings as the original model and stays correct under bf16 inference,
unlike the onnx/ folder shipped in the upstream repo.
Why this export exists
The upstream onnx/model.onnx was exported with PyTorch 2.1 at opset 11: every LayerNorm is
decomposed into ReduceMean / Sub / Pow / Sqrt / Div. In fp32 that is harmless. With
INFERENCE_PRECISION_HINT=bf16 on the OpenVINO CPU plugin (the usual way to double CPU
throughput on AVX-512 BF16 / AMX hardware), the variance of the decomposed LayerNorm is computed
in bf16 and XLM-RoBERTa's large residual outliers turn the output into noise: cosine 0.22-0.40
to the fp32 embeddings, for both the token_embeddings and the in-graph sentence_embedding
outputs. A graph with the fused LayerNormalization op (opset 17+) does not have the problem, the
backend keeps the statistics in fp32 inside the kernel.
This export (torch.onnx.export(dynamo=True), PyTorch 2.14, opset 18, eager attention, fp32
weights) measured against the upstream ONNX export in fp32 on 5 texts (CLS pooled, L2 normalized):
| runtime | cosine to upstream fp32 |
|---|---|
| onnxruntime CPUExecutionProvider, fp32 | 1.00000 |
| onnxruntime OpenVINOExecutionProvider, fp32 | 1.00000 |
| onnxruntime OpenVINOExecutionProvider, bf16 | 0.99998-0.99999 |
upstream onnx/model.onnx, OpenVINO bf16 |
0.22-0.40 |
Files
model.onnx+model.onnx.data(2.27 GB external weights) — encoder, fp32, opset 18, dynamic batch/sequence axes; inputsinput_ids,attention_mask; outputlast_hidden_state[batch, seq, 1024]- tokenizer and config files of the base model
export.py— the export script
Only the dense embedding is exported. Sparse (lexical) and ColBERT (multi-vector) heads of bge-m3 are not part of this graph.
Usage with infinity
bge-m3 uses CLS pooling and L2 normalization, so:
infinity_emb v2 --model-id ElXreno/bge-m3-onnx --engine optimum --pooling-method cls
bf16 on the OpenVINO CPU plugin (needs the -cpu image, onnxruntime-openvino):
infinity_emb v2 --model-id ElXreno/bge-m3-onnx --engine optimum --pooling-method cls \
--onnx-provider-options '{"load_config": "{\"CPU\": {\"INFERENCE_PRECISION_HINT\": \"bf16\"}}"}'
Usage with onnxruntime directly
import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("ElXreno/bge-m3-onnx")
sess = ort.InferenceSession("model.onnx")
enc = tok(["пример текста", "example text"], padding=True, return_tensors="np")
out = sess.run(["last_hidden_state"], {k: enc[k].astype(np.int64) for k in ("input_ids", "attention_mask")})[0]
emb = out[:, 0, :]
emb /= np.linalg.norm(emb, axis=1, keepdims=True)
License
MIT, same as BAAI/bge-m3.
- Downloads last month
- 37
Model tree for ElXreno/bge-m3-onnx
Base model
BAAI/bge-m3