bge-m3 ONNX (opset 18, fused LayerNormalization, last_hidden_state output)

ONNX export of the dense encoder of BAAI/bge-m3 that produces the same embeddings as the original model and stays correct under bf16 inference, unlike the onnx/ folder shipped in the upstream repo.

Why this export exists

The upstream onnx/model.onnx was exported with PyTorch 2.1 at opset 11: every LayerNorm is decomposed into ReduceMean / Sub / Pow / Sqrt / Div. In fp32 that is harmless. With INFERENCE_PRECISION_HINT=bf16 on the OpenVINO CPU plugin (the usual way to double CPU throughput on AVX-512 BF16 / AMX hardware), the variance of the decomposed LayerNorm is computed in bf16 and XLM-RoBERTa's large residual outliers turn the output into noise: cosine 0.22-0.40 to the fp32 embeddings, for both the token_embeddings and the in-graph sentence_embedding outputs. A graph with the fused LayerNormalization op (opset 17+) does not have the problem, the backend keeps the statistics in fp32 inside the kernel.

This export (torch.onnx.export(dynamo=True), PyTorch 2.14, opset 18, eager attention, fp32 weights) measured against the upstream ONNX export in fp32 on 5 texts (CLS pooled, L2 normalized):

runtime cosine to upstream fp32
onnxruntime CPUExecutionProvider, fp32 1.00000
onnxruntime OpenVINOExecutionProvider, fp32 1.00000
onnxruntime OpenVINOExecutionProvider, bf16 0.99998-0.99999
upstream onnx/model.onnx, OpenVINO bf16 0.22-0.40

Files

  • model.onnx + model.onnx.data (2.27 GB external weights) — encoder, fp32, opset 18, dynamic batch/sequence axes; inputs input_ids, attention_mask; output last_hidden_state [batch, seq, 1024]
  • tokenizer and config files of the base model
  • export.py — the export script

Only the dense embedding is exported. Sparse (lexical) and ColBERT (multi-vector) heads of bge-m3 are not part of this graph.

Usage with infinity

bge-m3 uses CLS pooling and L2 normalization, so:

infinity_emb v2 --model-id ElXreno/bge-m3-onnx --engine optimum --pooling-method cls

bf16 on the OpenVINO CPU plugin (needs the -cpu image, onnxruntime-openvino):

infinity_emb v2 --model-id ElXreno/bge-m3-onnx --engine optimum --pooling-method cls \
  --onnx-provider-options '{"load_config": "{\"CPU\": {\"INFERENCE_PRECISION_HINT\": \"bf16\"}}"}'

Usage with onnxruntime directly

import numpy as np, onnxruntime as ort
from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("ElXreno/bge-m3-onnx")
sess = ort.InferenceSession("model.onnx")
enc = tok(["пример текста", "example text"], padding=True, return_tensors="np")
out = sess.run(["last_hidden_state"], {k: enc[k].astype(np.int64) for k in ("input_ids", "attention_mask")})[0]
emb = out[:, 0, :]
emb /= np.linalg.norm(emb, axis=1, keepdims=True)

License

MIT, same as BAAI/bge-m3.

Downloads last month
37
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ElXreno/bge-m3-onnx

Base model

BAAI/bge-m3
Quantized
(301)
this model