Business Entity Resolution: LightGBM pair and meta models

Business Entity Resolution at Scale

These are two gradient-boosted tree models from our Amazon ML Challenge 2026 solution. The task is to link each business record in Source 1 to every record of the same real-world business in Sources 2 and 3: about 1.7M queries against about 10M noisy records from the US, India and France. The final system scored 0.9303 macro F0.5 on the public leaderboard.

Full pipeline (retrieval, feature generation, training, inference): https://github.com/sanjayrohith/business-entity-resolution Write-up notebook: https://www.kaggle.com/code/sanjayrohith/business-entity-resolution-writeup

Pipeline

Files

File Description
base_lightgbm.txt Pair classifier: 80 features, 600 trees, 31 leaves, learning rate 0.04. Threshold 0.642
meta_lightgbm.txt Score-context meta model: 41 base features + 14 per-query context features, 341 trees, 15 leaves. Threshold 0.70
base_model_selection.json Feature order, tuned thresholds, and tuning/confirmation metrics for LightGBM vs XGBoost vs logistic regression
meta_model_result.json Meta-model feature names, threshold curve, and baseline vs meta metrics
feature_importance.json LightGBM gain importance for the base model
inference.py Loads both models, builds the context features, and applies the thresholds and one-owner-per-target decoding
er_demo.py, core.py Synthetic end-to-end demo: retrieval, the exact 80 features, both models, and ownership decoding on made-up records
images/ Charts used in this card

How it works

  1. Multi-route sparse retrieval returns candidate target records per query. It combines hashed TF-IDF on names and addresses, exact/suffix/address-key joins, transliteration and missing-address rescue routes.
  2. Each (query, candidate) pair gets 80 features: name/address string similarity, numeric and postal agreement or conflict, transliteration, frequency, missingness, and route ranks/scores.
  3. The base model scores each pair.
  4. The meta model re-scores each pair using its base score plus the query's score context: rank, gap to the best candidate, counts above several thresholds, and the other source's best score.
  5. Pairs with a meta score ≥ 0.70 become links. If several queries claim the same target, the target is kept only for an owner that leads by ≥ 0.05 on meta score.
from inference import load_models, score_pairs, resolve_ownership, META_THRESHOLD

base_model, meta_model, feature_names = load_models(".")
# X80: (n_pairs, 80) float32 in feature_names order; q: query index per pair; source: 2 or 3 per pair
base, meta = score_pairs(X80, q, source, base_model, meta_model)
links = [(s1_ids[i], target_ids[i], meta[i]) for i in range(len(meta)) if meta[i] >= META_THRESHOLD]
predictions = resolve_ownership(links)   # {s1_id: [target_id, ...]}

Try it: synthetic end-to-end demo

The challenge data can't be shared, so er_demo.py runs the full pipeline at small scale on made-up businesses. It uses 8 designed hard-case queries and about 700 generated filler records, with the same R10 retrieval routes, the exact 80 production features, and both trained models.

pip install lightgbm scikit-learn scipy rapidfuzz anyascii
python er_demo.py

Synthetic demo

The demo shows abbreviated legal suffixes, Devanagari script, missing addresses and typos being matched. It also shows same-name businesses in other cities, co-located businesses and no-match queries being rejected. All 8 synthetic queries are resolved correctly. The synthetic cases are easier than the real data, so read this as an illustration, not a benchmark.

Important: the models expect the exact 80 features produced by the repository's pair_features.py, including retrieval route ranks and scores. They cannot score raw name/address strings by themselves.

Evaluation

The metric is macro-averaged per-query F0.5 (precision-weighted). A query with no true matches and no predictions scores 1.

Stage Held-out macro F0.5 Precision Recall
Base model, top-100 candidates (2,500 confirmation queries) 0.9453 0.9658 0.9054
Base model, fast R10 candidates (4,000 fresh queries) 0.9363 0.9665 0.8733
+ meta model (same 4,000 queries) 0.9438 0.9654 0.8986
Submission Public leaderboard F0.5
Base model 0.9176
+ meta model 0.9217
+ target-ownership decoding 0.9303

Leaderboard progression

Model comparison

Feature importance

Slice analysis

Confirmation queries were grouped by normalized name/address keys so that near-duplicates could not cross folds. All thresholds were chosen on a separate tuning fold before confirmation scoring.

Training data and limitations

  • Trained from scratch only on the official challenge training labels. The base model used 5,001 grouped fitting entities (1.09M sampled candidate pairs, including hard negatives); the meta model used 4,045 separate entities. No pretrained models or external data were used. The dataset is not redistributed here.
  • There are no French training labels, so performance on France is unmeasured.
  • The hardest slices are cross-script names (Indic script vs Latin) and records with missing addresses.
  • The models are specific to this dataset's schema and feature pipeline and are not a general-purpose entity matcher.

License

MIT

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Public leaderboard macro F0.5 on Amazon ML Challenge 2026 - Business Entity Resolution
    self-reported
    0.930
  • Local held-out macro F0.5 (meta model, R10 candidates) on Amazon ML Challenge 2026 - Business Entity Resolution
    self-reported
    0.944