Business Entity Resolution: LightGBM pair and meta models
These are two gradient-boosted tree models from our Amazon ML Challenge 2026 solution. The task is to link each business record in Source 1 to every record of the same real-world business in Sources 2 and 3: about 1.7M queries against about 10M noisy records from the US, India and France. The final system scored 0.9303 macro F0.5 on the public leaderboard.
Full pipeline (retrieval, feature generation, training, inference): https://github.com/sanjayrohith/business-entity-resolution Write-up notebook: https://www.kaggle.com/code/sanjayrohith/business-entity-resolution-writeup
Files
| File | Description |
|---|---|
base_lightgbm.txt |
Pair classifier: 80 features, 600 trees, 31 leaves, learning rate 0.04. Threshold 0.642 |
meta_lightgbm.txt |
Score-context meta model: 41 base features + 14 per-query context features, 341 trees, 15 leaves. Threshold 0.70 |
base_model_selection.json |
Feature order, tuned thresholds, and tuning/confirmation metrics for LightGBM vs XGBoost vs logistic regression |
meta_model_result.json |
Meta-model feature names, threshold curve, and baseline vs meta metrics |
feature_importance.json |
LightGBM gain importance for the base model |
inference.py |
Loads both models, builds the context features, and applies the thresholds and one-owner-per-target decoding |
er_demo.py, core.py |
Synthetic end-to-end demo: retrieval, the exact 80 features, both models, and ownership decoding on made-up records |
images/ |
Charts used in this card |
How it works
- Multi-route sparse retrieval returns candidate target records per query. It combines hashed TF-IDF on names and addresses, exact/suffix/address-key joins, transliteration and missing-address rescue routes.
- Each (query, candidate) pair gets 80 features: name/address string similarity, numeric and postal agreement or conflict, transliteration, frequency, missingness, and route ranks/scores.
- The base model scores each pair.
- The meta model re-scores each pair using its base score plus the query's score context: rank, gap to the best candidate, counts above several thresholds, and the other source's best score.
- Pairs with a meta score ≥ 0.70 become links. If several queries claim the same target, the target is kept only for an owner that leads by ≥ 0.05 on meta score.
from inference import load_models, score_pairs, resolve_ownership, META_THRESHOLD
base_model, meta_model, feature_names = load_models(".")
# X80: (n_pairs, 80) float32 in feature_names order; q: query index per pair; source: 2 or 3 per pair
base, meta = score_pairs(X80, q, source, base_model, meta_model)
links = [(s1_ids[i], target_ids[i], meta[i]) for i in range(len(meta)) if meta[i] >= META_THRESHOLD]
predictions = resolve_ownership(links) # {s1_id: [target_id, ...]}
Try it: synthetic end-to-end demo
The challenge data can't be shared, so er_demo.py runs the full pipeline at small scale on made-up businesses. It uses 8 designed hard-case queries and about 700 generated filler records, with the same R10 retrieval routes, the exact 80 production features, and both trained models.
pip install lightgbm scikit-learn scipy rapidfuzz anyascii
python er_demo.py
The demo shows abbreviated legal suffixes, Devanagari script, missing addresses and typos being matched. It also shows same-name businesses in other cities, co-located businesses and no-match queries being rejected. All 8 synthetic queries are resolved correctly. The synthetic cases are easier than the real data, so read this as an illustration, not a benchmark.
Important: the models expect the exact 80 features produced by the repository's pair_features.py, including retrieval route ranks and scores. They cannot score raw name/address strings by themselves.
Evaluation
The metric is macro-averaged per-query F0.5 (precision-weighted). A query with no true matches and no predictions scores 1.
| Stage | Held-out macro F0.5 | Precision | Recall |
|---|---|---|---|
| Base model, top-100 candidates (2,500 confirmation queries) | 0.9453 | 0.9658 | 0.9054 |
| Base model, fast R10 candidates (4,000 fresh queries) | 0.9363 | 0.9665 | 0.8733 |
| + meta model (same 4,000 queries) | 0.9438 | 0.9654 | 0.8986 |
| Submission | Public leaderboard F0.5 |
|---|---|
| Base model | 0.9176 |
| + meta model | 0.9217 |
| + target-ownership decoding | 0.9303 |
Confirmation queries were grouped by normalized name/address keys so that near-duplicates could not cross folds. All thresholds were chosen on a separate tuning fold before confirmation scoring.
Training data and limitations
- Trained from scratch only on the official challenge training labels. The base model used 5,001 grouped fitting entities (1.09M sampled candidate pairs, including hard negatives); the meta model used 4,045 separate entities. No pretrained models or external data were used. The dataset is not redistributed here.
- There are no French training labels, so performance on France is unmeasured.
- The hardest slices are cross-script names (Indic script vs Latin) and records with missing addresses.
- The models are specific to this dataset's schema and feature pipeline and are not a general-purpose entity matcher.
License
MIT
Evaluation results
- Public leaderboard macro F0.5 on Amazon ML Challenge 2026 - Business Entity Resolutionself-reported0.930
- Local held-out macro F0.5 (meta model, R10 candidates) on Amazon ML Challenge 2026 - Business Entity Resolutionself-reported0.944






