Instructions to use Hukyl/dfine-egret-x-rukopys with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Hukyl/dfine-egret-x-rukopys with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("object-detection", model="Hukyl/dfine-egret-x-rukopys")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForObjectDetection tokenizer = AutoTokenizer.from_pretrained("Hukyl/dfine-egret-x-rukopys") model = AutoModelForObjectDetection.from_pretrained("Hukyl/dfine-egret-x-rukopys", device_map="auto") - Notebooks
- Google Colab
- Kaggle
dfine-egret-x-rukopys
A D-FINE-X (DFineForObjectDetection, ~62.7M params) document-layout detector
for handwritten Ukrainian pages, fine-tuned on the
Rukopys
dataset with a 7-class head:
0 handwritten · 1 printed · 2 formula · 3 table · 4 annotation · 5 image · 6 graph.
TL;DR
| value | |
|---|---|
| Architecture | D-FINE (HGNetV2 backbone, RT-DETR-style decoder with fine-grained distribution refinement), DFineForObjectDetection |
| Parameters | ~62.7M |
| Init from | docling-project/docling-layout-egret-xlarge (17-class document-layout checkpoint) |
| Classes | handwritten, printed, formula, table, annotation, image, graph |
| Decoding | set prediction, 300 object queries, no NMS |
| Gold-val mAP@50 / mAP@50-95 | 0.6704 / 0.4441 (macro over 7 classes) |
| Input | a document page image, 640×640 |
| Output | class-labeled region bounding boxes |
Intended use
Detecting and classifying page regions on handwritten Ukrainian documents, upstream
of region recognizers (e.g.
Hukyl/trocr-large-rukopys for
text, Hukyl/trocr-base-rukopys-formula
for formulas). For a stronger single detector on the same task, see
Hukyl/doclayout-yolov10-rukopys.
How to use
import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForObjectDetection
repo = "Hukyl/dfine-egret-x-rukopys"
model = AutoModelForObjectDetection.from_pretrained(repo).eval()
processor = AutoImageProcessor.from_pretrained(repo)
image = Image.open("page.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
# threshold ~0.5-0.6 is a balanced operating range
target_sizes = torch.tensor([image.size[::-1]]) # (height, width)
result = processor.post_process_object_detection(
outputs, threshold=0.5, target_sizes=target_sizes
)[0]
for score, label, box in zip(result["scores"], result["labels"], result["boxes"]):
print(model.config.id2label[label.item()], round(score.item(), 3), box.tolist())
Files
model.safetensors, config.json, the shipped checkpoint — loads directly
preprocessor_config.json with from_pretrained
training_meta.json recorded recipe + the shipped checkpoint's metrics
metrics.json val metrics, aggregate + per-class
selection_comparison.json the two checkpoint-selection axes side by side
training_log.jsonl per-epoch loss / mAP / P / R curves
Head re-initialisation (17 → 7)
The backbone, encoder, decoder, and the class-agnostic box-regression head load
from the base checkpoint as-is; the classification head is re-initialised to the
7 Rukopys classes, with id2label/label2id baked into the config.
Training
Fine-tuned on the human-labeled Rukopys train split via the 🤗 Trainer. The
image processor applies a deterministic 640×640 resize with ImageNet
normalisation.
Hyperparameters (as launched)
| hyperparameter | value |
|---|---|
| epochs | 60 (early stopping, patience 15 on val mAP@50) |
| batch / image size | 32 / 640 |
| optimizer / schedule | AdamW, linear decay |
| learning rate | 1e-4 |
| warmup ratio | 0.1 |
| weight decay | 1e-4 |
| freeze | none |
| augmentation | none |
| eval / checkpoint cadence | every epoch |
| checkpoint | best val mAP@50 (epoch 24) |
| seed | 42 |
The full recorded recipe ships in training_meta.json; per-epoch curves in
training_log.jsonl.
Checkpoint selection
The run tracked two selection axes; val loss and val mAP@50 are anti-correlated across this run (Pearson ≈ −0.81). This repo ships the mAP@50-selected checkpoint.
| selected on | mAP@50 | mAP@50-95 | precision | recall |
|---|---|---|---|---|
| val mAP@50 (shipped) | 0.6704 | 0.4441 | 0.8971 | 0.8787 |
| val loss | 0.6413 | 0.4356 | 0.9184 | 0.8695 |
Class distribution (region counts)
| class | train | val |
|---|---|---|
| handwritten | 19,420 | 2,157 |
| printed | 266 | 42 |
| formula | 2,545 | 347 |
| table | 128 | 14 |
| annotation | 494 | 60 |
| image | 115 | 4 |
| graph | 54 | 5 |
| total | 23,022 | 2,629 |
Results
Held-out val split of the Rukopys train data: 133 pages / 2,629 regions (stratified holdout, seed 42). mAP is threshold-free with box matching at IoU 0.50 (mAP@50–95 averages IoU 0.50:0.95); precision and recall are reported at a fixed confidence ≥ 0.50 operating point (class-aware greedy matching at IoU ≥ 0.50). Per-class precision/recall are not recorded.
Aggregate (macro over 7 classes)
| metric | value |
|---|---|
| mAP@50 | 0.6704 |
| mAP@50-95 | 0.4441 |
| precision | 0.8971 |
| recall | 0.8787 |
Per class
| class | n | mAP@50 | mAP@50-95 |
|---|---|---|---|
| handwritten | 2,157 | 0.9329 | 0.5960 |
| printed | 42 | 0.6545 | 0.3685 |
| formula | 347 | 0.9044 | 0.6107 |
| table | 14 | 0.8522 | 0.4763 |
| annotation | 60 | 0.2561 | 0.1148 |
| image | 4 | 0.4515 | 0.4149 |
| graph | 5 | 0.6411 | 0.5272 |
We also acknowledge that printed/table/annotation/image/graph n is quite
small, so measuring detection metrics against them is quite noisy.
Limitations
annotation(mAP@50 0.26) andimage(mAP@50 0.45) are weak classes.- Trained and evaluated at 640×640; very small or dense regions on high-resolution scans may benefit from a higher inference resolution.
- Precision/recall are reported at the conf ≥ 0.50 operating point; calibrate the inference threshold against your own target metric.
- Handwritten Ukrainian school/archival-style pages only; behaviour on other document types is untested.
- Single seed and validation split — no across-run variance estimate.
Training data & attribution
| dataset | source | license | role |
|---|---|---|---|
| Rukopys | UkrainianCatholicUniversity/rukopys |
CC BY 4.0 | gold fine-tune |
Model weights are Apache-2.0, inherited from D-FINE and the base checkpoint.
- Downloads last month
- 127
Model tree for Hukyl/dfine-egret-x-rukopys
Base model
docling-project/docling-layout-egret-xlargeDataset used to train Hukyl/dfine-egret-x-rukopys
Evaluation results
- mAP@50 on Rukopys, gold valvalidation set self-reported0.670
- mAP@50-95 on Rukopys, gold valvalidation set self-reported0.444