dfine-egret-x-rukopys

A D-FINE-X (DFineForObjectDetection, ~62.7M params) document-layout detector for handwritten Ukrainian pages, fine-tuned on the Rukopys dataset with a 7-class head: 0 handwritten · 1 printed · 2 formula · 3 table · 4 annotation · 5 image · 6 graph.

TL;DR

value
Architecture D-FINE (HGNetV2 backbone, RT-DETR-style decoder with fine-grained distribution refinement), DFineForObjectDetection
Parameters ~62.7M
Init from docling-project/docling-layout-egret-xlarge (17-class document-layout checkpoint)
Classes handwritten, printed, formula, table, annotation, image, graph
Decoding set prediction, 300 object queries, no NMS
Gold-val mAP@50 / mAP@50-95 0.6704 / 0.4441 (macro over 7 classes)
Input a document page image, 640×640
Output class-labeled region bounding boxes

Intended use

Detecting and classifying page regions on handwritten Ukrainian documents, upstream of region recognizers (e.g. Hukyl/trocr-large-rukopys for text, Hukyl/trocr-base-rukopys-formula for formulas). For a stronger single detector on the same task, see Hukyl/doclayout-yolov10-rukopys.

How to use

import torch
from PIL import Image
from transformers import AutoImageProcessor, AutoModelForObjectDetection

repo = "Hukyl/dfine-egret-x-rukopys"
model = AutoModelForObjectDetection.from_pretrained(repo).eval()
processor = AutoImageProcessor.from_pretrained(repo)

image = Image.open("page.jpg").convert("RGB")
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
    outputs = model(**inputs)

# threshold ~0.5-0.6 is a balanced operating range
target_sizes = torch.tensor([image.size[::-1]])  # (height, width)
result = processor.post_process_object_detection(
    outputs, threshold=0.5, target_sizes=target_sizes
)[0]
for score, label, box in zip(result["scores"], result["labels"], result["boxes"]):
    print(model.config.id2label[label.item()], round(score.item(), 3), box.tolist())

Files

model.safetensors, config.json,   the shipped checkpoint — loads directly
preprocessor_config.json          with from_pretrained
training_meta.json                recorded recipe + the shipped checkpoint's metrics
metrics.json                      val metrics, aggregate + per-class
selection_comparison.json         the two checkpoint-selection axes side by side
training_log.jsonl                per-epoch loss / mAP / P / R curves

Head re-initialisation (17 → 7)

The backbone, encoder, decoder, and the class-agnostic box-regression head load from the base checkpoint as-is; the classification head is re-initialised to the 7 Rukopys classes, with id2label/label2id baked into the config.

Training

Fine-tuned on the human-labeled Rukopys train split via the 🤗 Trainer. The image processor applies a deterministic 640×640 resize with ImageNet normalisation.

Hyperparameters (as launched)

hyperparameter value
epochs 60 (early stopping, patience 15 on val mAP@50)
batch / image size 32 / 640
optimizer / schedule AdamW, linear decay
learning rate 1e-4
warmup ratio 0.1
weight decay 1e-4
freeze none
augmentation none
eval / checkpoint cadence every epoch
checkpoint best val mAP@50 (epoch 24)
seed 42

The full recorded recipe ships in training_meta.json; per-epoch curves in training_log.jsonl.

Checkpoint selection

The run tracked two selection axes; val loss and val mAP@50 are anti-correlated across this run (Pearson ≈ −0.81). This repo ships the mAP@50-selected checkpoint.

selected on mAP@50 mAP@50-95 precision recall
val mAP@50 (shipped) 0.6704 0.4441 0.8971 0.8787
val loss 0.6413 0.4356 0.9184 0.8695

Class distribution (region counts)

class train val
handwritten 19,420 2,157
printed 266 42
formula 2,545 347
table 128 14
annotation 494 60
image 115 4
graph 54 5
total 23,022 2,629

Results

Held-out val split of the Rukopys train data: 133 pages / 2,629 regions (stratified holdout, seed 42). mAP is threshold-free with box matching at IoU 0.50 (mAP@50–95 averages IoU 0.50:0.95); precision and recall are reported at a fixed confidence ≥ 0.50 operating point (class-aware greedy matching at IoU ≥ 0.50). Per-class precision/recall are not recorded.

Aggregate (macro over 7 classes)

metric value
mAP@50 0.6704
mAP@50-95 0.4441
precision 0.8971
recall 0.8787

Per class

class n mAP@50 mAP@50-95
handwritten 2,157 0.9329 0.5960
printed 42 0.6545 0.3685
formula 347 0.9044 0.6107
table 14 0.8522 0.4763
annotation 60 0.2561 0.1148
image 4 0.4515 0.4149
graph 5 0.6411 0.5272

We also acknowledge that printed/table/annotation/image/graph n is quite small, so measuring detection metrics against them is quite noisy.

Limitations

  • annotation (mAP@50 0.26) and image (mAP@50 0.45) are weak classes.
  • Trained and evaluated at 640×640; very small or dense regions on high-resolution scans may benefit from a higher inference resolution.
  • Precision/recall are reported at the conf ≥ 0.50 operating point; calibrate the inference threshold against your own target metric.
  • Handwritten Ukrainian school/archival-style pages only; behaviour on other document types is untested.
  • Single seed and validation split — no across-run variance estimate.

Training data & attribution

dataset source license role
Rukopys UkrainianCatholicUniversity/rukopys CC BY 4.0 gold fine-tune

Model weights are Apache-2.0, inherited from D-FINE and the base checkpoint.

Downloads last month
127
Safetensors
Model size
62.7M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Hukyl/dfine-egret-x-rukopys

Finetuned
(2)
this model
Quantizations
1 model

Dataset used to train Hukyl/dfine-egret-x-rukopys

Evaluation results