Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Formeze form-field labeler β€” LoRA adapter (struct-v1, 2026-09)

Structure-aware labeling adapter for the Formeze PDF pipeline. For each detected box it generates one string carrying the verbatim printed caption, whose information the box holds, and the choice group it belongs to:

<caption> | <SUBJECT> [| e:<entity>] [| g:<group question>] [| one]

e.g. Name | OTHER_PERSON | e:Emergency Contact, Married | SELF | g:Marital status | one. The serving pipeline parses it (formfieldai/structured_label.py); legacy bare captions still parse. SUBJECT is one of SELF, SPOUSE, DEPENDENT, PARENT_GUARDIAN, CO_APPLICANT, OTHER_PERSON, EMPLOYER, ORGANIZATION, PREPARER, OFFICE_USE, NONE.

Recipe

  • Base: microsoft/Florence-2-base @ 5ca5edf5bd017b9919c05d08aebef5e4c7ac3bac
  • LoRA r=64, alpha=128, dropout 0.05, targets q/k/v/o_proj + fc1/fc2, PEFT 0.18.0
  • Continuation of the previous production adapter (lora-ocr-cleangt-v1e3, revision 5e2968d7) on structured targets
  • Prompts: <REGION_TO_DESCRIPTION> + region tokens + k=12 nearest OCR tokens (unchanged serving contract); decode budget 96 new tokens
  • Targets: 2026-09 Gemini 3.7-flash structure relabel of the AcroForm boxes in bugsiesegal/form-fields-for-layout-labeled-pages @ b2d482e0… (224,481 targets on 7,926 pages; verbatim captions, subjects, entities, choice groups), template-grouped train split, 2 epochs (178,429 samples)
  • The immutable test split was never used for training or model selection.

Validation metrics (template-grouped validation split, 759 images, INT8 T4)

Against the structured (verbatim-caption) references it was trained toward:

metric struct-v1 previous adapter
caption semantic agreement on matched boxes 0.852 0.598
caption exact agreement on matched boxes 0.743 0.371
end-to-end (found and correctly captioned) 0.733 0.514
subject accuracy 0.661 β€”
choice-group pair F1 / exclusive accuracy 0.62 / 0.80 β€”

Against the dataset's original label strings (the release benchmark's references): matched semantic accuracy 0.557 vs 0.605 and end-to-end 0.479 vs 0.520 β€” the two reference sets disagree on 62.5% of fields, mostly in wording. At the level of field meaning (both references and predictions mapped to value concepts by the Formeze label canonicalizer), on boxes a personal profile can fill: 0.730 vs 0.717 (original references) and 0.885 vs 0.773 (structured references).

The immutable test-split benchmark has not been run for this adapter.

Review-aid use only: labels are proposals requiring explicit user approval; no calibrated label confidence exists. See the Formeze model card and model-manifest.json for the confidence contract and release gating.

Downloads last month
18
Safetensors
Model size
0.2B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for bugsiesegal/form-field-labeling-florence

Adapter
(8)
this model