fa_dep_news_lg / README.md
Phazel's picture
Regenerate model card: frontmatter, install, usage, tier cross-links
f630b96 verified
|
Raw History Blame Contribute Delete
4.56 kB
metadata
language:
  - fa
license: cc-by-sa-4.0
library_name: spacy
pipeline_tag: token-classification
tags:
  - spacy
  - token-classification
  - dependency-parsing
  - persian
  - farsi

fa_dep_news_lg

Persian dependency pipeline optimized for CPU. Components: tok2vec, tagger, morphologizer, trainable_lemmatizer, parser. No NER, see fa_core_news_lg.

Install

pip install https://huggingface.co/Phazel/fa_dep_news_lg/resolve/main/fa_dep_news_lg-3.8.0-py3-none-any.whl

Use

import spacy

nlp = spacy.load("fa_dep_news_lg")
doc = nlp("شرکت ایران خودرو اعلام کرد که تولید خود را افزایش می‌دهد.")

print([(t.text, t.pos_, t.tag_, t.lemma_, t.dep_) for t in doc])

Accuracy

Scored with spacy benchmark accuracy on the held-out PerDT test split.

Metric Score
Tokenization accuracy 99.96
XPOS tag accuracy 96.55
UPOS tag accuracy 96.68
Morphological features 96.70
Lemma accuracy 98.08
Unlabelled attachment (UAS) 90.96
Labelled attachment (LAS) 86.60
Sentence segmentation F 99.18

Other packages in this family

Pipeline Tier LAS ENTS_F Wheel
fa_dep_news_sm sm 85.15 - 7.9 MB
fa_core_news_sm sm 85.15 71.87 13.5 MB
fa_ent_news_sm sm - 71.87 5.9 MB
fa_dep_news_md md 86.34 - 62.6 MB
fa_core_news_md md 86.34 74.71 68.5 MB
fa_ent_news_md md - 74.71 60.6 MB
fa_dep_news_lg (this one) lg 86.60 - 229.3 MB
fa_core_news_lg lg 86.60 75.94 235.2 MB
fa_ent_news_lg lg - 75.94 227.3 MB
fa_core_news_trf trf 90.79 82.89 608.2 MB

Tiers: sm hash embeddings, no vectors, md 50k x 300d floret vectors, lg 200k x 300d floret vectors, trf fine-tuned transformer, GPU recommended.

Standalone vector tables, usable as --paths.vectors for your own training:

Vectors Rows Used by Wheel
fa_floret_400k 50,000 md tier 54.5 MB
fa_floret_full_wiki 50,000 no shipped pipeline 54.9 MB
fa_floret_wiki_200k 200,000 lg tier 221.3 MB

Training scripts, configs and evaluation: https://github.com/Fazel94/spacy-persian.

Sources

Source Author Licence
UD_Persian-PerDT (PerUDT v1.0) Mohammad Sadegh Rasooli, Pegah Safari, Amirsaeid Moloodi, Alireza Nourian CC BY-SA 4.0
spaCy lang/fa language data (stop words originally from HAZM) Explosion and spaCy contributors MIT
fa_floret static vectors (lg tier: 200k rows x 300d floret table trained on the full Persian Wikipedia dump, 5 epochs, with floret-torch) Kiyarash Fazeli CC BY-SA 4.0

Notes

Trained on UD_Persian-PerDT, licensed CC BY-SA 4.0; this pipeline is therefore distributed under CC BY-SA 4.0 with attribution to the treebank authors. Multiword tokens (pronominal clitics, enclitic copulas) were merged with spacy convert --merge-subtokens, so a small number of XPOS tags are composite (e.g. N_IANM_PR_JOPER) and ~1.5% of lemmas contain a space. doc.noun_chunks under-fires on this pipeline: spacy/lang/fa/syntax_iterators.py matches ClearNLP labels that do not exist in Universal Dependencies, see docs/upstream/fa-noun-chunks.md. This is the lg tier: identical architecture to sm/md but a larger static floret vector table (200,000 rows x 300 dimensions, minn=maxn=5, hash_count=2) trained on the full Persian Wikipedia dump for 5 epochs with floret-torch. Same zero-OOV rationale as md (see docs/MODELS.md): floret hashes subwords into a fixed table, so token.has_vector is always True despite Persian's ZWNJ (U+200C) inconsistency.