Instructions to use Phazel/fa_dep_news_lg with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- spaCy
How to use Phazel/fa_dep_news_lg with spaCy:
!pip install https://huggingface.co/Phazel/fa_dep_news_lg/resolve/main/fa_dep_news_lg-any-py3-none-any.whl # Using spacy.load(). import spacy nlp = spacy.load("fa_dep_news_lg") # Importing as module. import fa_dep_news_lg nlp = fa_dep_news_lg.load() - Notebooks
- Google Colab
- Kaggle
|
Download README.md from Phazel/fa_dep_news_lg: direct link, hf CLI and curl.
- Browser
- Download file 4.56 kB
-
https://huggingface.co/Phazel/fa_dep_news_lg/resolve/main/README.md
- Command line
-
hf download hf://Phazel/fa_dep_news_lg/README.md
-
curl -L -o README.md https://huggingface.co/Phazel/fa_dep_news_lg/resolve/main/README.md
4.56 kB
| language: | |
| - fa | |
| license: cc-by-sa-4.0 | |
| library_name: spacy | |
| pipeline_tag: token-classification | |
| tags: | |
| - spacy | |
| - token-classification | |
| - dependency-parsing | |
| - persian | |
| - farsi | |
| # fa_dep_news_lg | |
| Persian dependency pipeline optimized for CPU. Components: tok2vec, tagger, morphologizer, trainable_lemmatizer, parser. No NER, see fa_core_news_lg. | |
| ## Install | |
| ```bash | |
| pip install https://huggingface.co/Phazel/fa_dep_news_lg/resolve/main/fa_dep_news_lg-3.8.0-py3-none-any.whl | |
| ``` | |
| ## Use | |
| ```python | |
| import spacy | |
| nlp = spacy.load("fa_dep_news_lg") | |
| doc = nlp("شرکت ایران خودرو اعلام کرد که تولید خود را افزایش میدهد.") | |
| print([(t.text, t.pos_, t.tag_, t.lemma_, t.dep_) for t in doc]) | |
| ``` | |
| ## Accuracy | |
| Scored with `spacy benchmark accuracy` on the held-out PerDT test split. | |
| | Metric | Score | | |
| | --- | ---: | | |
| | Tokenization accuracy | 99.96 | | |
| | XPOS tag accuracy | 96.55 | | |
| | UPOS tag accuracy | 96.68 | | |
| | Morphological features | 96.70 | | |
| | Lemma accuracy | 98.08 | | |
| | Unlabelled attachment (UAS) | 90.96 | | |
| | Labelled attachment (LAS) | 86.60 | | |
| | Sentence segmentation F | 99.18 | | |
| ## Other packages in this family | |
| | Pipeline | Tier | LAS | ENTS_F | Wheel | | |
| | --- | --- | ---: | ---: | ---: | | |
| | [`fa_dep_news_sm`](https://huggingface.co/Phazel/fa_dep_news_sm) | `sm` | 85.15 | - | 7.9 MB | | |
| | [`fa_core_news_sm`](https://huggingface.co/Phazel/fa_core_news_sm) | `sm` | 85.15 | 71.87 | 13.5 MB | | |
| | [`fa_ent_news_sm`](https://huggingface.co/Phazel/fa_ent_news_sm) | `sm` | - | 71.87 | 5.9 MB | | |
| | [`fa_dep_news_md`](https://huggingface.co/Phazel/fa_dep_news_md) | `md` | 86.34 | - | 62.6 MB | | |
| | [`fa_core_news_md`](https://huggingface.co/Phazel/fa_core_news_md) | `md` | 86.34 | 74.71 | 68.5 MB | | |
| | [`fa_ent_news_md`](https://huggingface.co/Phazel/fa_ent_news_md) | `md` | - | 74.71 | 60.6 MB | | |
| | `fa_dep_news_lg` (this one) | `lg` | 86.60 | - | 229.3 MB | | |
| | [`fa_core_news_lg`](https://huggingface.co/Phazel/fa_core_news_lg) | `lg` | 86.60 | 75.94 | 235.2 MB | | |
| | [`fa_ent_news_lg`](https://huggingface.co/Phazel/fa_ent_news_lg) | `lg` | - | 75.94 | 227.3 MB | | |
| | [`fa_core_news_trf`](https://huggingface.co/Phazel/fa_core_news_trf) | `trf` | 90.79 | 82.89 | 608.2 MB | | |
| Tiers: `sm` hash embeddings, no vectors, `md` 50k x 300d floret vectors, `lg` 200k x 300d floret vectors, `trf` fine-tuned transformer, GPU recommended. | |
| Standalone vector tables, usable as `--paths.vectors` for your own training: | |
| | Vectors | Rows | Used by | Wheel | | |
| | --- | ---: | --- | ---: | | |
| | [`fa_floret_400k`](https://huggingface.co/Phazel/fa_floret_400k) | 50,000 | `md` tier | 54.5 MB | | |
| | [`fa_floret_full_wiki`](https://huggingface.co/Phazel/fa_floret_full_wiki) | 50,000 | no shipped pipeline | 54.9 MB | | |
| | [`fa_floret_wiki_200k`](https://huggingface.co/Phazel/fa-floret-wiki-vectors) | 200,000 | `lg` tier | 221.3 MB | | |
| Training scripts, configs and evaluation: <https://github.com/Fazel94/spacy-persian>. | |
| ## Sources | |
| | Source | Author | Licence | | |
| | --- | --- | --- | | |
| | [UD_Persian-PerDT (PerUDT v1.0)](https://github.com/UniversalDependencies/UD_Persian-PerDT) | Mohammad Sadegh Rasooli, Pegah Safari, Amirsaeid Moloodi, Alireza Nourian | CC BY-SA 4.0 | | |
| | [spaCy lang/fa language data (stop words originally from HAZM)](https://github.com/explosion/spaCy/tree/master/spacy/lang/fa) | Explosion and spaCy contributors | MIT | | |
| | [fa_floret static vectors (lg tier: 200k rows x 300d floret table trained on the full Persian Wikipedia dump, 5 epochs, with floret-torch)](https://huggingface.co/Phazel/fa-floret-wiki-vectors) | Kiyarash Fazeli | CC BY-SA 4.0 | | |
| ## Notes | |
| Trained on UD_Persian-PerDT, licensed CC BY-SA 4.0; this pipeline is therefore distributed under CC BY-SA 4.0 with attribution to the treebank authors. Multiword tokens (pronominal clitics, enclitic copulas) were merged with `spacy convert --merge-subtokens`, so a small number of XPOS tags are composite (e.g. N_IANM_PR_JOPER) and ~1.5% of lemmas contain a space. doc.noun_chunks under-fires on this pipeline: spacy/lang/fa/syntax_iterators.py matches ClearNLP labels that do not exist in Universal Dependencies, see docs/upstream/fa-noun-chunks.md. This is the `lg` tier: identical architecture to `sm`/`md` but a larger static floret vector table (200,000 rows x 300 dimensions, minn=maxn=5, hash_count=2) trained on the full Persian Wikipedia dump for 5 epochs with floret-torch. Same zero-OOV rationale as `md` (see docs/MODELS.md): floret hashes subwords into a fixed table, so `token.has_vector` is always True despite Persian's ZWNJ (U+200C) inconsistency. | |