Models / hbku

hbku/arabic-gov-ner

Arabic Gov NER

PublicLimited riskOpenArabic NLPTransformersArabic1.4.0
4.3K 88K/30d 210

Macro F1

0.89

F1 (PERSON)

0.92

F1 (LAW_REF)

0.84

QID redaction recall

0.97

About this model

Token-classification model that extracts persons, organisations, law and decree references, dates, and QID mentions from Arabic government documents. Ships with a redaction mode that masks QIDs and personal names for safe document sharing.

Intended use

Entity extraction and PII redaction in document-processing pipelines for circulars, decisions, and correspondence across government entities.

Training data lineage

Fine-tuned on 62,000 annotated sentences drawn from the arabic-gov-docs-corpus, with QID patterns validated against qid-specimen-templates.

Limitations & bias notes

Recall on transliterated foreign organisation names is noticeably lower than on Arabic-native entities. Law references using informal shorthand (e.g. omitting the year) are frequently missed or partially tagged.

Evaluation metrics

Macro F10.89
F1 (PERSON)0.92
F1 (LAW_REF)0.84
QID redaction recall0.97
#arabic-nlp#ner#redaction#pii#government-documents

Try it

Live sandbox

Feedback

Owning entity

hbHBKU
Updated
2026-05-06
Latest version
1.4.0
License
Open
Access
Open