Models / mcit

mcit/arabic-dialect-sentiment

Arabic Dialect Sentiment

PublicLimited riskOpenArabic NLPTransformersArabicv3.2
12K 418K/30d 486

Macro-F1 (Gulf test set)

0.894

Accuracy

0.907

F1: negative class

0.871

Dialect ID accuracy

0.942

Median latency

38 ms

About this model

Sentiment classifier tuned for Gulf Arabic, with first-class support for the Qatari dialect alongside Modern Standard Arabic. Trained on anonymized citizen feedback and public social content, it distinguishes positive, negative, and neutral sentiment with dialect-aware tokenization. The reference model behind national service-satisfaction dashboards.

Intended use

Classifying citizen feedback, service reviews, and public sentiment streams for government entities. Not intended for decisions about individuals.

Training data lineage

Fine-tuned from CAMeLBERT-Mix on the Citizen Feedback Corpus (v3.1) plus 210K weakly-labelled Gulf social posts, with human-verified evaluation splits.

Limitations & bias notes

Accuracy drops on Maghrebi and Levantine dialects (−9 to −14 F1 versus Gulf Arabic). Sarcasm and mixed code-switched Arabic-English messages remain the largest error class; scores should be aggregated, never used to act on a single message.

Evaluation metrics

Macro-F1 (Gulf test set)0.894
Accuracy0.907
F1: negative class0.871
Dialect ID accuracy0.942
Median latency38 ms
#arabic#sentiment#gulf-dialect#citizen-feedback#camelbert#text-classification

Try it

Live sandbox

Used in workflows

Retiring this asset would require these chains to be re-pointed first.

Feedback

5.0(1)

Owning entity

mcMCIT
Updated
2026-07-14
Latest version
v3.2
License
Open
Access
Open