mcit/arabic-dialect-sentiment
Arabic Dialect Sentiment
Macro-F1 (Gulf test set)
0.894
Accuracy
0.907
F1: negative class
0.871
Dialect ID accuracy
0.942
Median latency
38 ms
About this model
Sentiment classifier tuned for Gulf Arabic, with first-class support for the Qatari dialect alongside Modern Standard Arabic. Trained on anonymized citizen feedback and public social content, it distinguishes positive, negative, and neutral sentiment with dialect-aware tokenization. The reference model behind national service-satisfaction dashboards.
Intended use
Classifying citizen feedback, service reviews, and public sentiment streams for government entities. Not intended for decisions about individuals.
Training data lineage
Fine-tuned from CAMeLBERT-Mix on the Citizen Feedback Corpus (v3.1) plus 210K weakly-labelled Gulf social posts, with human-verified evaluation splits.
Limitations & bias notes
Accuracy drops on Maghrebi and Levantine dialects (−9 to −14 F1 versus Gulf Arabic). Sarcasm and mixed code-switched Arabic-English messages remain the largest error class; scores should be aggregated, never used to act on a single message.
Evaluation metrics
| Macro-F1 (Gulf test set) | 0.894 |
| Accuracy | 0.907 |
| F1: negative class | 0.871 |
| Dialect ID accuracy | 0.942 |
| Median latency | 38 ms |
Try it
Live sandboxUsed in workflows
Retiring this asset would require these chains to be re-pointed first.
Feedback
5.0(1)Owning entity
- Updated
- 2026-07-14
- Latest version
- v3.2
- License
- Open
- Access
- Open
Trained on
citizen-feedback-corpus →