8.4K downloads
Updated monthly. Last update 9 days ago. Next expected in 21 days.
About this dataset
Anonymized Arabic feedback texts submitted by residents through government digital channels, annotated with sentiment, topic, and dialect labels by a double-annotation workflow. Names, contact details, and identifiers are removed and texts are lightly redacted before release. The reference corpus for Arabic citizen-voice NLP in Qatar.
Usage rights
Open for research and model training with attribution to MCIT. Texts were anonymized under a PDPPL-reviewed protocol; attempting to re-identify authors is prohibited.
Lineage
Sourced from feedback forms across government portals and the national contact centre, passed through automated PII scrubbing, then double-annotated with adjudication (inter-annotator agreement 0.86 Cohen's kappa).
3 models trained on this dataset
Owning entity
- Updated
- 2026-07-21
- Cadence
- Monthly
- Format
- JSONL
- Latest version
- v4.1