Datasets / mcit

Citizen Feedback Corpus

PublicOpenStatisticsJSONL412K rows1.1 GB
94Data Health

8.4K downloads

Updated monthly. Last update 9 days ago. Next expected in 21 days.

About this dataset

Anonymized Arabic feedback texts submitted by residents through government digital channels, annotated with sentiment, topic, and dialect labels by a double-annotation workflow. Names, contact details, and identifiers are removed and texts are lightly redacted before release. The reference corpus for Arabic citizen-voice NLP in Qatar.

Usage rights

Open for research and model training with attribution to MCIT. Texts were anonymized under a PDPPL-reviewed protocol; attempting to re-identify authors is prohibited.

Lineage

Sourced from feedback forms across government portals and the national contact centre, passed through automated PII scrubbing, then double-annotated with adjudication (inter-annotator agreement 0.86 Cohen's kappa).

#arabic-nlp#sentiment#citizen-feedback#annotated-corpus#open-data#dialect

3 models trained on this dataset

Owning entity

mcMCIT
Updated
2026-07-21
Cadence
Monthly
Format
JSONL
Latest version
v4.1