Datasets / mcit

Khaliji Speech Corpus

Government-RestrictedGov ReuseStatisticsAudio/WAV + JSONL1,800 hours1.6 TB
87Data Health

260 downloads

Updated quarterly. Last update 52 days ago. Next expected in 38 days.

About this dataset

1,800 hours of Gulf-dialect Arabic speech with verbatim transcripts, speaker metadata, and acoustic condition tags, recorded from consented volunteers and de-identified broadcast material. Balanced across Qatari, broader Khaliji, and MSA registers. The primary training corpus for sovereign Arabic ASR.

Usage rights

Restricted to approved ASR development under recorded speaker consent; voice data is treated as biometric-adjacent personal data under the PDPPL and may not leave the sovereign environment.

Lineage

Collected through consented recording campaigns and licensed broadcast archives, transcribed by dialect-trained annotators with 5% double-transcription QA, and segmented into utterance-level shards.

#speech#asr#gulf-dialect#arabic#audio#transcripts

2 models trained on this dataset

Owning entity

mcMCIT
Updated
2026-06-08
Cadence
Quarterly
Format
Audio/WAV + JSONL
Latest version
v2.2