Khaliji Speech Corpus
260 downloads
Updated quarterly. Last update 52 days ago. Next expected in 38 days.
About this dataset
1,800 hours of Gulf-dialect Arabic speech with verbatim transcripts, speaker metadata, and acoustic condition tags, recorded from consented volunteers and de-identified broadcast material. Balanced across Qatari, broader Khaliji, and MSA registers. The primary training corpus for sovereign Arabic ASR.
Usage rights
Restricted to approved ASR development under recorded speaker consent; voice data is treated as biometric-adjacent personal data under the PDPPL and may not leave the sovereign environment.
Lineage
Collected through consented recording campaigns and licensed broadcast archives, transcribed by dialect-trained annotators with 5% double-transcription QA, and segmented into utterance-level shards.
2 models trained on this dataset
Owning entity
- Updated
- 2026-06-08
- Cadence
- Quarterly
- Format
- Audio/WAV + JSONL
- Latest version
- v2.2