wahaj-labs/arabic-invoice-extractor
Arabic Invoice Extractor
Field F1 (header)
0.93
Field F1 (line items)
0.85
Character error rate
2.1%
Straight-through rate
0.78
About this model
Document-understanding pipeline that extracts structured fields, supplier, CR number, line items, VAT-ready totals, dates, and payment terms, from Arabic, English, and mixed-language invoices. Combines a layout-aware OCR stage with a field-tagging transformer robust to scans, photos, and PDFs.
Intended use
Accounts-payable automation and procurement digitisation for government entities and enterprises processing bilingual supplier invoices.
Training data lineage
Layout backbone pre-trained on document pages from the arabic-gov-docs-corpus, then fine-tuned on 96,000 annotated invoices contributed by launch partners.
Limitations & bias notes
Handwritten amendments and stamps overlapping key fields remain the dominant failure mode. Line-item extraction accuracy drops on multi-page invoices with carried-over subtotals and on low-resolution WhatsApp photos.
Evaluation metrics
| Field F1 (header) | 0.93 |
| Field F1 (line items) | 0.85 |
| Character error rate | 2.1% |
| Straight-through rate | 0.78 |
Try it
Live sandboxinvoice_gulf_supplies_04417.pdf
fixtureFeedback
Owning entity
- Updated
- 2026-07-03
- Latest version
- 2.1.0
- License
- Revenue Share
- Access
- Open
Trained on
arabic-gov-docs-corpus →