A preprint on historical Arabic manuscripts reports a 75-fold difference between its annotation workflow and manual work, but the comparison is tied to a specific test. Across seven fully page-validated books, the workflow processed 3,000 lines per hour, compared with 40 lines per hour for manual annotation.
The result comes with an important condition: RefLAM starts with manuscript page images paired with clean reference transcriptions. It is designed to produce validated line-level ground truth, meaning text and page-line links that have been checked by people.
A fast pass still keeps people in the loop
The pipeline has five stages. It segments the page into lines, uses a structured multimodal language-model OCR step to read them, normalises the text, anchors it to the page and then uses greedy fuzzy alignment to match it with the reference transcription. Confidence scores guide the subsequent human review.
The study gives a precise meaning to its top confidence score. A score of 100 means that the normalised strings are identical character by character under the study's indel-similarity test, and the authors report no counterexample across the full corpus. It does not certify raw diacritization on the page, pixel-perfect geometry or a unique placement of the reference text in the corpus.
Human checking remains part of the process. Every line was reviewed before entering the released dataset: confidence-100 lines received a rapid visual comparison, while lines below 100 received detailed review.
The corpus is broad in count, selective in confidence
The resulting AraMS-28k corpus contains 14 historical Arabic books spread across 3,043 pages. It includes 27,971 main-text lines and 629 lines written in the margins.
The validation was divided into seven books checked at the page level and seven further books handled in a line-validation phase. In that second phase, 16,533 confidence-100 main-text lines were retained within one week. Lines below the threshold were excluded rather than manually corrected, so the later set is a selected high-confidence collection rather than a full correction of every candidate line.
Margin material proved harder to connect to the reference text. Insertion anchors were assigned to 191 of the 629 margin lines, or approximately 30%, leaving most margin lines without a confident attachment point.
Recognition results varied by model and script
The researchers also used the corpus for a downstream handwritten-text recognition test. They fine-tuned two Muharaf-pretrained models on nine books containing 19,739 training lines, then evaluated them on three held-out books containing 6,746 lines under matched data conditions.
Kraken recorded the lower overall weighted character error rate in this comparison. Its rate was 23.31%, compared with 26.74% for HATFormer. The result describes this evaluation setup and does not establish that one model is universally better.
The same pattern appeared across the three scripts tested. For Kraken and HATFormer respectively, the character error rates were 11.65% and 13.26% for Ruq‘ah, 22.62% and 25.37% for Naskh, and 32.71% and 37.88% for Maghrebi. Ruq‘ah had the lowest rates for both models, while Maghrebi had the highest.
A measured claim about scale
Taken together, the findings describe a workflow that can compare manuscript images with existing clean transcriptions at high reported throughput while retaining human review. They do not show that review can be eliminated, and the confidence-100 rule certifies normalised-string identity rather than raw page detail or a unique placement of the reference text.
The authors report releasing the resources under the CC BY-NC-SA 4.0 licence. The document is an arXiv preprint, version 1, dated 25 August 2026, and its peer-reviewed publication status is not reported. Funding information is not reported in the supplied document.
Paper data and sources
Original title: RefLAM: A Reference-Grounded Line Annotation Pipeline for Historical Arabic Manuscripts
Authors: Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib, Sahraoui Dhelim
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-25
DOI: Not available
Original paper · Full text