Preprint

Arabic NLP research surges, but dialect coverage remains uneven

An arXiv version 2 preprint dated 25 August 2026 maps 7,120 papers and finds core-task coverage centered on Modern Standard Arabic.

A preprint analysis of Arabic natural-language-processing research finds that most papers in its final corpus appeared recently, while coverage of Arabic dialects varied widely. Approximately 82% of the 7,120 papers were published after 2020. In the core task-dialect matrix, Modern Standard Arabic (MSA) appeared in over 50% of papers across all listed tasks, while Hejazi and Hassaniya had zero detected coverage. The figures describe records captured by the study; zero detected coverage does not prove that no relevant research exists.

The work is a multi-source bibliometric and topic-based study of publication trends, themes, citation predictors, collaboration and geographic contributions. The document is an arXiv version 2 preprint dated 25 August 2026.

A large, recent record

The final analysis set comprised 7,120 unique papers published between 1960 and 2026. The study reports that approximately 82% of papers were published after 2020. It also notes that data for 2026 were incomplete.

To build the corpus, the analysis drew records from arXiv, ACL Anthology, Semantic Scholar, Crossref, OpenAlex and a targeted OpenAlex subset, covering multiple publication types. A paper passed the term-based relevance filter when its title or abstract contained at least two curated Arabic NLP terms.

DOI-based deduplication produced 9,490 unique papers. Grouping records without DOIs by title and year then produced 9,141 unique papers with abstracts. After the filtering pipeline, 7,120 remained for the final analysis.

The field’s broad themes

To sort the literature into broad themes, the study compared BERTopic, a transformer-based topic model, with LDA, a traditional bag-of-words baseline. After the outlier category was excluded, BERTopic presented 19 interpretable topics. Its largest topic, grouped around text, speech, translation and recognition, contained 2,942 papers, or 41.3% of the corpus.

On the reported quality measures, BERTopic scored higher than LDA on both diversity and coherence. Diversity was 0.868 for BERTopic versus 0.720 for LDA; coherence was 0.772 versus 0.028.

The topic mix also varied across the reported time periods. A supplementary chi-squared test gave a value of 75.04 with 8 degrees of freedom and a p-value below 0.001.

Citations offer a narrower signal

Older papers tended to have more citations in the available corpus: paper age and citation count were positively correlated, with r = 0.245 and p < 0.001. That relationship is observational, and citation data were incomplete.

An ordinary least-squares model, a calculation that estimates how recorded features relate to citation counts, was statistically significant, with F = 92.83 and p < 0.001. Yet it explained only about 10.5% of the variation in citation counts, with R-squared of 0.105 and adjusted R-squared of 0.104. Its residuals were right-skewed and its standard errors were non-robust, limiting the model’s explanatory reach.

In the reported coefficients, Semantic Scholar indexing (+5.455), OpenAlex indexing (+11.086), institutional presence (+8.730) and matched-term count (+1.208) were positive. Centered publication year (-1.234) and OpenAlex extra (-9.042) were negative. These are associations in an observational model; the results do not show that indexing, institutional presence or matched-term count caused citation differences.

Concentration by place and dialect

The geographic analysis covered 2,568 papers with an identified country and counted each unique country once per paper. Saudi Arabia led the reported affiliation counts with 519, followed by the United States with 463 and Egypt with 266.

In the institutional counts, King Saud University ranked first with 142 papers, followed by Cairo University with 67 and Columbia University with 64. The institutional data covered only part of the corpus, so these figures describe records with available affiliation information.

Dialect mentions in titles and abstracts were dominated by Modern Standard Arabic, Egyptian and Maghrebi/Darija: 1,553, 462 and 341 mentions respectively. Hejazi appeared once, Sudanese 20 times, Yemeni 28 times and Hassaniya not at all.

The same imbalance appeared in the core task matrix. Modern Standard Arabic appeared in over 50% of papers across all listed tasks; Hejazi and Hassaniya had zero detected coverage, while Sudanese and Yemeni appeared only sporadically. Because the matrix used keyword detection, zero coverage should be read as no detected coverage in these records, not as proof that no relevant work exists.

Specialized summarization was narrower still in the reported matrix: 76.9% of coverage was MSA, 15.4% Egyptian and 7.7% Gulf, with zero coverage for the other listed dialects. These percentages came from row-normalized keyword mentions rather than full-text verification.

A map of the captured literature

Publication volume and citation impact did not line up neatly. The largest topic had the highest H-index, 90, but averaged 14.47 citations per paper. Smaller topics recorded higher averages: legal BERT had 29.90, cultural translation 24.90 and patient-related topics 21.84.

The co-authorship analysis reported Nizar Habash as having the highest betweenness centrality, with a score of 0.025. The study also released a curated corpus of 9,141 papers with unified metadata and the complete analysis code.

Taken together, these findings map publication, topic, citation, collaboration, geographic and task-dialect patterns in records retrieved from selected platforms and filters. Dialect and task results came from exact keyword matching in titles and abstracts, while country, institutional and citation findings relied on records with the relevant fields available. The study therefore describes the captured literature; it does not provide a complete census, show that metadata variables caused citation differences, or prove that zero detected task-dialect coverage means no research exists.

Paper data and sources

Original title: A Comprehensive Analysis of Arabic Natural Language Processing Research: Trends, Topic Evolution, and Research Gaps -- A Bibliometric and Topic-Based Study
Authors: Mullosharaf K. Arabov
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-24
DOI: Not available
Original paper · Full text

Versions and corrections

  1. Published automatically after legal-source, freshness, evidence, and independent-verification gates passed.