Skip to content
Open access

Exploration of Arabic Collocation Patterns in the Indonesia Al-Youm News Corpus

Jul 2026 · Alinea: Jurnal Bahasa, Sastra, dan Pengajaran · 0 citations · 38 references

Abstract

This study investigates the structural and semantic characteristics of Arabic collocations in the Indonesia Al-Youm digital news corpus, a corpus compiled from Arabic-language online news texts reporting Indonesian social, political, cultural, and economic issues. Using Sketch Engine’s Multiword Term extraction feature, the study identifies and classifies the top 100 recurrent multiword units according to their morphosyntactic structures, grammatical functions, and semantic domains. The findings show that Arabic journalistic discourse in this corpus is dominated by nominal constructions, particularly idafah, noun–adjective patterns, and Masdar-based formations. Semantically, the collocations cluster around six domains: media identity and digital technology, politics and diplomacy, social protection and administration, economy and development, tourism, culture, and geography, and proper names. The novelty of this study lies in its corpus-based mapping of Arabic collocation patterns within an Indonesian digital news context, an area that remains underexplored in Arabic corpus linguistics. The results contribute to the description of Arabic journalistic phraseology and provide empirical input for Arabic for Specific Purposes, especially journalistic Arabic and media translation instruction.

Read PDF

Similar papers

Open access Jul 2026

When Morphology Indexes Prestige: Arabic-English Code-mixing as Linguistic and Social Practice in Saudi Arabia

This study explores the intersection of linguistic form and social meaning within Arabic-English code-mixing in relation to the morphological patterns that arise in mixed speech and their association with prestige and modern identity. Based on naturally occurring spoken and digital data and questionnaires from educated Arabic-English speakers in Saudi Arabia, the analysis reveals that morphological adaptation strategies include the introduction of English lexical items into Arabic morphological patterns, affixal incorporation, and the creation of hybrid lexical forms, whose recurrent patterns provide evidence of systematic linguistic innovation, rather than mere borrowing. These morphological adaptations are analyzed through the lens of indexicality and social meaning (Silverstein, 2003; Eckert, 2008) that correlate language choice with social aspiration, education, and symbolic capital. The occurrence of English insertions in the data is often related to prestige, global orientation and identification with a modern lifestyle, although these meanings are contextually inferred, and Arabic morphology is used as a sign of authenticity and local identity. The study argues that morphological choices in code-mixing are socially motivated, and they serve as a site for negotiating status, identity, and belonging. By drawing on morphological and sociolinguistic approaches, this study illustrates how form and meaning converge to shape prestige meanings in contemporary Arabic-English speech. 

W. Alshammari · 0 citations
Open access Jul 2026

Analysis of Morphological and Semantic Transformations of English Words with Indonesian Affixes in Digital Communication

Nowadays, the increasing use of English words combined with Indonesian affixes in digital communication has given rise to a unique form of linguistic innovation that reflects the interaction between the local and global language. This research aims to analyze how Indonesian affixes influence the morphological structure and semantic meaning of English words which are used in digital communication. This research uses qualitative descriptive to analyze 45 lexical items found from Instagram and YouTube comment sections. The data selected from real online interactions containing English words attached to Indonesian affixes. Data were analyzed by following four stages, such as data reduction, classification, interpretation, and conclusion. The findings show that the Indonesian suffix -nya is the most frequently used representing 66.67% of the collected data, followed by the prefixes di-, ke-, nge-, ter-, se-, and the suffix -an. These affixes modify English words by adapting them to Indonesian grammatical structures while preserving or extending their meanings. Unlike previous studies that primarily examined English-Indonesian lexical borrowing through corpus data or magazine texts, this study contributes new insights by examining the morphological and semantic adaptation of English lexical items in spontaneous digital interactions on Instagram and YouTube.  The research determines that Indonesian-English hybrid expressions showcase linguistic innovation, facilitate efficient digital interaction, and exemplify the active development of bilingual language usage in modern Indonesian online conversations.

Prihatin Puji Astuti, Ria Antika · 0 citations
Open access Jul 2026

Errors in Indonesian Writing in One Tribun Jabar News Article: A Case Study of Spelling, Punctuation, and Standard Vocabulary

This study was motivated by the continued occurrence of inaccuracies in Indonesian writing in online news, particularly in the aspects of spelling, punctuation, and standard word usage. Linguistic accuracy in news portals is important because online media not only convey information but also function as models of Indonesian language use in the digital public sphere. This study aims to identify and describe the forms of writing errors found in one selected Tribun Jabar news article, classify the types of errors, and provide corrected forms in accordance with Indonesian language rules. The study employed a descriptive qualitative approach using documentation, close reading, and note-taking techniques. The data were analysed through identification, classification, interpretation, and the formulation of alternative corrections based on the categories of spelling, punctuation, standard word usage, diction, and sentence effectiveness. The findings reveal 18 instances of writing errors. The most dominant errors occur in standard word usage, diction, and sentence effectiveness, with 8 instances (44.45%). Errors in spelling, typing, formatting, and writing consistency account for 6 instances (33.33%), while punctuation and spacing errors account for 4 instances (22.22%). These findings indicate that language problems in online news are not limited to technical inaccuracies but also involve imprecision in word choice and sentence construction. The study implies the need to strengthen journalistic language quality control, particularly in the editing process of online news production, and highlights the value of news texts as authentic materials for Indonesian language learning. However, because the analysis is limited to one selected news article, the findings cannot yet be generalized to the overall language practices of the Tribun Jabar news portal.

Heni Heryani, Susie Kusumayanthi, Rika Widawati · 0 citations
Review Open access Jul 2026

Mapping the Evolution of Motifs in Indonesian Literature 1918-2025 Using Distant Reading

This study presents a comprehensive overview of a century of Indonesian poetry, short stories, and novels in terms of the development of literary motifs, namely a series of words that are repeatedly deployed and serve as the formal basis of the theme. This study aims to address the scarcity in the scholarship of motifs in Indonesian literature over the long term by aiming to present a diachronic analysis of the changing meaning associations of the most frequently appearing motifs in Indonesian poetry and prose over the past century. This very broad scope does not hinder the analysis because this study uses digital humanities approach, specifically Franco Moretti's distant reading, which is a computational analysis of the statistics of word occurrences based on a digitized literary corpus. This paper analyzes a corpus of 193 award-winning or critically acclaimed Indonesian poetry and prose books (novels and short stories) published between 1918 and 2025, divided into 10 periods, each covering 11 years (with the exception of the final period, 2017-2025, which spanned 9 years). This paper addresses the following research question: is there a pattern underlying the evolution of motifs in Indonesian poetry and prose from 1918 to 2025? The findings of this study are: (1) there is a trend of increasing lexical diversity in Indonesian literature, (2) a number of words that are central motifs in Indonesian literature have experienced changes in their meaning associations to become more abstract, and (3) there are stylistic similarities between works based on similarities in motifs across periods.

Martin Suryajaya, Hilmar Farid, C. Dewi · 0 citations
Review Aug 2026

Romanized Arabic Across Dialects: Views, Usage Patterns, and Linguistic Variation

Arabizi refers to Arabic written in Latin script. Although previous studies have shown that the prevalence and usage of Arabizi vary by factors such as region and age group, most NLP research on Arabic texts treats it as a temporary phenomenon resulting from limited technological support for the Arabic script. In this work, we engage with Arabic speakers to collect insights on their perceptions and usage of Arabizi. We further examine writing norms among speakers of different dialects, focusing on Algerian, Egyptian, Lebanese, Moroccan, and Tunisian Arabic. To this end, we release two resources. First, a character-level alignment of Arabic words to study inter- and intra-dialectal variation across these five dialects, based on words transliterated by survey participants, finding systematic intra-dialectal regularity and inter-dialectal variation. Second, to study Arabic speakers'ability to identify this stylistic variation at the sentence-level, we build a manually curated parallel corpus of sentences written in Arabic script alongside multiple Arabizi transliterations, collected from speakers of the same five dialects. Our study presents the largest human-centered, cross-dialectal study of Arabizi's perceptions and practices to date.

Amr Keleg, Ahmed Amine Ben Abdallah, Taha Yassine et al. · 0 citations
Review Open access Jul 2026

Grammatical Gender Patterns in Contemporary Arabic Medical Terminology: A Morphological Analysis of Arabic Newspaper Articles

The dynamics of the development of medical terminology in Arabic raise grammatical issues related to gender, specifically the distinction between masculine and feminine. The difference between morphological forms and syntactic behaviour leads to variations in gender usage in contemporary Arabic media texts. This trend marks the transformation of the lexical system, necessitating a review of the gender classification framework. Therefore, a systematic morphological and syntactic analysis is needed to explain the consistency and function of grammatical gender in modern scientific terms. This research aims to categorize the types of masculine and feminine gender in medical terms at the morphological and syntactic levels, and to describe the characteristics of gender classification in Arabic, particularly in loanwords, through a corpus-based approach that combines morphological and syntactic analysis within a unified analytical framework. This research employs a qualitative descriptive approach grounded in corpus linguistics. Data were collected from the newspapers An-Nahār and Akhbārul Yaum, then analyzed using the observation method, specifically the free conversation observation technique and the note-taking technique. Data analysis was conducted using the intralingual equivalence method and the distributional method with a top-down technique. The entire process of data collection and analysis is supported by Sketch Engine as a web-based linguistic corpus software, enabling the empirical mapping of gender usage patterns in modern Arabic journalistic discourse. The research findings indicate that morphologically, the types of mudzakkar are haqiqi, majazi, lafdzi, ma’nawi, and dzati, while the types of mu`annats are haqiqi, majazi, lafdzi, ma’nawi, dzati, and sima’i. The classification at the morphological level is influenced by biological sex, the presence of gender markers, gender forms in meaning, and language conventions. Meanwhile, at the syntactic level, the types of mudzakkar and mu`annats are divided into two types, namely ta’wili and hukmi, which are determined by lexical collocation. This research enriches the study of modern Arabic grammar by providing a grounded understanding of noun gender based on corpus data from medical terminology. In addition, the findings of this research can serve as a reference for Arabic language enthusiasts and speakers in consistently and accurately determining the gender of modern terms. This research also has implications for language planning and the compilation of terminology dictionaries based on empirical evidence.

A. A. Musthofa, Sangidu Sangidu, Hendrokumoro Hendrokumoro et al. · 0 citations