Skip to content
Open access

The Lexical Analysis of Postgraduate Artificial Intelligence Academic Texts

TL;DR

The findings demonstrate that the lexical demands of authentic postgraduate AI readings are more nuanced than estimates of vocabulary coverage alone imply and cannot be generalised even within a single discipline.

Abstract

This thesis investigates the vocabulary and incidental vocabulary learning opportunities in authentic academic texts for postgraduate students of Artificial Intelligence. Book chapters and journal articles from five courses for taught masters of Artificial Intelligence at Victoria University of Wellington were collected to compile the corpus of Artificial Intelligence reading texts (CAIRT). The thesis consists of three studies. The first study investigates the vocabulary profile of texts in CAIRT, assessing how much vocabulary is needed to reach 95% and 98% coverage using Nation’s (2012) 25 BNC/COCA word lists with five supplementary lists. The findings show that 4,000 and 6,000 word families plus supplementary lists are needed to reach 95% and 98% coverage, respectively. However, lexical demands vary across courses and text types, indicating that the vocabulary load cannot be generalised even within a single academic discipline. Additionally, the first 3,000 word families account for the largest proportion of the corpus, while mid-frequency word families and supplementary-list words make comparable contributions, with low-frequency word families accounting for the smallest proportion. The second study explores the repetition and distribution of word families from each category (high-, mid-, low-frequency, and supplementary lists). Although high-frequency word families are most likely to recur, the majority of word families occur only a small number of times. While many word families are shared across courses and trimesters, a substantial proportion are restricted to individual sub-corpora, indicating uneven opportunities for incidental vocabulary learning across the texts and courses. The third study investigates which words are elaborated within texts that may facilitate vocabulary learning and reading comprehension. Using Hyland’s (2005) taxonomy of code glosses, 57 single words and 188 multiword units (MWUs) are identified as elaborated within two courses. The elaborated single words include high-, mid-, and low-frequency words, abbreviations, proper nouns, transparent compounds, and other words outside the BNC COCA 25,000 word families. The corpus frequency rate of these elaborated words varies, and their distribution is inconsistent across texts. Many of these elaborated words occur in only one text, and a small number of them recur across multiple texts. Five single words and seven of the MWUs are elaborated multiple times, and most of the elaboration occurs within a single text. Overall, the findings demonstrate that the lexical demands of authentic postgraduate AI readings are more nuanced than estimates of vocabulary coverage alone imply and cannot be generalised even within a single discipline. Although authentic disciplinary texts repeatedly recycle a relatively small core vocabulary, they provide uneven opportunities for incidental vocabulary learning through repetition and lexical elaboration. Compared with the adapted reading materials and authentic texts used in experimental studies of incidental vocabulary learning through reading, authentic AI readings provide these lexical conditions less consistently, as their primary purpose is to communicate disciplinary knowledge rather than to facilitate vocabulary learning. These findings contribute to a more contextualised understanding of incidental vocabulary learning in authentic disciplinary reading and have implications for vocabulary profiling, course sequencing, and pedagogical support for postgraduate AI students.

Read PDF

Similar papers

Open access Aug 2026

Exploring Lexical Density in English Reading Passages for Senior Vocational High School Students

This study explores the lexical density of an English reading text designed for 11th-grade vocational high school students, with the aim of evaluating its suitability for educational use. Lexical density refers to the ratio of content words to the total number of words in a text, which serves as an indicator of text co...

Tiara Putri Aishwaryadi, Rahma Dania, Rosi Kumala Sari · 0 citations

Artificial Intelligence and academic writing in higher education: a comparative analysis of linguistic indicators in student texts

The results revealed relevant differences between the analyzed papers: texts produced in 2023 showed greater stylistic variation, the presence of authorial markers, and irregularities typical of human writing, whereas texts from 2025 presented a higher concentration of indicators associated with linguistic standardizat...

G. D. da Silva, I. Rhuan, G. Tardo et al. · 0 citations
Open access Jul 2026

Lexical diversity and CEFR vocabulary in critical academic reading passages: A conceptual replication

Background and Purpose: Critical academic reading skills demand learners to engage in higher-order reading skills which require advanced lexical knowledge. The two dimensions of lexical knowledge, namely lexical diversity and vocabulary variety, however, are mostly assessed in isolation using unstable measurement instr...

Anealka Aziz, Tuan Sarifah Aini Syed Ahmad, Suryani Awang et al. · 0 citations
Open access Jul 2026

A Computational Corpus Study of English Varieties: Comparing Native and Non-Native Academic Discourse

This study examines similarities and differences between native and non-native academic English through a computational corpus-based framework to identify variation in lexical choices, lexical bundles, grammatical patterns, collocations, and academic stance.

N. Akhter, Fatima Khan, Hammad Malik et al. · 0 citations
Open access Aug 2026

Lexical Richness and Academic Writing Performance in a Thai University Academic English Test

Although lexical richness is an important feature of L2 writing quality, research on its predictive ability in university entrance tests, particularly in the Thai EFL context, is limited. Following Read’s (2000) framework, this study analyzes the relationship between lexical richness and academic writing performance in...

Zi-Xuan Hu, Chawin Srisawat, K. Poonpon · 0 citations
Review Open access Jul 2026

Artificial intelligence tools for the development of writing skills in English language learning

This research analyzes English teachers’ perceptions of the use of artificial intelligence tools as pedagogical support for developing writing skills. It also examines the changes observed in the written production of A2 English as a Foreign Language students after using ChatGPT and Grammarly. The study involved 19 fir...

María Valentina Loor Santos, Henry Xavier Mendoza Ponce · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.