Aug 2026· Linguistics Vanguard· 0 citations· 19 references
Abstract
Abstract This study investigates how nonexperts conceptualize linguistics and what they most want to know about language. We recruited 812 English-speaking participants from Canada, the United States, and the United Kingdom via Prolific to complete four open-ended surveys (N = 406 on perceived professional foci; N = 406 on desired research topics). All responses were segmented into 2,684 statements and analyzed with unsupervised natural language processing: sentence embeddings (all-MiniLM-L6-v2) followed by k-means clustering, yielding 29 topic clusters. Results show that respondents broadly recognize core subfields (e.g., syntax, phonetics/phonology, morphology, semantics, historical linguistics) and also reproduce familiar misconceptions (e.g., linguists as translators; linguists as language or literature teachers). Public curiosity, however, concentrates on applied and experiential themes – language learning and learnability, bilingualism, dialect differences, and the origins, spread, and loss of languages. Topics linking variation, cognition, and development (sociolinguistics, acquisition, psycho-/neurolinguistics) form shared ground. We discuss how these high-interest domains can serve as entry points for communicating foundational linguistic concepts.
This paper examines the types of English-language questions posed by learners on HiNative, a global peer-to-peer language learning platform. While online question-and-answer platforms have been widely adopted for language learning, the linguistic focus and distribution of learner-generated questions in self-access digital environments remain underexplored. Drawing on a linguistic content analysis approach informed by learner autonomy, the study systematically classifies learners’ questions across key linguistic domains. Using qualitative content analysis, a dataset of 783 learner-generated inquiries was categorized into key linguistic domains (grammar, vocabulary, usage, pronunciation, idioms, translation, comparison, formality/contextual use, and other). The findings reveal a predominant focus on meaning-oriented questions, particularly translation (34.2% of all questions) and vocabulary (21.5%), while grammar and pronunciation queries were comparatively rare. This distribution suggests that learners prioritize communicative meaning, lexical expansion, and pragmatic appropriateness over formal rule-focused learning. Such inquiry patterns highlight learners’ reliance on their first language as a bridge to the target language and their pursuit of precise and natural usage in authentic contexts. These results contribute to second language acquisition research by offering insights into learner priorities in informal, technology-mediated settings, with implications for the design of adaptive learning resources and learner support systems.
R. Alsharif· Theory and Practice in Langu...· 0 citations
The Cross-Lingual Comprehension Gap (CLCG) is defined as the reduction in response quality when the same content and question are presented in a target language rather than in English.
Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.
Recently, contact linguistics has become increasingly interested in multiword units. At the same time, the code-copying framework (CCF) includes the notion of mixed copies (MCs) that are in-between global copies (‘borrowing’) and selective copies (‘structural change’) and illustrate the transition between the lexicon and grammar. The research question is: What types of MCs occur in English-Estonian bilingual speech?
The data were transcribed, and MCs identified, annotated, and classified according to their structure. English items were searched for in Estonian dictionaries to establish their Estonian equivalents or conventionalization of such items. The frequencies of MCs and their Estonian equivalents were also searched on Google to determine whether the MCs occur outside the corpus.
Three datasets were analysed: written texts from 44 blogs (385,124 tokens), spoken data from 10 vlogs (117,555 tokens), and 8 podcasts (77,277 tokens). Quantitative analyses of the various MC types were conducted, followed by a qualitative analysis of representative examples.
Compound nouns constitute the majority of MCs, followed by idioms, phrasal compounds, and a small number of compound verbs. No frame-changing MCs (i.e., MCs resulting in grammatical change) were attested. Since compound nouns and analytic verbs occur in both languages, structural similarity may be a facilitating factor in copying.
The notion of MCs is not widely used. Research typically focuses on particular types of items (e.g., compound nouns or verbs); here, however, the question is reversed: which types of items yield MCs?
It was established that the proportion of MCs in the data is comparable to that of selective copies. Within MCs, the globally copied element renders the remaining part more specific, highlighting the importance of meaning in contact-induced language change. MCs are also present on the Estonian internet and, in some cases, outnumber their Estonian equivalents, if such equivalents exist.
A. Verschik, H. Kask· International Journal of Bil...· 0 citations
Comparison of five widely used large language models suggests that AI-generated language may shape how culturally situated perspectives are expressed, with differences across models indicating that AI-generated language may shape how culturally situated perspectives are expressed.
Ashkan Goudarzi, Aylar Naderi Zonouz· Digital Studies in Language...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.