This paper concludes that machines may assist translators but that they will not, by principle, be able to reach an almost perfect level and object to the huge amount of money spent on software development for systems with that objective.
Abstract
This paper takes up the important questions whether (i) machines—by means of software—can translate texts, thereby imitating translations made by humans, and (ii) publications in the fields of scientific domains and engineering can be translated by machines to an almost perfect level. This paper examines some issues in this context, among them the question if software can be compared with the minds processed by human brains, if machines handle information or something else, if knowledge is found in software (including some deliberations on what knowledge is), and if the latest software developments are what they declare. This paper concludes that machines may assist translators but that they will not, by principle, be able to reach an almost perfect level. Accordingly the specific aim (I shall not label it objective) is to point to theoretical and analytic (parsing) obstacles to the idea that machines can produce acceptable translations in all contexts. I demonstrate that certain combinations of words will present systems with insurmountable challenges; challenges which also present human translators with questions that have almost unattainable answers. Therefore I object to the huge amount of money spent on software development for systems with that objective. I also point to the option of developing more relevant alternatives, including simpler solutions. If you ask for new insights—like you would do when perusing a research paper in order to update your professional frame of reference—there are, basically, no new insights in this theoretical scientific article. I just make it transparent what all linguists know, i.e., that in the field of machines handling language, less focus should be on what electronic systems can, or cannot, do and more focus on what details in language, and languages, will make certain kinds of handling particularly arduous; to the verge of being insurmountable.
Tokenization is a process that breaks down text into smaller units called tokens. It serves as the initial step in NLP for dissecting the text so that the machines can understand human languages. With the latest LLMs, tokenization is extremely crucial because it is at the basis of how text can be interpreted, kept, and produced. This paper covers the concept of tokenization, its role in AI language systems and the problem of token limits in modern models. LLMs have a fixed number of tokens they can handle. If we exceed those, for instance, in summarization, translation, and conversational AI, they can give only a part of the answer, forget the context, and be less accurate. The paper describes various tokenization techniques word-based, subword-based, and character-based and weighs the advantages and disadvantages of each in practical situations. The author(s) merges the theoretical part with the evaluation of the case study to demonstrate the impact of token limits on the performance of the model and the user experience. On top of that, the piece of writing comes up with some solutions to these issues such as prompt optimization, chunking, context management, and advanced compression techniques. The key findings show that correct token handling results in not only computing efficiency but also higher quality of responses and longer retained context in machine systems. In conclusion, the paper highlights the growing significance of adaptable tokenization techniques and renderable architectures for the continued development of intelligent language models. These insights aid in gaining a deeper understanding of how tokenization affects both the capabilities and the limitations of communication systems based on AI and at the same time offers hands-on tips to researchers, developers, and companies that use the latest NLP technologies.
Madhurima Kommuru· International Journal of Mac...· 0 citations
This paper was written during the course “Introduction to Linguistic Theories and Analysis,” offered by the PPGL – Graduate Program in Linguistics at UFSCAR – Federal University of São Carlos, in February 2026. During the course, we had the opportunity to develop a paper, following the conclusion of the course, with the aim of defining the semantic fields of the two terms mentioned, as well as examining their use within the field of linguistics. Methodologically, the text is structured as a literature review of works addressing science, linguistics, and language, drawn from both physical and digital collections. The primary reference consulted in this approach is Karl Popper’s, The Logic of Scientific Discovery (2013). As intended outcomes of this study, we hope to compile a set of arguments that will, in principle, serve to integrate some of the theoretical frameworks into our doctoral dissertation (in progress, UFSCAR, 2024-2028).
Marcelo Pessoa de Oliveira, Dirceu Cleber Conde· Revista de Estudos Interdisc...· 0 citations
In this paper, we are exclusively concerned with the part of grammar that deals with the structure of sentences. This is called syntax. Not only the grammatical units of language were explored but the division of selected sentences into constituents (units) was also analyzed. To achieve this feat, sentences were separated into words and finally, words were regrouped on the basis of relationship between them. This paper has gone further to explain how the (agent) or subject of a sentence is identified through grammatical units. The grammatical units were introduced on the hierarchical order [down-up]. The general syntactic framework we have adopted is inspired by the theories of language developed by Noam Chomsky. The choice of this study is based on the assumption that English and French are “closely related and well documented languages” (Tanja et all, 2010:110) and the two (duo) “constitute a minimal pair suitable for micro-comparison (Kayne, 2005). Most learners that were not always comfortable with the syntax and structure of the two languages would be familiar with their components and syntax of English and French languages armed with this work.
T. A. Balogun, G.S. Idowu· Eureka-Unilag· 0 citations
ABSTRACT This study is a qualitative bibliographic essay that aims to question and seek to understand the (non)existence of otherness and the (non)production of meaning in the processes of interaction between humans and machines. To this end, we conducted a synthesis that integrated the following topics: the dialogic conception of language, artificial intelligence (AI), machine learning (ML) and natural language processing (NLP). During our analysis, we raised questions that led us to an unstable ground, and we concluded that we cannot completely exclude the idea of meaning production by machines, given that we also cannot affirm the existence of otherness. Therefore, we foresee that for these topics to be explored accurately, it is necessary to decentralize the human and adopt a hybrid view of humans, society, and language.
Renan Monezi Lemes· Bakhtiniana: Revista de Estu...· 0 citations
Automatic translation systems, from CAT tools to MT, overwhelmingly treat translation as a sentence-by-sentence act. This paper asks whether LLMs can be moved beyond that paradigm through whole-document, corpus-informed translation. We present PAT (Pragmatic Auto-Translator), a RAG-based system that pairs user-configured specifications with context from a comparable corpus of authentic longform texts in U.S. English and Latin American Spanish, passing retrieved paragraph-, section-, and document-level examples to an LLM for whole-document generation. The goal is draft translation for professional verification: target texts reformulated to fit their Spanish-language context, where discourse organization, rhetorical style, and pragmatic norms differ meaningfully from English. We evaluated six automatic translations of essays on generative AI across three projects using a customized MQM typology, assessed by two trained evaluators working from U.S. English into LATAM and Mexican Spanish. Results show that a limited prompt produced no meaningful reformulation, and specifications and corpus-informed translations at times showed substantial reformulation, though not always to effect. We find that LLMs can be moved toward reformulation and away from the sentence-by-sentence paradigm, though more work is needed to improve the effectiveness of those reformulations. In this paper, we discuss considerations related to automatic translation system design, corpus construction, and translation quality evaluation methodology and results.
The article describes texts generated by artificial intelligence, with an emphasis on Ukrainian-language material, which remains understudied. The aim of the study was to identify specific identifying characteristics of AI texts by comparing their semantic, structural, and stylistic parameters with texts written by humans. The study combined systematic analysis and synthesis, a comparative approach, content analysis, and quantitative linguistic methods (Type–Token Ratio, syntactic complexity analysis by T-unit), as well as semantic modeling, fact verification, and semantic-stylistic analysis. It has been proven that Ukrainian-language AI texts formally meet the basic criteria of textuality (cohesion, coherence, articulation), which are implemented through the probabilistic combination of templates rather than the author’s cognitive and communicative activity. Typical markers of machine generation have been identified: template composition (introduction – main part (3–5 subtopics) – conclusion), homogeneous paragraphs of “average” length, predominance of direct word order, presence of passive constructions, excessive frequency of formal connectors, structural and lexical monotony, errors in word usage. Semantic analysis revealed a combination of formal correctness with factual “hallucinations,” low information density, a predominance of neutral style, superficial, statistically determined expressiveness, and emotional masking. It was concluded that texts generated by artificial intelligence constitute a separate linguistic phenomenon with its own set of characteristics, which requires a special typology and a flexible, updatable analysis methodology. The risks to linguistic norms, speech culture, and information security are emphasized, as well as the need to develop critical competence in users regarding the perception of AI content.
L. Kravets, Viktória Stefuca, N. Libak et al.· Computer and Decision Making...· 0 citations