Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Abstract
Rhetorix: Distance-Aware Ordinal Learning for Classical Arabic Rhetoric Official open-source release of the Rhetorix computational framework, accompanying the research paper: "Distance-Aware Ordinal Learning for Processing-Effort Classification in Classical Arabic Rhetorical Texts" 📌 Overview Rhetorix is an end-to-end computational framework designed for predicting graded processing effort in Classical Arabic figurative discourse. It addresses the challenge of modeling ordered linguistic constructs under severe class imbalance by introducing a task-specific distance-aware optimization objective directly into transformer fine-tuning. 🚀 Key Features & Components Distance-Aware Ordinal Loss ($\mathcal{L}_{Ordinal}$): A unified PyTorch objective combining weighted Categorical Cross-Entropy ($L_{CE}$) with an expected metric distance penalty ($L_{Dist}$): $$\mathcal{L}{Ordinal} = L{CE} + \lambda L_{Dist}$$ Built-in cost-sensitive inverse class weighting to counteract severe empirical skew without altering input representations. Pretrained Transformer Pipelines: Ready-to-use fine-tuning and evaluation modules across five architectures: ARBERT (Large-scale Modern Standard Arabic) MARBERTv2 (Web-scale Dialectal Arabic) CAMeLBERT-CA (Classical Arabic / Heritage texts) AraBERTv2 (Standard Formal Arabic baseline) XLM-RoBERTa (Multilingual cross-lingual baseline) Robust & Leak-Free Validation Suite: Nested Stratified 5-Fold Cross-Validation with source-level grouping. Internal-validation checkpoint selection optimizing Quadratic Weighted Kappa (QWK). Strict partition isolation: weights and hyperparameters are computed solely from inner training data. Comprehensive Evaluation Metrics: Automated calculation of chance-corrected agreement (QWK), error magnitude (MAE), class-balanced recovery (Macro-F1), and nominal Accuracy. Error severity profiling tools distinguishing adjacent misclassifications (Severity-1) from extreme endpoint deviations (Severity-2). Non-parametric Wilcoxon signed-rank diagnostic test for paired instance-level error analysis. 📂 Dataset: RhetoriScale ($N = 965$) An expert-annotated corpus of Classical Arabic figurative expressions categorized into three ordered inferential levels: Low (0): 45 instances (4.66%) Medium (1): 764 instances (79.17%) High (2): 156 instances (16.17%) Fully pre-partitioned into leak-free cross-validation folds. Software Stack: Python 3.10+ | PyTorch 2.x | Hugging Face Transformers
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The results show that the embedded industry has been able to apply agile methods in its development processes and that the appreciation of the agile methods and their individual practices appears to increase once adopted and applied in practice.
O. Salo, P. Abrahamsson· IET Software· 238 citations· ⚡9
Consequences of happiness and unhappiness that are beneficial and detrimental for developers' mental well-being, the software development process, and the produced artifacts are found.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· Journal of Systems and Softw...· 236 citations· ⚡13
The Mobile-D approach is briefly outlined here and the experiences gained from four case studies are discussed, which helped develop an agile development approach for mobile application development.
P. Abrahamsson, Antti Hanhineva, H. Hulkko et al.· Conference on Object-Oriente...· 225 citations· ⚡18
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 6, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.