LLMs have high diagnostic accuracy when challenged with published nephrological case scenarios, and the interpretation is limited, however, because previous exposure of the models to individual vignettes, or parts thereof, cannot be ruled out.
Chalid Hasan, Göran R. Boeckel, S. Siam et al.· Deutsches Ärzteblatt Interna...· 0 citations
Abstract Data combining field observations with information from farm sensors, remote sensing, weather services, agricultural machinery, mobile apps, and e-commerce are increasingly underpinning new, intelligent ways of supporting agriculture decisions. This paper reviews the evidence for improvements in farm productiv...
Abstract Data combining field observations with information from farm sensors, remote sensing, weather services, agricultural machinery, mobile apps, and e-commerce are increasingly underpinning new, intelligent ways of supporting agriculture decisions. This paper reviews the evidence for improvements in farm productiv...
Abstract Four small language models of 124M–1.1B parameters were trained on the same five-task stream with five configurations: full fine-tuning, LoRA (r=8 and 16), bottleneck adapters, and LoRA r=8 with 2% replay. We evaluate task accuracy, backward transfer, forgetting, forward transfer, and learning plasticity. Acro...
Vedant Rajendra Vare, Jatin Nayak, Mihir Jadhav et al.· Zenodo (CERN European Organi...· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Abstract Four small language models of 124M–1.1B parameters were trained on the same five-task stream with five configurations: full fine-tuning, LoRA (r=8 and 16), bottleneck adapters, and LoRA r=8 with 2% replay. We evaluate task accuracy, backward transfer, forgetting, forward transfer, and learning plasticity. Acro...
Vedant Rajendra Vare, Jatin Nayak, Mihir Jadhav et al.· Zenodo (CERN European Organi...· 0 citations
Understanding gender biases in large language models (LLMs) is increasingly important as these systems become embedded in decision-support tools with real consequences. Prior research has focused only on a small set of models, leaving open the extent to which gender biases are common and heterogeneous across LLMs. We a...
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have v...
Zhen-Yu Wang, Tian-Ze Wang, Lin-Jun Zhang et al.· 0 citations
MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all is introduced.
Nagham Omar, Mahmoud Jabarin, Kinan Ibraheem et al.· 0 citations
This paper revisits the standard state-only formulation of value estimation and proposes PPO, a self-privileged actor-critic framework that consistently improves value-estimation quality by a substantial margin and outperforms representative actor-critic and critic-free RLVR baselines on challenging mathematical reason...
Kun Liang, Chenming Tang, Clive Bai et al.· 0 citations
Jev answers MMLU's calculation-heavy questions more accurately than other MMLU questions (94% vs. 91%), whereas both open models, and all three on C-Eval, find them harder.
Tobias Deußer, L. Sparrenberg, R. Sifa· 4 citations
Speculative decoding accelerates autoregressive generation by using a smaller drafter to propose tokens for batched verification by a larger target. However, conventional speculative decoding couples drafting to the target's evolving verified prefix, serializing drafting and verification. We ask whether this dependency...
Yun-Zhe Li, Kyoungjun Park, Hong-Zi Zhu et al.· 0 citations
Test-time compute scaling has emerged as a cornerstone of advanced machine reasoning, yet performing iterative deliberation directly within continuous latent representation spaces reveals a catastrophic pathology: the Deliberation Drift Cliff. While unconstrained recurrent latent models achieve initial reasoning gains...