Cost-driven interest in running large language models (LLMs) on non-mainstream silicon outpaces the maturity of the surrounding software stacks. This paper uses one such platform, the AMD BC-250 (a repurposed cryptocurrency-mining board with a GFX1013 "Cyan Skillfish" accelerated processing unit (APU), 16 GB of unified...
Artur Andrzejczak· Zenodo (CERN European Organi...· 0 citations
Fine-grained image editing requires more than producing a visually plausible result: an editor must execute the requested attribute change precisely while leaving everything else intact. However, existing benchmarks leave a critical gap between realism and verifiability: benchmarks built on realistic images typically r...
Mu-Yao Wang, Chen Zhu, Shi-Qi Yang et al.· 0 citations
Vision-language models (VLMs) achieve strong zero-shot transferability but remain vulnerable to target-domain shifts at inference time. Test-time adaptation (TTA) offers a practical remedy, yet most existing VLM-TTA methods follow a prediction-side adaptation paradigm. They use test samples to adjust logits, prototypes...
Masked diffusion language models (dLLMs) generate text by iteratively denoising masked positions, re-predicting each token multiple times before it is committed. An autoregressive decoder exposes an answer's distribution once, at the step that commits it; a dLLM exposes it at every denoising step before commitment, and...
Sarim Hashmi, Mukul Ranjan, Abdelrahman W. A. Elsayed et al.· 0 citations
Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as"On the table."is usually produced as four separate predictions for the preposition (On), article (the), noun (table),...
On-policy distillation (OPD) has emerged as an effective post-training paradigm for language models, with recent efforts extending it to block diffusion language models (BDLMs). However, existing studies focus almost exclusively on small block sizes, leaving distillation into student models with larger blocks underexpl...
Zai-Quan Yang, Fei Wei, Yong Wang et al.· 0 citations
When documents supporting an agent's derived facts are revoked or replaced, should it repair memory or re-read current evidence? We introduce an evidence-revision evaluation on medication- and problem-list tasks from public ICU records. Under revocation, replacement and control events, we compare full and source-filter...
Reinforcement learning post-training for language models relies on two reward designs: human preferences (RLHF, DPO) and binary verifiers (RLVR). Clinical question answering fits neither. Near-correct answers differ by a single substituted entity, and no executable check decides clinical correctness. We instantiate a s...
CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when sco...
Guang-Yuan Li, Tian-Ming Du, Yan Jiang et al.· 0 citations
Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M par...
Jaehoon Lee, Harry B. Partridge, M. Jayasekara et al.· 0 citations
Generative large language models (LLMs) inherit undesirable behaviors from pre-training, including demographic bias and toxic generation, that often emerge only after deployment and affect a small subset of inputs. A repair should eliminate the identified defect, preserve the model's overall functionality and, ideally,...
Hsin-Ling Hsu, Min-Yue Chen, Nai-Chia Chen et al.· 0 citations