2026· SemEval@ACL· pp. 2057-2064· 1 citation· 29 references
Computer Science
TL;DR
This system leverages LLMs both for extracting high-level aspects and to encode them with state-of-the-art generalist embedding models, and applies a series of embedding post-processing steps to fit the embedding space with a Mahalanobis-like diagonal metric.
Abstract
The Narrative Story Similarity and Narrative Representation Learning (NSNRL) task measures the narrative similarity between two stories based on three core aspects: the abstract theme, the course of action, and the outcomes. Our system leverages LLMs both for extracting high-level aspects and to encode them with state-of-the-art generalist embedding models. We then apply a series of embedding post-processing steps and learn to fit the embedding space with a Mahalanobis-like diagonal metric. We show that some of these techniques should not be applied universally, as they do not necessarily increase performance or overfit, depending on the base encoder. Our system outperforms the baseline only in Track B, ranking twelfth out of twenty-seven on the final leaderboard, while performing lower than the baseline accuracy in Track A.
MAJEPPA is presented, a self-supervised framework to learn piano performance representations that span the full skill spectrum, from beginner practice sessions to virtuoso concert recordings, and a suite of downstream tasks spanning quality assessment, competition ranking, mistake and technique classification are introduced.
Jin-Wen Zhou, Huan Zhang, Weixin Zhai et al.· 0 citations
Across OLMo-2, Llama-3.1, and Qwen-3, under both MEMIT and AlphaEdit and in batch and sequential regimes, Moir consistently extends preservation in the most vulnerable domains, suggesting that aligning the preservation distribution with the model's operative distribution is a key factor in non-destructive editing and that the model itself may be the most accessible source of that distribution for deployed systems.
Jea Kwon, Jiwon Kim, Dong-Kyum Kim et al.· arXiv.org· 0 citations
We construct a corpus of 1,262 verse--commentary (urai) pairs from five Classical Tamil source sections, ranging from technical grammatical prose to modern paraphrase, and ask what information representation learning can recover. We train recurrent and Transformer encoders, a Siamese-style pair-matching network, an mBART-style encoder--decoder, and a decoder-only language model. Each analysis is interpreted against an appropriate control on the same data. TF-IDF provides a strong no-training lexical retrieval baseline, alongside representation analyses and generation controls for the learned models. A fixed string containing the 25 most frequent commentary words scores higher on generation overlap than the decoder-only model. Canonical correlation reaches 1.000 on Gaussian noise at these sample sizes, token-F1 spans only about 0.02--0.20 on this corpus, and the encoder--decoder continues to lower training loss for sixteen epochs after validation loss has begun to rise. One narrow result remains: the decoder-only model prefers authentic word order in 107 of 112 minimal-pair comparisons (95.5%), but does not reproduce held-out commentary content. We release the extraction and evaluation protocol; redistribution of the source commentaries remains subject to permission.
Temporal reasoning benchmarks for Video-LLMs are often mediated by language, leaving room for linguistic shortcuts from option wording, answer correlations, or language priors. To reduce such shortcuts, we introduce TempCloze, a video cloze benchmark for evaluating visual temporal reasoning in Video-LLMs. Given the beginning and ending clips of a video, models must identify the true missing middle from four candidates. TempCloze contains 1,521 carefully filtered videos from seven sources, mainly long-take and egocentric videos. We construct same-source distractors along three dimensions: Semantic asks what event should happen, Alignment probes when it should occur, and Progression tests how it should unfold, while shared scenes and objects reduce appearance cues. Our evaluation of 10 proprietary and 21 open-source Video-LLMs reveals Alignment as the primary bottleneck: models often recognize plausible semantic content and local event progression but struggle with temporal alignment. We further conduct error pattern and behavioral sensitivity analyses on TempCloze-Mixed and TempCloze-Hard with four representative models to examine where errors arise and how candidate order, context direction, visible span, frame density, and test-time scaling influence model choices.
Wenqi Pei, Henry Hengyuan Zhao, Yi-Lai Liu et al.· 0 citations
Text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point, showing that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.
Dense embeddings are foundational to contemporary natural language processing, information retrieval, recommendation, and retrieval-augmented generation. Nevertheless, a single general-purpose vector typically superimposes multiple relations—topic, entailment, sentiment, part–whole structure, evidential role, temporality, and domain-specific constraints—within one geometry and one similarity function. This article presents layered semantic refinement (LSR) as a framework and evaluation protocol rather than as a single universally validated algorithm. LSR retains a general encoder as a transferable semantic substrate while explicit modules reorganize, augment, or index its representations for a defined operational objective. The framework covers learned projections, nonlinear adapters, graph message passing, hierarchical aggregation, deterministic rule channels, clustering, multi-view representations, and prompt-conditioned embeddings.
Pedro Emílio Amador Salomão· Nexus Science Review· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.