LifeSciBench is introduced, a benchmark of 750 expert-authored tasks designed to evaluate whether language models can handle realistic life science research work, with each constituent task paired with a human expert-written rubric.
Amelia Liu, Andrew Ho, Anne Marie Droste et al.· bioRxiv· 2 citations
A rapidly advancing precision-therapy pipeline-including antisense oligonucleotides to upregulate the intact allele, AAV-based gene replacement, CRISPR-mediated transcriptional activation, epigenetic modulators, and rational pathway-targeted small molecules-offers realistic prospects for disease modification.
TestifAI, a deep learning testing framework for efficient and accurate estimation of robustness against combinations of perturbations, is proposed and partial model tomography is introduced, a novel approach to reconstructing model behaviour in a multi-perturbation space from tests that apply only a small number of perturbations.
Arooj Arif, T. Hartung, E. Botoeva et al.· 1 citation
This paper test representative models regarding their knowledge of Chinese and Japanese Buddhist history on consumer hardware and finds that, in Spring 2026, the Qwen series emerges as a winner for models in the 30B range, while the larger Kimi2.5 models lead in a cloud-based setup.
M. Bingenheimer· Yin-Cheng Journal of Contemp...· 0 citations
This study presents the first systematic literature review of research on Estonian Sign Language (EVK) published between 1985 and 2024.
Using Okoli's eight-step systematic literature review (SLR) model, we identified and screened 52 sources. Twenty-one were excluded on the basis of predefined criteria, leaving 31 studies for analysis. We examined how data-collection methods, annotation practices, and analytical frameworks have shaped the development of EVK linguistics.
The findings reveal a gradual shift from descriptive documentation toward more empirical and methodologically reflective research. Many studies reflect the methodological conditions of their time, including a reliance on elicited rather than naturalistic data, non-standardized transcription practices until 2006, and analytical frameworks shaped by earlier linguistic traditions of the 1960s–2000s. Across the reviewed literature, methodological conditions varied considerably, reflecting different historical phases in which formal training opportunities in linguistics, research teams with EVK proficiency and corpus-based infrastructures were not yet consistently available. Only a small number of studies reported the use of ELAN annotation or systematically documented participatory practices involving Deaf collaborators or heritage signers.
The review highlights significant progress in methodological awareness but underscores persistent gaps in the Estonian Sign Language research landscape, linguistic infrastructure and involvement of Deaf researchers. By tracing EVK research across historical phases, the review identifies priorities that could inform a collaborative, ethically grounded and corpus-based research agenda. Such an agenda could support cooperation among universities, Deaf community organisations, researchers, educators and language-policy actors, while enhancing data transparency, reproducibility and the sustainable development of EVK research.
Jari Pärgma, Christian Rathmann, Péter Zalán Herbszt-Romanek· Frontiers in Communication· 0 citations
This work presents the first complete system for automated six degrees of freedom (6DOF) satellite pose estimation from spatially resolved, ground-based, adaptive optics (AO)-corrected imagery, addressing a key challenge in Space Domain Awareness (SDA). The approach mitigates the need for human labeling by directly regressing satellite orientation and position from blurry, noisy, and deeply shadowed imagery. A multi-stage deep neural network pipeline localizes the satellite, predicts pose, and optionally applies temporal filtering. Networks are trained exclusively on fully synthetic imagery generated from a CAD model, yet generalize effectively to real data, bridging the Sim2Real domain gap. On 137 real, human-labeled test images of Seasat, the model achieved a mean rotation error of 5° and a mean image-plane translation error of 21 cm. Slant range error was quantitatively evaluated on synthetic data due to unknown real-sensor parameters. Qualitative evaluation of additional real Seasat imagery rated 177 of 199 predicted poses as “ground truth equivalent” or “high-confidence match,” with zero catastrophic failures. The system was extended to seven degrees of freedom (7DOF) for satellites with articulating components and demonstrated on real Hubble Space Telescope (HST) imagery, achieving 5.5° rotation error, 51 cm image-plane translation error, and 8° symmetry-adjusted solar array error on a 249-frame pass with causal temporal filtering. Across 586 real test images from Seasat and HST (captured over multiple decades under diverse conditions) the system consistently performed well. Full 6DOF performance was quantified on a high-fidelity wave optics (HFWO) synthetic test set of Seasat, where the model achieved 8.4° mean rotation error, 34 cm image-plane translation error, and 1.4% line-of-sight range error at r0=6 cm and 1031 km range. In a limited 200-image benchmark, the model demonstrated 48% lower mean rotation error than a single human labeler while operating ∼800× faster. It required <40 h and a single A100 GPU to generate data and train. The approach was also demonstrated for ARGOS, a smaller satellite with highly symmetric geometry. An exploratory General Image-Quality Equation-based image quality metric (AO-IQ) was introduced as an empirical correlate for pose accuracy. General-purpose models like GPT-4o and Depth Anything V2 failed across most SDA tasks, but rapid gains in vision-language models warrant continued monitoring. These results establish a new operational baseline for practical, real-time satellite pose estimation from AO SDA imagery.
Thomas J. Dickinson, Dawson Friesenhahn, Justin Fletcher et al.· Aerospace· 0 citations
Among the six evaluated model configurations, GPT-5.2 achieved significantly higher agreement with the clinician-derived composite reference than the other tested models for both ASA-PS and CCI in post hoc paired analyses with multiplicity correction.
Marco di Maio, G. Stopper, Vincenzo Di Matteo et al.· Bioengineering· 0 citations
It is argued that realizing deep learning’s full potential requires not only architectural innovation but also domain-aware representations that encode the statistical ensemble nature of polymers, evaluation protocols aligned with discovery scenarios, and physically grounded inductive biases.
Nassima Aleb· International Journal of Com...· 0 citations
Results align with a diagnostic perspective on chunking: using evidence at a task-appropriate level of granularity can improve grounding, auditability, and answer quality, but the observed patterns should be interpreted within the HotpotQA distractor setting, fixed generator, and tested context budgets.
IT help desks at large organizations face a high volume of recurrent, well-documented user requests that nevertheless require human-written replies, creating a persistent staff workload that is repetitive in content but non-trivial in tone and procedural correctness. We present FroLineR, short for Front-Line Response, a system that drafts the initial staff reply to such tickets in the login and account-activation category and integrates into a human-in-the-loop ticketing workflow on a Romanian-language ticketing platform. The generator is an unmodified instruct model augmented with retrieval from a small set of hand-curated guide documents, using a Romanian system prompt refined over several rounds of staff review. To evaluate and refine the prompt without manual labeling, we cluster the first user message of every historical thread with both BERTopic and Semantic Signal Separation (S3), score configurations along coherence and lexical-diversity axes, and extract a 200-message evaluation set from the winning model. Prompt convergence was certified by several rounds of manual review by support staff. The production system is quantized to Q4_K_M GGUF, served through llama-cpp-python behind a small Flask API, and deployed with GPU offloading on the target server, reducing end-to-end per-answer latency from approximately 830 s on the server’s CPU to roughly 61 s once layers are offloaded to the GPU, with no observable degradation in answer quality.
Alexandru Dima, Marian Mihailescu, Darius Mihai et al.· Applied Informatics· 0 citations
Self-reported prior IPE experience showed a clearer and more consistent association with perceived collaboration than did the single-item measure of self-rated IPE knowledge, whose association was small and less consistent across domains.
Viktorija Xharra, Rosario Caruso, Florian Spada et al.· Healthcare· 0 citations
Comparison of GPT-4, BERT (bidirectional encoder representations from transformers), Gemini, and DeepSeek large language models (LLM), focusing on architectures, training methodologies, and real-world applications reveals GPT-4 excels in natural language generation and complex reasoning, supporting up to 128K tokens with moderate latency and higher costs making it effective for conversational artificial intelligence (AI).
Kavish Sanghvi, Aparna S. Sharma, Surbhi Hooda· Computer Science and Informa...· 0 citations
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.