Learning with class-conditional label noise often relies on a transition model from latent clean classes to observed annotations. Forward correction embeds this transition in the likelihood, yet finite-sample networks may still memorize corrupted labels. The corrected likelihood also induces a reverse posterior over the clean classes that could explain each annotation. When its leading class differs from the annotation, the model and transition matrix provide evidence against that annotation, but the leading alternatives can remain nearly tied. We introduce Selective Posterior Margin Regularization (SPMR), which preserves the Forward objective and converts this disagreement into a graded update on the clean classifier. SPMR selects the leading reverse-posterior class, scales a detached pairwise margin by the separation between the two leading posterior classes, and assigns correspondingly little influence to diffuse conflicts. The gap factorizes into transition- adjusted pairwise separation and the posterior mass carried by the leading pair. The active margin follows the locally minimum-norm logit direction that enlarges the selected pairwise margin. Across five known-transition benchmarks, SPMR improves full-length Forward by 2.5-7.0 percentage points and remains 0.7-2.5 percentage points above Forward with Mixup and early stopping. Matched interventions support distinct gains from the posterior-space coefficient, transition-adjusted target, and pairwise action. The same design transfers to estimated transitions, human annotations, architectural changes, and stronger Forward recipes. The formulation uses latent-class evidence already available inside Forward correction without promoting every posterior conflict to a corrected label.
Ze-Xing Zhang, Ji-Chao Li, Tian-Yang Lei et al.· 0 citations
ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue, is introduced and two tasks are defined: hallucination detection and minimal evidence set localization.
Zexing Zhang, Tianyang Lei, Kewei Yang et al.· Proceedings of the 32nd ACM...· 0 citations
Cross-phase collaborative management in complex product systems (CoPS) development is inherently challenged by heterogeneous organizational coupling and stochastic disturbances. Although digital twin (DT) and model-based systems engineering (MBSE) technologies provide foundations for physical–virtual synchronization and model traceability, existing approaches remain fragmented in three respects: insufficient requirements-traceable architectural integration, limited cross-phase semantic interoperability and runtime evolution, and weak operational links between semantic reasoning and adaptive decision models. To address these gaps, this paper proposes an MBSE-driven digital twin framework with semantic enhancement. First, a four-layer architecture is derived using the MagicGrid methodology, encompassing physical–virtual mapping, semantic reasoning, decision support, and service interaction. Second, a collaboration-oriented SysML profile is developed to standardize the representation of tasks, resources, materials, disturbances, and management constraints across engineering phases. Third, a knowledge-driven adaptive collaboration mechanism maps runtime disturbance inputs into semantic states, propagates their cross-phase impacts, supports process-topology reconfiguration, and generates decision-ready constraints for adaptive management. A case-based prototype for aero-engine turbofan blade development demonstrates the feasibility of the mapping–reasoning–decision chain and provides case-level evidence of improved cross-phase coordination under controlled disturbance scenarios. The results indicate a feasible engineering pathway from perceptive DT functions toward reasoning-enabled collaborative decision support.
Zhu Xiang, Minghao Li, Tianyang Lei et al.· Systems· 0 citations
Large language models (LLMs) and large multimodal models (LMMs) are increasingly applied in scientific dialogue, but it remains unclear whether they can reliably ground specific dialogue statements to paper-based evidence. A central challenge is paper-grounded hallucination under a paper-as-truth setting: statements that are contradicted by, not found in, or otherwise not decidable from the paper PDF. These hallucinations can be caused by both human misinterpretations and model-generated assertions, ultimately undermining the efficiency, fairness, and credibility of scientific dialogue. Existing benchmarks often overlook this issue, focusing either on subjective macro-level quality assessments or lacking cross-modal evidence localization. We introduce ADAM-Bench (Auditing Dialogue Assertions with Multimodal Evidence), a benchmark for paper-grounded hallucinations in scientific dialogue. Starting from around 27,000 papers, ADAM-Bench is a multi-layer benchmark with three tiers: Scale, Core, and Gold. ADAM-Bench pairs approximately 1 million atomic claims with over 7 million multimodal evidence objects extracted from the corresponding PDFs. We build it through a four-stage pipeline of claim atomization, candidate evidence recall, model-assisted pre-alignment, and human verification. Based on this dataset, we define two tasks: hallucination detection and minimal evidence set localization. Additionally, to avoid the brittleness introduced by single-rationale supervision, we formalize minimal evidence as a set of equivalent evidence sets and evaluate localization by best-matching against multiple gold evidence sets. We conduct a comprehensive benchmark of 34 LLMs and 10 LMMs, spanning large proprietary models (Claude-Opus-4-6, GPT-5.2) and open-source models (Qwen3-235B, GLM-4.6V 106B). Results are markedly low (25.2%--51.1%), indicating that grounding conversational hallucinations in real multimodal papers remains far from solved. We hope this benchmark will contribute to building scientific assistants that make calibrated judgments, cite minimal, auditable evidence, and mitigate the impact of hallucinations in scientific discovery evaluation.
Zexing Zhang, Tianyang Lei, Kewei Yang et al.· Proceedings of the 32nd ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.