Generating metamodel-conforming instance models is a recurring task in Model-Driven Engineering (MDE), yet it remains tedious and tool-bound: instances are serialized as verbose, deeply nested XMI that today’s large language models (LLMs) cannot reliably produce or edit. We present AMIGO, a general, harness-agnostic ar...
Maximilian Hummel, Julian Roßkothen, Nathan Hagel et al.· 0 citations
This work investigates whether relational specifications can ground SLM behaviour by mathematical proof rather than statistical regularities alone, and gives practitioners a route to deploy verifiable SLMs in safety-critical settings.
Together, these results suggest a practical division of labor: ALM rubrics are effective for broad perceptual qualities that are difficult to formalize, while specialized objective rewards remain preferable when reliable measurements are available.
Ping Wang, Guang Yang, Shao-Rong Su et al.· 0 citations
Conformal Interval-Driven Self-Evolution (CISE) is proposed, which constructs candidate-specific reward intervals using conditional conformal inference and iteration-wise online density-ratio estimation and returns candidates only when all required property intervals lie entirely within their respective feasible region...
Kangjun Noh, Soyu Kim, Kyungwoo Song· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Video multimodal large language models (MLLMs) keep climbing video question answering benchmarks, yet shuffling the frames, masking the segment that supports the answer, or occluding the target object barely changes their predictions. The accuracy rests on appearance and language priors, not on the temporal evidence th...
Zhaolu Kang, Shi-Yu Liu, Tai-Long Luo et al.· 0 citations
🌟 Summary Version 8.4.173 adds native AMD Xilinx export for Versal AI Edge Series Gen 2 NPUs, opening a new path to deploy Ultralytics models on AMD edge hardware. 📊 Key Changes 🚀 New xilinx export format — Export models through AMD Quark to a quantized ONNX model and Vitis AI configuration. The format also accepts...
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23132634. Thank you for this paper. You took 30 Korean-language artifacts (code modules, tutorials and presentation scripts, generated by Claude Opus 4.6) with 150 planted errors, a...
Evgeny V. Arsentyev· Zenodo (CERN European Organi...· 0 citations
Big J v1.0 is a math engine with an English front end. It turns English into 27 numbers, runs one fused calculation on them, and answers in plain English built only from those numbers. No language model writes any reply, and the same input always gives the same answer. The fused calculation (one loop, 2000 steps): the...
James Jardine· Zenodo (CERN European Organi...· 0 citations
AIproposesCodeVerifies is a local tool that checks whether a text rewritten or translated by an AI assistant preserves the data the user has chosen to protect. Principle: the AI proposes, the code verifies, the person decides. The tool compares the original text with the AI's proposal and detects certain classes of cha...
Jose Ranero García· Zenodo (CERN European Organi...· 0 citations
This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23129468. Summary This paper attacks what it calls the compliance fiction: the industry practice of treating regulatory conformity as a binary verdict declared at deployment time, w...
Karmendra Pandey· Zenodo (CERN European Organi...· 0 citations
User-generated Uzbek texts — social media comments, customer reviews and citizens’ appeals to e-government portals — contain a high proportion of spelling errors that degrade the quality of downstream social-media monitoring tasks. The paper formulates spelling correction for Uzbek as a noisy-channel decoding problem i...
Sevinch Kenjayeva· Zenodo (CERN European Organi...· 0 citations
Operational evidence can contain the facts needed to diagnose a failure and identifiers or temporaldetails that need not be disclosed to an external model. We study diagnostic-invariant privacy-oriented transformation: changing the representation sent to a large language model while preservingthe relationships required...