Skip to content
Conference

Meaningful Human-in-the-Loop Checking of GenAI Synthesis for Restricted Languages

2026 · European Conference on Object-Oriented Programming · pp. 22:1-22:31 · 0 citations · 93 references
Computer Science

TL;DR

Pick is a significant improvement over showing users the candidate expressions, and also helps catch situations where no output is a match, and also helps catch situations where no output is a match.

View source

Similar papers

Review Aug 2026

Refine After Generation: Toward Correct and Concise Patches in LLM-based Program Repair

This paper identifies patch verbosity as a major yet overlooked concern in LLM-based APR and proposes RECAP, a lightweight, plug-and-play adapter that attaches to existing repair frameworks after generation that achieves a substantially better size-correctness tradeoff.

Wen-Qiang Luo, J. Keung, Xiaoyu Shi et al. · 0 citations
Preprint Aug 2026

Unreliable in Practice? A Comprehensive Study of Errors in LLM-Generated Code

It is observed that generated code often omits basic input validation or memory-safety checks, which can lead to overflows, resource exhaustion, or other reliability/security issues, and even the largest models frequently make simple mistakes.

Rodrigo Pato Nogueira, Marco Vieira, João R. Campos · 0 citations
Book Open access Aug 2026

Discovery, Validation and Editing of Large Language Models Mechanisms: Recent Advances and Future Perspectives

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks, yet their internal mechanisms remain largely opaque, making it difficult to understand, predict, or control their behavior. As LLMs are increasingly deployed in high-stakes settings, this lack of transparency raises ser...

Yinhan He, Wendy Zheng, Tianyi Zhao et al. · 0 citations
Preprint Aug 2026

Evaluating Language Models on Cross-Language Code Functional Equivalence

This work investigates whether LLMs can accurately judge functional equivalence across different programming languages in human-written code, a setting that requires deeper reasoning beyond superficial similarity, and identifies a difficulty-dependent breakdown in equivalence judgment.

Hui Sun, Anderson G. Uchôa, Rohit Gheyi et al. · 0 citations
2026

Panbench: A Comparative Benchmarking Tool for Dependently-Typed Languages

This work creates an “over language” in which to express all the information the authors need to be able to output correct and idiomatic syntax for each of their targets, and details the design of the extensible system.

Reed Mullanix, J. Carette · 0 citations
Preprint Aug 2026

On the Robustness of LLMs'Internal Representation of Code Correctness

This work studies an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice.

Francisco Ribeiro, Sohaila Abdulsattar, R. Gonzalez et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.