Skip to content

Author

Martin T. Vechev

We have 6 of 311 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

AutoMark: Enabling Autoresearch to Discover Better LLM Watermarks

With LLM watermarking being deployed commercially and now required by regulations, improving its reliability and effectiveness has become crucial. Yet, recent progress in the field of LLM watermarking has increasingly been driven by improving details of existing methods, an effort fundamentally limited by the pace of h...

Thibaud Gloaguen, Robin Staab, Martin T. Vechev · 0 citations
#artificial intelligence Review Feb 2026

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents?

This work evaluates coding agents' completion performance in two complementary settings: established SWE-bench tasks from popular repositories, with LLM-generated context files, and a novel collection of issues from repositories containing developer-committed context files.

Thibaud Gloaguen, Niels Mündler, Mark Niklas Müller et al. · 28 citations · ⚡1

SABER-Math: Automated Benchmark for Information Retrieval Evaluation in Mathematics

SABER-Math is introduced, the first fully automated benchmark for evaluating mathematical IR without expert annotation, and it is shown that general-purpose IR benchmarks such as MTEB do not reliably predict mathematical performance, especially for recent embedding models, highlighting the need for math-specific retrie...

N. Georgiev, Maria Drencheva, Kseniia Ibragimova et al. · 0 citations
#artificial intelligence Preprint May 2025

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

An attack is proposed, FAB (Finetuning-activated Adversarial Behaviors), which compromises an LLM via meta-learning techniques that simulate downstream finetuning, explicitly optimizing for the emergence of adversarial behaviors in the finetuned models.

Thibaud Gloaguen, Mark Vero, Robin Staab et al. · 4 citations

Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation

This work uses classic instruction tuning, supervised fine-tuning without reasoning traces, on the RLM to improve RLM performance in both verifiable and hard-to-verify domains, including coding and text summarization, while preserving RLM capabilities across other domains.

Yuanning Feng, Niels Mündler-Sasahara, Mark Vero et al. · 0 citations
Jul 2026

Generative Compilation: On-the-Fly Compiler Feedback as AI Generates Code

It is shown that generative compilation reduces non-compiling outputs and improves functional correctness, relative to standard post-generation feedback, by detecting a broad range of errors close to their source and early during generation, thereby reducing errors cascades and enabling focused diagnostics.

Niels Mündler-Sasahara, Hristo Venev, Dawn Song et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.