Skip to content
Preprint

Invariant Pretraining for Robust Code Representations

Aug 2026 · 0 citations · 50 references
Computer Science

TL;DR

InvPT applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs with invariant-contrast pairs for positives of varying difficulty.

Abstract

Encoder-based code representation models remain widely deployed for discriminative tasks such as clone detection and code classification, where their small size and low inference cost are decisive. Their robustness, however, is fragile: under invariant programs, semantically equivalent code written in different syntactic forms, learned representations degrade substantially even though program behavior is unchanged. We present an empirical study of this robustness gap across four encoder baselines, two downstream tasks, and four datasets, together with a minimal code-only continued pretraining recipe that closes much of it. Our method, invariant pretraining (InvPT), applies semantic-preserving transformations to the corpus and combines masked language modeling with multi-positive supervised contrastive learning that treats all augmentations of the same source function as positives, mixing self-contrast pairs (same code, different masks) with invariant-contrast pairs (transformed code) for positives of varying difficulty. Unlike prior contrastive code encoders, InvPT does not require paired natural-language data. Across our evaluation, InvPT improves robustness on transformed test sets by up to 11 percentage points on clone detection and 19 on code classification while matching or improving standard accuracy, and our ablations isolate multi-positive invariant contrast as the main source of the gains. Our aim is not a new objective but a careful measurement of where encoder robustness breaks and how far a simple, code-only recipe can recover it.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

General-purpose code embeddings power tools for code search, classification, and retrieval. Compact transformer encoders for code typically rely on either human-written docstrings (labor-intensive and inconsistent) or mined structural signals such as execution traces (setting-specific and costly to collect). We empiric...

K. Paulsen, Florian Tambon, Mike Papadakis et al. · 0 citations
Aug 2026

CodeLite: A Low-Cost Framework for Code Classification Tasks via Enhanced Code Representations and Lexical Fusion

This work proposes ‘CodeLite’, a relatively lightweight framework that integrates moderately sized code models, such as CodeBERT, with traditional frequency-based techniques, achieving better performance and computational efficiency compared to larger models.

Aman Swaraj, Sandeep Kumar · 0 citations
Preprint Aug 2026

Instruction Alignment for Binary Code Representation Learning

Binary code representation learning is a fundamental problem in software security and reverse engineering. Existing methods mainly learn function-level embeddings that capture coarse-grained semantic relationships between binary functions, but they largely ignore fine-grained instruction-level correspondences. This lim...

Huaijin Wang, Shuai Wang · 0 citations
#machine learning Preprint Sep 2026

Type-IV Code Clone Detection via Layer-Wise Non-Contrastive Representation Learning

LWVIC4Code is proposed, a non-contrastive representation learning approach specifically designed for Type-IV clone detection that achieves competitive or superior performance without negative samples, benefits from layer-wise supervision, and generalizes effectively from Python to other languages, particularly Java and...

Luciano Marchezan, Kévin Delcourt, Eugene Syriani et al. · 0 citations
#artificial intelligence Preprint Sep 2026

UNBIND: UNlearning By INference-time Directional Steering for Code LLMs

Code large language models acquire programming capabilities from large code corpora, but can also memorize implementations that later require removal. Code unlearning is needed to control their continued reproduction when copyright or security concerns arise. However, targeted and retained code share computational patt...

Zhengyang Shan, Jia-Yu Xin, Yan-Jun Lin et al. · 0 citations
Open access Sep 2026

DuaLoc: Dual-Encoder Bug Localization with Bug-Report-Conditioned Attention and Contrastive Learning

DuaLoc combines two pre-trained language models: UniXcoder for the semantic understanding of source code and GraphCodeBERT for awareness of data-flow structure and fine-tuned with a contrastive objective that shapes the embedding space around the localization task.

Amany AlBatlaa, M. Abdullah-Al-Wadud · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.