Skip to content

Similar papers

Preprint Aug 2026

Limitations of Synthetic Data Generation in Specialized Data-Scarce Domains

Advances in diffusion-based generative models have motivated the use of synthetic image generation to alleviate data scarcity in vision tasks. While this strategy has shown promise in natural image benchmarks such as ImageNet, its effectiveness in sparse, high-variance real-world domains remains unclear. In this work, we focus on domains where images differ substantially from common image datasets and additional data are expensive to obtain. Against non-generative data augmentation baselines, we evaluate the downstream classifier performance improvements yielded by two schools of generative sparse data extension: distribution modeling and sample perturbation. Across five trauma classification tasks using subject-wise train--validation splits, no generative approach consistently outperforms a strong non-generative baseline. Feature-space analysis reveals recurring failure modes: memorization or collapse, distributional drift, and generation of visually plausible but simplified canonical instances that are easier to classify than real data.

Edward Zhang, Marcel Hussing, Tanay Tandon et al. · 0 citations
Preprint Aug 2026

NAE: Normalizing AutoEncoder

This work proposes Normalizing Autoencoder (NAE), which employs a novel conditional loss that aligns the surrogate loss gradient with that of reconstruction loss, directly improving upon the current standard.

Muhammad Abdur Rafae, Niels Landwehr · 0 citations

On the Transfer of Output Diversity via Synthetic Data in Language Models

This paper finds that output-distribution tendencies can be partially transmitted through unrelated synthetic data, and suggests that diversity-related behavior is shaped, at least in part, by the latent model state rather than solely by decoding-time randomness.

Adisesh Venkatesh Sanklapur · 0 citations
#natural language process... Preprint Sep 2026

Debias-SparseGPT: Bias-Aware Pruning for Large Language Models

Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models, with outputs varying significantly depending on persona cues in the prompt. In this paper, we introduce Debias-SparseGPT, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs. We perform empirical validation of our method over a wide range of generative LLMs. Across models and sparsity regimes (25%, 50%, and structured 2:4 sparsity), Debias-SparseGPT consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy. Under the most restrictive 2:4 structured sparsity pattern, which most aggressively degrades model quality, augmenting the calibration set with long-context, content-rich examples further improves both downstream performance and fairness. Overall, Debias-SparseGPT advances the bias-performance trade-off while preserving the computational efficiency of sparse models.

Irina Proskurina, Guillaume Metzler, Antoine Gourru et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.