Skip to content
Preprint

The Learning Objective Governs Perceptual Narrowing: A Cross-Lingual, Layer-Wise, Ten-Seed Study of Self-Supervised Speech Encoders

Aug 2026 · 0 citations · 21 references
Computer Science Engineering

TL;DR

It is concluded that the objective, not the architecture, is the first-order determinant of narrowing-shaped representational change.

Abstract

Perceptual narrowing---the developmental loss of non-native phoneme discrimination in the first year of life \citep{werker1984}---is a canonical developmental finding, yet \emph{what learning objective produces it} remains open. We train a \(\sim\)7\,M-parameter Transformer encoder on child-directed and read speech and evaluate phoneme ABX in English, French, and Mandarin over ten seeds, the seed as the unit of replication. Six results. \textbf{(1)}~The objective sets the direction of cross-lingual transfer: reconstruction (masked mel-prediction) degrades non-native discrimination, prediction (frame-contrastive) improves it---a same-encoder, same-data gap of \(+0.051\) in first-layer Mandarin ABX (\(p=3\times10^{-8}\)), unanimous in sign across twenty runs. \textbf{(2)}~That decline combines a large arm-intrinsic difficulty gradient with a smaller language-specialization effect (matched vs.\ mismatched \(+0.022\), \(p=10^{-4}\), all four layers). \textbf{(3)}~Against a language-symmetric raw-mel floor, reconstruction pushes the first layer \emph{below} the discriminability of its input; prediction pushes it \emph{above}. \textbf{(4)}~Read speech gives a \(3.6\times\) steeper non-native decline than child-directed speech. \textbf{(5)}~The customary three-seed budget cannot see this reliably: an effect unambiguous at ten seeds is called significant by as few as 70\% of three-seed subsets. \textbf{(6)}~Six objective configurations---sharpening, compression, consolidation, their composition, and word-level semantic grounding in two forms---fail to produce the full developmental signature (native improves \emph{and} non-native declines): a single objective moves both languages the same way because it acts on a shared representation. We conclude that the objective, not the architecture, is the first-order determinant of narrowing-shaped representational change.

View source

Similar papers

Open access Aug 2026

Detecting Self-Repairs from Spontaneous Speech with Prompt Ablation Across LLMs and Fine-Tuned Encoder

This work compares the capability of generative LLMs under a five-condition prompt ablation against a fine-tuned DistilBERT token classifier at detecting self-repairs and suggests that a locally deployable encoder, given sufficient in-domain annotation, is a more plausible route to clinical self-repair detection than s...

R. Wu, S. Pugh, K. O'Connor et al. · 0 citations
#machine learning Preprint Sep 2026

The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100

The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phonem...

Francesco Mantegna, Gereon Elvers, D. Jayalath et al. · 0 citations
Preprint Aug 2026

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance.

Yexing Du, Kaiyuan Liu, You-Cheng Pan et al. · 0 citations
Preprint Sep 2026

AudioICL-Bench: A Benchmark for Large Audio Language Model In-Context Learning

In-context learning (ICL) promises training-free adaptation for audio, where labeling every new condition is costly. Yet existing audio ICL studies largely measure Task Recognition, where demonstrations merely cue pre-trained capabilities, rather than Task Learning, where a genuinely new input-label mapping must be inf...

Jia-Hung Chen, Yi-Cheng Lin, Kai-Wei Chang et al. · 0 citations
Jul 2026

Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text

State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On...

Lucas Zamora Vera, Jose A. Gonzalez-Lopez · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.