Does Becoming Exceptional at One Domain Reduce Transfer Elsewhere
Abstract
Foundation models derive their value from broad general capability across domains, yet deployment usually rewards specialization - creating a fundamental question for general intelligence: when performance is pushed upward in one region of capability space, is competence elsewhere conserved, redistributed, or destroyed? Synthesizing evidence from continual learning, transfer learning, multi-task optimization, parameter-efficient adaptation, model merging, vision-language adaptation, alignment, and 2023–2026 large-language-model studies, we find that the evidence rejects a universal specialization tax: domain-adaptive pretraining can produce positive transfer, whereas sequential fine-tuning, narrow supervised adaptation, and conflicting objectives can cause catastrophic forgetting, feature distortion, degraded zero-shot transfer, weakened instruction following, or loss of safety behavior depending on task relatedness, update locality, data mixture, optimization geometry, and model capacity. We propose the Generality-Specialization Frontier (GSF), a deployment-oriented framework that treats specialization gain and transfer retention as a Pareto problem, introducing distance-stratified transfer evaluation, invariant retention tests, worst-case regression reporting, and a normalized transfer-elasticity measure. We further propose G-S Bench, an evaluation protocol comparing full fine-tuning, replay, parameter-efficient updates, modular routing, weight interpolation, and non-parametric alternatives under matched target gains, concluding that becoming exceptional at one domain reduces transfer elsewhere only when specialization overwrites shared representations faster than the system preserves broadly useful structure - meaning the tradeoff is an architectural and optimization choice rather than an inevitable law of intelligence.