Uniform stability is a classical tool for controlling the generalization error of a learning algorithm. Bousquet, Klochkov, and Zhivotovskiy (2020) showed that the problem can be reduced to a moment inequality for a sum of weakly interacting functions of independent random variables. Their bound contains an additional factor $\log n$, and they asked whether this factor can be removed. We answer this upper-bound question affirmatively. More specifically, let $Z=(Z_1,\ldots,Z_n)$ have independent coordinates and let $g_i(Z)$ satisfy $$ \mathbb E[g_i(Z)\mid Z_{-i}]=0, \qquad \left| \mathbb E[g_i(Z)\mid Z_i]\right|\le M, \qquad \forall i = \overline{1, n} $$ while changing any coordinate $Z_j$, $j\neq i$, changes $g_i$ by at most $\beta$ and $Z_{-i}$ denotes all coordinates except $Z_i$. We prove that, for every $p\ge2$, $$ \left\| \sum_{i=1}^n g_i(Z)\right\|_p \le 16pn\beta+M\sqrt{2pn}. $$ This removes the $\log n$ factor from the previous bound and matches the lower bound of Bousquet, Klochkov, and Zhivotovskiy up to universal constants in the range covered by their construction. Our proof first establishes the required estimate on the Rademacher cube, then transfers it to arbitrary product distributions by a two-copy randomization argument.
Uniform stability controls how much one training example can change the loss at any test point. A new logarithmic-free upper bound shows that a $\gamma$-uniformly stable algorithm with loss in $[0,L]$ has generalization gap at most $O \left(\gamma\log(1/\delta) +L\sqrt{\frac{\log(1/\delta)}{n}}\right)$ with probability $1-\delta$. Whether an actual bounded-loss learning algorithm can realize the linear dependence on $\log(1/\delta)$ has remained open. The known construction realizes it only for auxiliary weakly dependent random variables whose pointwise range grows with $n$. The known learning lower bound holds only at constant probability. We close this gap. For every $n$, stability level $\gamma$, and loss bound $L$, we construct one deterministic $\gamma$-uniformly stable learning problem whose tail satisfies, simultaneously for $1\le p\le c n$, $\mathbb P \left( R(A_S)-R_S(A_S) \ge c'\min \left\{L,\gamma p+L\sqrt{p/n}\right\} \right)\ge e^{-p}.$ The construction is ordinary bounded absolute-loss regression with constant labels. Its key is a multiscale collection of rare Rademacher features. A coordinatewise ramp is stable in sup norm, while an odd symmetrized maximum converts a unique extreme feature into a gap of order $\gamma p$ without violating the loss bound. Geometrically spaced ramps put all confidence levels into the same problem. Together with the logarithmic-free upper bound, this determines the optimal high-probability and moment dependence of uniform stability up to universal constants.
Consider $n$ independent, non-negative, mean at most one random variables, $X_1,X_2,\ldots$. We show the following bound on the probability of their sum exceeding a threshold $t$: \[ \mathbb{P}\left[\sum_{i=1}^n X_i\ge t\right] \leq 1-\left(1-\frac{1}{t}\right)^n \text{ for all } t\ge 2n+1 \,. \] To prove this, we consider a relaxed optimization problem over a set of sequences of ordered, but non-independent random variables. This allows us to reformulate it recursively as dynamic programming problem. The bound becomes an equality for the binary i.i.d.~random variables satisfying $\mathbb{P}\left[X_i=0\right]= 1-\frac{1}{t}$ and $\mathbb{P}\left[X_i=t\right]=\frac{1}{t}$, which remains the maximizer in the relaxed problem.
The Johnson--Lindenstrauss lemma asserts that every set of $n$ points in $d$-dimensional Euclidean space embeds into $O(\varepsilon^{-2}\log n)$-dimensional Euclidean space with distortion at most $1+\varepsilon$. Larsen and Nelson conjectured that the optimal target dimension throughout the full range of the parameters $n,d, \varepsilon$ is \[ \Theta\left(\min\left\{d,n-1,\frac{\log(2+\varepsilon^2n)}{\varepsilon^2}\right\}\right). \] We resolve this conjecture in the affirmative. In fact, we prove the stronger statement that the upper bound is attained by a linear map. The matching lower bound, due to Larsen--Nelson and Alon--Klartag, holds even for nonlinear embeddings.
We establish a power-sum convergence principle for randomly weighted means. Let $P(t)=(P_j(t))$ be random finitely supported subprobability weight sequences, independent of iid centered integrable marks. If the expected total mass converges and the expected power sum of some order $r\in(0,1)$ is uniformly bounded, then convergence, for one fixed mark law, of $\sum_jP_j(t)X_j$ to a nondegenerate law forces $\mathbb{E}\sum_jP_j(t)^p$ to converge to a positive limit for some $p\in(1,2]$. The proof combines normal-family compactness for Mellin transforms of characteristic-function remainders, inversion at frequencies $\pm u$, and Landau's theorem at the abscissa of convergence; a uniform Abelian estimate handles the endpoint $p=2$. For weights obtained by normalizing a Poisson-sized iid sample of nonnegative variables, a gap below one for the logarithmic slope of the Laplace exponent yields the required power-sum bound of order below one. A Poissonized ratio Tauberian theorem then identifies the common tail of the unnormalized variables as regularly varying, with index $-\beta$ for a unique $\beta\in[0,1)$. As an application, this proves the remaining necessity direction in Breiman's 1965 conjecture for centered integrable marks. Combined with Breiman's sufficiency theorem, it completes the conjecture.
Let $X_1,X_2,\ldots$ be independent $\mathrm{Bernoulli}(\theta)$ random variables, and let $\bar X_n = n^{-1}(X_1 + \cdots + X_n)$. We prove that, for every real $p \geq 1$, the sequence $\{\mathsf{E}(\bar X_n^p)\}_{n \geq 1}$ is log-convex. This proves the Bernoulli case of a conjecture of Lamkin and Tkocz [Canad. Math. Bull., 65(2):271-278, 2022]. The proof conditions on the total number of successes among $2n$ trials and reduces the desired inequality to a convex-order comparison between two normalized quadratic functions of hypergeometric random variables. All the log-convexity inequalities are strict for $p>1$ and $0<\theta<1$.
The quadratic inverse large sieve problem predicts that the examples sharp at the square-root threshold are essentially quadratic. Hanson proved the first unconditional result in this direction: if $A\subseteq[N]$, $|A|\gg\sqrt N$, and $|A_p|\le p/2+O(1)$ for every prime $p$, then $A$ contains $\gg\log N$ elements in the image of a single quadratic. We significantly improve this lower bound to \[ \exp\left(c\frac{\sqrt{\log N}}{\log\log N}\right). \] We also prove density-dependent variants, including a two-set version motivated by Green--Harper's robust inverse large sieve conjectures and their connection with the inverse Goldbach problem. Combined with a theorem of Elsholtz--Harper on hypothetical decompositions of the primes, our results show that any such decomposition would force both summands to have large intersections with quadratic images. Our proof combines a weighted entropy argument with sieve estimates, inspired by the recent work of Croot--Mao--Pohoata--Sheffer--Yip.