Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated, differentially private, adaptively monitored retrieval-augmented generation. Clients release only Gaussian-perturbed score and loss histograms. Record-indexed and noise...
Agents that answer questions from compressed or retrieved memory must recognize when the evidence a query needs is no longer in memory. Benchmarks for this task usually create insufficient-evidence examples by deleting supporting passages. We show that this construction leaks the label through memory size: on MuSiQue,...
Joyanta J. Mondal, Md. Shifatul Ahsan Apurba, Mridul Banik et al.· 0 citations
This research describes this certification frontier for Gaussian LoRA posteriors and separates three interventions: changing the prior, changing the stochastic predictor, and changing only how its complexity is counted.
Joyanta J. Mondal, Ibne Farabi Shihab· 0 citations
This work introduces StepCOPS, which uses an independent proposal split to nominate one lower-tail floor per candidate, exact binomial tests on a fresh certification split, and Holm's step-down procedure to certify a set of floors.
The demonstrated benefit is robust sparse advice volume on a useful task, not a proven per-state placement advantage, because the response-contingent metareasoning problem is formulated as a response-contingent metareasoning problem.
Ibne Farabi Shihab, Md Najmus Swaqeeb, Abu Sa-Adat Mohamed Moon-Im Al Ahsan· 0 citations
The perturbation analysis motivates pre-norm attention and MLP components under explicit local assumptions; results on AdaLN, U-shaped, convolutional, and cross-attention blocks are empirical transfer, not certified guarantees.
Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Anuj Sharma· 0 citations
Task-Aware Spectral Pruning (TASP), a post-training framework that calibrates module-level spectral descriptors against measured task-specific ablation effects, closes grouped-query-attention and SwiGLU dependencies during sparse-mask construction, and routes each user turn to one compiled mask that remains fixed throu...
BLADE is presented, a variational model that separates latent truth from graph recording and distills offline language-model judgments into a frozen teacher regularizer and claims calibration only for the declared candidate distributions, not for all unobserved triples.
Ibne Farabi Shihab, Rabeya Bosri Tamanna, Abdo El Karaky et al.· 0 citations
A scalar recalibration map fitted for an LLM judge on one task can fail when the task distribution changes, but the source-target accuracy gap is often treated as a proxy for that failure. We test what this gap can predict and what it can certify across thirteen judges, two generators, eight domains, and 1,176 predecla...
A finite-width conditional bound linking the inverse participation ratio of squared singular values to central pre-softmax logit kurtosis is derived and a pointwise query--key product-tail target is defined and compared with independent factor surgery with a product-targeted factorization that preserves native attentio...
Calibration-Preserving Pruning augments a base pruning score with nonconformity-gradient saliency and uses disjoint pruning, validation-selection, conformal-calibration, and test splits to obtain smaller valid prediction sets.
An explicit finite lower certificate for communicating MDPs is derived and an auditable composition rule for a span-constrained optimistic learner is given, and valid expectation conversion and constant comparability are formalized.
Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.