Strong image generation models are conditioned on class labels, aligned to frozen pretrained encoders, or built on separately trained autoencoders. While effective, generation then depends on supervision or pretraining: labels must be annotated, and encoders or autoencoders pretrained for the target domain. We study jo...
Generative foundation models are attracting interest for their ability to produce desired outputs from demonstrations given at inference time, without updating parameters. However, since a few demonstrations cannot uniquely identify the intended task, the challenge is how to learn and sample from an output distribution...
Guoji Fu, Tomoya Wakayama, Ryotaro Kawata et al.· 0 citations
Risk regulation imposes directional constraints on scores; we adopt their strict per-input form -- the score monotone non-decreasing in every exposure input -- as a normative commitment. Deployed pipelines -- monotone hand-crafted aggregates feeding sign-constrained gradient boosting -- already satisfy it by compositio...
The theoretical understanding of multi-layer neural networks is largely confined to overparameterized settings, which obscure parameter identifiability and incur high sample complexity. Neural tangent kernel (NTK) provides a general theory for wide networks, but does not offer efficient sample-complexity guarantees. Re...
Jinqi Tang, Qian Chen, Shihong Ding et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
To mitigate the quadratic complexity bottleneck of the Transformer, sparse attention has emerged as a pivotal technology. Despite the extensive empirical success of sparse Transformers, the theoretical understanding of sparse attention remains fragmented. In particular, two fundamental questions remain unclear: (1) How...
Extreme winter weather has repeatedly disrupted gas-fired power generation in the United States, yet the plant-level data needed to systematically quantify outage risk remain proprietary. Using publicly available weather and electricity demand data together with anonymized generator contingency records from the North A...
Pretrained flow models are now widely used as generative priors in science and vision, where inference-time guidance enables test-time constraints without retraining. Existing methods use projection, posterior sampling, or iterative optimization of the generative trajectory. We propose LyapuFlow, an alternative based o...
Minseon Gwak, Hans Hao-Hsun Hsu, Danielle C. Maddix et al.· 0 citations
Most feature-based choice models, classical and deep, score items and apply a single softmax. We introduce consideration circuits (CC), feature-based models of multi-stage choice defined by directed acyclic graphs of multinomial logit (MNL) units. Source units assign probabilities to menu items, and internal units comb...
Large language models (LLMs) are increasingly used as classifiers, yet they operate as opaque systems whose decisions are difficult to interpret, which complicates their use in regulated domains such as credit scoring or medical diagnosis. We propose an evolutionary framework that iteratively discovers natural language...
Jack Butler, Zainab Afolabi, Nikita Kozodoi· 0 citations
Clinical language models increasingly operate over electronic health records (EHRs), yet patient records are not stored as temporally grounded trajectories. Clinical notes describe symptoms, assessments, and disease progression, but often compress or narratively reorder events. Structured EHR rows provide timestamps fo...
Sayantan Kumar, Shahriar Noroozizadeh, Juyong Kim et al.· 0 citations
Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and...
Mumin Jia, Yilin Chen, Divya Sharma et al.· 0 citations
Causal abstraction offers a principled framework for mechanistic interpretability, aligning a high-level causal model with low-level neural computation through interchange intervention analysis. Finding such an alignment, however, often requires fitting and evaluating separate learned mappings across many candidate neu...
Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity. The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.