Across AI-driven science, generative capacity has outpaced investigative judgment. It is now possible to produce millions of plausible candidates in hours, but not yet to decide which merit expensive validation. This asymmetry defines the central bottleneck of scientific discovery for the next decade. This blue sky paper casts scientific investigation as an information market. A strategically intelligent agent must exploit cost-information asymmetries across fidelity levels, purchasing cheap inferences to decide whether expensive ones are worth acquiring. The paper formalizes this principle as Inferential Arbitrage: a multi-fidelity Markov Decision Process in which agents maximize inferential return on investment. The formulation is pluralistic by design, admitting reinforcement learning, evolutionary algorithms, and multi-armed bandits. The framework is developed in protein design, where the vendor hierarchy from sequence embeddings to wet-lab assays spans several orders of magnitude, but the same structure recurs in materials science and autonomous chemistry. Three grand challenges (cross-modal calibration, asynchronous decision logic, interpretable strategy) anchor a 2030 benchmark: an agent that discovers a functional protein at 1% of the compute and 10% of the lab cost of human-designed pipelines. The ultimate measure of an autonomous AI scientist is not what it can compute, but what it chooses not to.
Amarda Shehu· Proceedings of the 32nd ACM...· 0 citations
There is ongoing academic debate on whether one can teach AI literacy to undergraduate students across majors, and if yes, how. This article reports a case study: a three-week midterm project embedded in an undergraduate “AI-for-all” course. Students designed reasoning tasks, ran controlled comparisons across widely used chatbots, and evaluated both answer correctness and explanation validity. Through field experience, students with no STEM background learned what consumer chatbots can and cannot do, documenting systematic brittleness across models that “sounded right but reasoned wrong.” More critically, students built understanding of how to evaluate AI outputs. The midterm gave them agency as investigators rather than passive users. Eager to share their discoveries, they are co-authors of this article. Together, we offer here to educators and the broader scientific community a concrete example of the operationalization of AI literacy as experimental practice. The method, however, is not specific to the classroom. It shows any user how to test an AI system rather than trust it blindly. In three-week midterm project, students investigated whether AI literacy can be taught to undergrads.
Amarda Shehu, Adonyas Ababu, Asma Akbary et al.· Communications of the ACM· 0 citations
Stochastic Gradient Descent (SGD) and its variants are single-objective optimizers focused on minimizing training loss, often failing to address the generalization gap in deep learning. In this paper we introduce EvoBatch, a novel Hybrid Evolutionary Algorithm (HEA) that re-frames deep network optimization as an explicitly multi-objective problem. EvoBatch leverages a two-stage selection process, guided by both training loss and validation performance, to directly optimize for generalization. To overcome the historical computational barrier of EAs, EvoBatch uses mini-batch evolutionary local search, restricting each individual to local updates on unique, random data subsets. Theoretically, this M-ELS mechanism acts as a robust implicit regularizer by injecting heterogeneous noise, promoting the discovery of broader, flatter minima. This stability is the necessary condition that allows the explicit multi-objective selection to systematically and monotonically reduce the generalization gap. Empirically, EvoBatch demonstrates superior generalization profiles and consistently outperforms gradient-based baselines across image classification (ResNet, ViT) and language understanding (BERT) benchmarks. Our findings establish that compute-constrained, multi-objective evolutionary optimization offers both an effective and an efficient alternative to single-objective gradient methods when generalization is critical.
Toki Tahmid Inan, Shahana Shultana, Amarda Shehu· Annual Conference on Genetic...· 0 citations
A three-week midterm project embedded in an undergraduate “AI-for-all” course investigated whether AI literacy can be taught to undergrads, and shows any user how to test an AI system rather than trust it blindly.
Amarda Shehu, Adonyas Ababu, Asma Akbary et al.· Communications of the ACM· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.