This work introduces the problem of personalized auto-research, which conditions every stage of the research process on a representation of the individual researcher, and proposes a general and flexible framework that threads a graph-grounded researcher context through retrieval, hypothesis search, experimentation, writing, and review.
Abstract
AI co-scientists that generate hypotheses, retrieve related work, design experiments, execute code, and draft full papers are beginning to change how research is carried out. Despite this rapid progress, state-of-the-art systems remain researcher-agnostic: given a research goal, they optimize novelty, validity, or reviewer score while ignoring the individual scientist who will use the output. This overlooks a fundamental fact about research, namely, that what counts as novel, valuable, or feasible depends on the researcher, including their prior work, methodological repertoire, and the collaborators and communities in which they are embedded. In this work, we introduce the problem of personalized auto-research, which conditions every stage of the research process on a representation of the individual researcher. We argue that personalization is not a convenience layer, but rather the fundamental property that allows an AI system to serve as a genuine co-scientist rather than a generic instrument. To address this problem, we propose a general and flexible framework that threads a graph-grounded researcher context through retrieval, hypothesis search, experimentation, writing, and review. The framework consists of three fundamental components: (i) graph-grounded researcher representations, (ii) personalization across the full research pipeline, and (iii) evaluation grounded in the individual. Notably, we highlight a one-size-fits-all failure mode where distinct researchers issuing the same goal receive essentially the same research, erasing the tacit knowledge through which novel ideas arise. Finally, we discuss fundamental open problems and challenges.
The replicability of papers is a cornerstone of scientific knowledge, ensuring the reliability of existing results and providing a base for further experiments. The act of replication typically illuminates details that were previously underspecified, and thus requires similar hypothesis-driven exploration to open-ended research. In this work, we develop Replica, a scalable task space for paper replication. To provide reward signal, we introduce an auto-generated rubric-based judge that has low noise and agrees with human assessment of replication quality. We post-train Faraday, a 27B-parameter"AI Scientist"agent that leverages coding agents as tools, surpassing the performance of Claude Opus 4.8 and GPT-5.5 on held-out replication tasks. Qualitative analysis of individual rollouts reveals that Faraday adopts a more scientifically-principled approach. We believe that our results provide a stepping stone towards AI agents capable of long-horizon scientific innovation without requiring complex harnesses.
Damon Falck, Samer Sabri, Anja Surina et al.· 0 citations
Inspired by search and recommender systems, this work builds Find, Attempt, and Recommend (FAR), a literature-to-review cascade that automates the search for suitable problems and focuses human attention on artifacts that have passed several stages of filtering.
Zeyu Zheng, Shengtong Zhang, Jeremy Avigad et al.· 0 citations
This article examines the emerging paradigm of agentic AI for scientific discovery, traces the conceptual shift from tools to agents, lays out a six-stage workflow spanning literature synthesis to manuscript generation, and reviews practical systems in chemistry, equation discovery, materials science, and general machine learning research.
Alexander Taktakidze· Longevity Horizon· 0 citations
Artificial intelligence (AI) is increasingly embedded in scientific work, but researchers may not evaluate its use uniformly across research tasks. This study examines task-specific attitudes towards AI among an international, self-selected sample of 3,785 PhD students in STEM and medical and health sciences who participated in Nature's Graduate Survey 2025. We analyse respondents'comfort with using AI for writing a research article, collecting and analysing data, designing experiments, tracking scientific literature, and summarising it. Latent class analysis identifies four distinct attitudinal profiles. The dominant profile reflects a"division of labour,"in which AI is widely accepted for literature-related tasks but resisted in activities closely associated with intellectual contribution, such as writing, data analysis, and experimental design. A"status quo"profile is broadly uncomfortable across tasks, an"all-purpose"profile is broadly comfortable, and an"undecided"profile expresses substantial uncertainty. These patterns suggest that attitudes towards AI in research are organised less around a simple acceptance-rejection divide than around task-specific boundaries, likely concerning delegation, authorship, and responsibility. Because the survey measures comfort rather than legitimacy, the profiles are best interpreted as attitudinal configurations with a normative dimension. The findings highlight the importance of task-specific approaches to AI governance, doctoral training, disclosure, and research evaluation.
The results raise concerns that LLM-assisted evaluation may under-select proposals that human reviewers identify as highly novel, potentially reflecting the statistical logic of next-token prediction trained on past scientific outputs.
The Little Scientist is presented, a framework in which a LLM agent stepping through the scientific method can discover both new algorithms and new ensemble strategies that outperform prior solutions, and is demonstrated on two problems that require fundamentally different modes of discovery.
T. Smith· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.