Skip to content

Author

Pressfront Research

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

Which Publications Do AI Assistants Cite When Recommending Businesses? An Exploratory Study

Consumers increasingly ask conversational large language models (LLMs) which product or company to choose. When web-search-enabled assistants answer, they ground their recommendations in sources retrieved at query time, yet little is documented about which publications they draw on. We conducted an exploratory study of the sources cited by two widely used assistants, OpenAI's ChatGPT and Google's Gemini, across 100 open "best/top X for a small business" questions spanning ten commercial categories, issued twice to each assistant (398 usable responses). For every response we recorded the source domains the assistant grounded on, excluding search-action links, and separated editorial publications from platform properties. The two assistants drew on almost entirely different sources: among editorial domains cited in three or more responses, their overlap was only 12%. ChatGPT leaned on technology-review and financial press (TechRadar was its most-cited source, in 16% of its responses; followed by CNBC, Tom's Guide, and Yahoo Finance). Gemini grounded heavily on platform properties (YouTube in 53% of its responses, Reddit in 25%) and on consumer-review and comparison sites (Forbes, Trustpilot, ConsumerAffairs, G2). Press-release wire domains appeared in under 3% of responses, marginally more often via ChatGPT. Important limitation: Gemini's grounding interface returns redirect labels rather than verifiable source URLs, so its reported sources cannot be independently confirmed, and part of the observed divergence reflects how each platform reports sources, not only what it reads. We release the full dataset. Given the exploratory sample, results are directional, not definitive.

Pressfront Research · 0 citations
#large language models Open access Sep 2026

Run-to-Run Consistency of Business Recommendations from Web-Search-Enabled Large Language Models: An Exploratory Study

Consumers increasingly ask conversational large language models (LLMs) which product or company to choose, yet little is documented about how stable those recommendations are under repetition. We conducted an exploratory study of run-to-run consistency for two widely used web-search-enabled assistants, OpenAI's ChatGPT and Google's Gemini. Ten open "best X for a small business" questions spanning ten commercial categories were each issued three times to each assistant within a single day (60 responses total), with web search enabled throughout. From each response we extracted the set of recommended businesses and measured consistency as the mean pairwise Jaccard overlap of these sets across the three repetitions. Overall consistency was 69.5%, i.e. roughly one recommended business in three changed between identical queries. Consistency differed by assistant (ChatGPT 87.2% ± 17.5; Gemini 51.9% ± 11.8) but the instability was not uniform: of 127 distinct business-slots, 56% appeared in every repetition (a stable core) while 27% appeared in only one of three (a volatile tail). We discuss implications for reproducibility research and for the emerging practice of generative engine optimization, and we release the full method and response data. Given the small sample the results are directional, not definitive.

Pressfront Research · 0 citations
#large language models Open access Sep 2026

Run-to-Run Consistency of Business Recommendations from Web-Search-Enabled Large Language Models: An Exploratory Study

Consumers increasingly ask conversational large language models (LLMs) which product or company to choose, yet little is documented about how stable those recommendations are under repetition. We conducted an exploratory study of run-to-run consistency for two widely used web-search-enabled assistants, OpenAI's ChatGPT and Google's Gemini. Ten open "best X for a small business" questions spanning ten commercial categories were each issued three times to each assistant within a single day (60 responses total), with web search enabled throughout. From each response we extracted the set of recommended businesses and measured consistency as the mean pairwise Jaccard overlap of these sets across the three repetitions. Overall consistency was 69.5%, i.e. roughly one recommended business in three changed between identical queries. Consistency differed by assistant (ChatGPT 87.2% ± 17.5; Gemini 51.9% ± 11.8) but the instability was not uniform: of 127 distinct business-slots, 56% appeared in every repetition (a stable core) while 27% appeared in only one of three (a volatile tail). We discuss implications for reproducibility research and for the emerging practice of generative engine optimization, and we release the full method and response data. Given the small sample the results are directional, not definitive.

Pressfront Research · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.