Physician Portraits Built with AI from Their Clinical Searches: An Exploratory Evaluation
Abstract
Administrative registries record who is licensed to practice medicine, but not how each physician practices. The search history a physician leaves in a clinical search engine is, by contrast, behavior recorded as it happens: what they searched, when, and how they phrased it. This whitepaper explores whether that history can be used to build a portrait that describes a specific individual, rather than just the type of physician they belong to. We used four mutually blind language-model agents, each reading the same history from a different perspective (clinician, educator, information-behavior researcher, and skeptic), to generate a portrait with a median length of about 10,700 tokens per physician. We tested the resulting portraits on a cohort of 339 physicians and 34,241 searches from Arkangel AI, a clinical search engine, with three evaluations. Identification. We split each physician's history in two, built an independent portrait from each half, and asked whether the portrait from one half matched the portrait from the other half of the same physician better than those of the other 333 candidates. Random guessing would score 0.3%; the portrait scores 78.4% with a TF-IDF representation. However, simply concatenating the raw search text, without any language-model calls, scores 85.0% and outperforms the portrait in every volume stratum, reaching 100% among the highest-volume physicians. Specialty. The portrait correctly identifies the physician's self-declared specialty in 49.9% of cases (169/339), versus 41.3% from raw text and 41.0% from always guessing the most common category (general practice). The advantage comes from generalists: among 200 specialists, no difference is detected (p = 1.00). Future query generation. Given only the portrait, the model generates 20 queries, and the correct physician is ranked higher than with raw text (better for 157 physicians versus 96, p < 0.0001), but the first-place difference is not statistically significant (p = 0.17), and a context-length confound remains uncontrolled. Agreement between portraits built from different halves of each history is limited: the median of the 22 dimension-specific ICCs is 0.496, with only one above 0.75. The results come from the run of 10 and 11 August 2026; a later code review found a training-test conversation overlap affecting 3.7% of the full temporal holdout, which may have favored generation results. Raw text appears to win where exact vocabulary matters and the portrait where synthesis matters. All findings are exploratory, and whether a portrait justifies building it remains to be tested in a prospective study with validated instruments. Keywords: digital phenotyping; large language models; clinical information seeking behavior; health workforce; data governance; Colombia