A task-level expertise database for ISCO-08, the international standard that allows for cross-country comparisons, is built and shows that generative AI reaches the tasks that make an occupation expert in some occupational groups but not others.
A growing literature measures how far occupations are exposed to AI, but these measures capture where AI could perform tasks, not whether workers have adopted it. We propose a new layer of exposure, delegated exposure, which records whether a worker has committed a task to AI by building it into a workflow. We operationalize it as the Agentic Adoption Index (AAI), which measures how closely an occupation's tasks match the agentic routines practitioners have already built and shared. We embed roughly 53,000 agent skill specifications from the Manus Skills Marketplace, compute their semantic similarity to about 18,000 O*NET task statements, and aggregate to the occupation level. Three findings follow. First, the occupations where delegation concentrates differ sharply from those pre-AI frameworks identified as most at risk. Second, the AAI tracks what AI could do more closely than what workers currently use it for. Third, the AAI peaks below the top of the wage distribution and at the bachelor's level, declining at both extremes. Technical availability explains most of this variation, but not the shortfall among the most educated occupations, so feasibility alone cannot account for who adopts. That shortfall may reflect work that resists advance specification, or professional discretion over the pace of codification. Distinguishing the two, and tracking how these measures diverge over time, will require repeated measurement.
KV-Skill is introduced, a design space of external factorized operators that a frozen language model reads through a lightweight interface that shows that task knowledge can be acquired from text or experience, compressed into an external operator, and deployed separately from the backbone.
Zhaowei Han, Xiang Zhang, Bing Han et al.· 1 citation
This work defines a grouping metric, specify a harness, and shows how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires.
NameRank, a [0,1] recognition score, is built: recognition is paid to named, indexable artifacts, not to credentials or titles, because recognition attaches to the artifact's own distinctive name, not to the roster behind it.
A new measure of curricular exposure to large language models is constructed by combining task-level estimates of LLM capabilities with course descriptions from more than 1,000 U.S. colleges and universities, showing that colleges have recognized the instructional challenge posed by generative AI but have made limited observable changes to how student learning is assessed.
JacobLight, David Autor, Nick Bloom et al.· 0 citations
A lot of research attention has been devoted to checking whether large language models (LLMs) are politically biased. This work has largely focused on high-level ideological dimensions, such as left--right or progressive--conservative, and it has been shown that while LLMs are predominantly left and progressive leaning, largely mimicking the biases in the training data, they can be to some extent steered to change their preferences in post-training. In this short note, we check if LLMs have robust stances with regard to major substantive societal issues, on which members of the same ideological camp are often in disagreement, summarised in a novel dataset \textsc{HardChoices}. We show that, faced with this line of questioning, LLMs, both large and small, surprisingly rarely declare neutrality, are often incoherent, and demonstrate a remarkable degree of agreement on issues where they do take stances.
Dmitry Nikolaev· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.