Hundreds of millions of people interact with language models (LMs) every day, using them for tasks such as writing assistance and information seeking. As their range of applications grows, it becomes increasingly essential to consider the role of social context in how these systems are designed and deployed. In particular, this thesis focuses on LMs' sociolinguistic competence—how associations between linguistic variation and social dimensions can be incorporated into LMs, how they are learned and manifested, and how they can lead to harm. In the first part of the thesis, we develop computational methods that improve LMs' sociolinguistic competence by explicitly injecting social context into the model architecture. We focus on two forms of social context—social networks and geographic location—and draw on recent advances in graph neural networks and multi-task learning to integrate them into LMs. Across a range of benchmarks, the proposed methods yield substantial gains. In the second part of the thesis, we explore whether LMs acquire sociolinguistic competence as a by-product of pretraining and posttraining, without being explicitly conditioned on social context. Experiments on dialectal variation and ideological framing suggest that LMs indeed learn associations between linguistic variation and social dimensions, albeit with varying levels of detail. Beyond these sociolinguistic associations, we also examine the question of how LMs' outputs reflect ideological leanings more generally, finding substantial evidence of instability. In the third part of the thesis, we investigate the harms that associations between linguistic variation and social dimensions can produce in LMs. Focusing on African American English, we find that LMs associate its speakers with pernicious stereotypes triggered by linguistic features alone, and that current posttraining practices do not address this covert racism. Preventing such harms is a critical goal for future research to ensure safe and equitable language technology. Finally, we release new analysis tools and datasets that facilitate broader empirical study of social context in natural language processing and computational social science, supporting subsequent work in these areas.
Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoption of Agile methods in general, and Scrum in particular. Little, if anything, is empirically known about the application and adoption of Scrum in a multi-team and multi-project situation. The authors carried out an ethnographically informed longitudinal case study in industrial settings and closely followed how the Scrum method was adopted in a 20-person department, working in a simultaneous multi-project R&D environment. Altogether 10 challenges pertinent to the case of multi-team multi-project Scrum adoption were identified in the study. The authors contend that these results carry great relevance for other industrial teams. Future research avenues arising from the study are indicated.
A. Marchenko, P. Abrahamsson· Agile Conference· 59 citations· ⚡11
A comprehensive taxonomy of the challenges faced when a medium-scale organization decided to adopt software platforms is provided, namely: business challenges, organizational challenges, technical challenges, and people challenges.
Yaser Ghanam, F. Maurer, P. Abrahamsson· Information and Software Tec...· 41 citations· ⚡3
It is shown that high article processing charges are not sufficiently justified by the publishers, which often lack transparency and may prevent authors from adopting OA.
D. Graziotin, Xiaofeng Wang, P. Abrahamsson· Scientometrics· 21 citations· ⚡1
MCGLPPI, a novel geometric representation learning framework that combines graph neural networks (GNNs) with the MARTINI molecular coarse-grained (CG) model to predict overall PPI properties accurately and efficiently, offers an effective and efficient solution for PPI overall property predictions.
Yang Yue, Shu Li, Yihua Cheng et al.· bioRxiv· 15 citations
PepPCBench enables a robust evaluation of PFNN-based methods and supports their continued development for peptide-protein structure prediction, and highlights the influence of peptide length, conformational flexibility, and training set similarity on prediction accuracy.
Si-Long Zhai, Huifeng Zhao, Ji-Ke Wang et al.· Journal of Chemical Informat...· 13 citations· ⚡1
OmniMol is presented, a framework using hypergraphs to improve predictions of molecular properties, addressing challenges of imperfect data annotation and enhancing model explainability, and achieves state-of-the-art performance in properties prediction.
Assistant Professor Pat Pataranutaporn describes a new interface that lets everyday users glimpse inside an AI's neural network before their chatbot ever says a word.
Microsoft Research Blog· microsoft.comJul 13, 2026
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduJul 6, 2026
PhD student Rachel Sava, winner of the Envisioning the Future of Computing Prize, explores transformative improvements and dystopian risks of neural technology.