Sep 2026· Zenodo (CERN European Organization for Nuclear Research)
Data Visualization and Analytics
Abstract
Existing LLM graph benchmarks typically ask models to answer graph-theoreticquestions or compute symbolic solutions rather than construct spatial layouts.Within-task difficulty is also primarily stratified by vertex count. However, existingresearch also suggests that task difficulty is more closely related to the number ofconstraints imposed by the edges than to the number of vertices being arranged.We introduce PlanarBench, a benchmark that asks models to produce crossing-free ASCII drawings of planar graphs given only an edge list. Across 91 modelconfigurations and 199 non-isomorphic connected planar graphs with 2–7 vertices,edge count is more strongly associated with mean task score than vertex count(r = −0.85 versus r = −0.47) and remains strongly associated after controllingfor vertex count (rpartial = −0.80). PlanarBench provides a controlled settingfor separating these two difficulty axes. In addition, neither drawing area nortotal response length demonstrated a meaningful correlation with score, which isevidence against a simple output-size explanation. Performance varies widely: thebest model scores 159.5 out of 199, most models below 30B parameters scoreunder 25, and substantial failures remain among frontier systems.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The perception of the impact of agile methods is predominantly positive, and several challenge areas were discovered, but based on this study, agile methods are here to stay.
M. Laanti, O. Salo, P. Abrahamsson· Information and Software Tec...· 260 citations· ⚡20
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.
MIT News · Artificial Intelligence· news.mit.eduSep 9, 2026
A weeklong summer workshop brought higher education faculty to campus to explore how AI and machine learning materials can be adapted for their classrooms.
MIT News · Artificial Intelligence· news.mit.eduSep 2, 2026
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.