Oct 2026· Artificial Life and Robotics· 12 references
Abstract
Abstract This study highlights the potential of image-based reinforcement learning methods for addressing swarm-related tasks. In multi-agent reinforcement learning, effective policy learning depends on how agents sense, interpret, and process local inputs. Traditional approaches often rely on handcrafted feature extraction or raw vector-based representations, which limit the scalability and efficiency of learned policies concerning input order and size. In this work, we propose an image-based reinforcement learning method for decentralized control of a multi-agent system, where observations are encoded as structured visual inputs that neural networks can process, extracting spatial features and producing decentralized motion control rules. We evaluate our approach on a multi-agent gathering task of agents with limited-range and bearing-only sensing that aim to preserve visibility-graph connectivity during the aggregation. The algorithm’s performance is evaluated against two benchmarks: an analytical solution proposed by Bellaiche and Bruckstein, which ensures gathering success, but progresses slowly, and VariAntNet , a neural network-based framework that gathers much faster, but shows moderate success rates in adverse initial agent configurations. Our method achieves high gathering success, with a speed nearly matching that of VariAntNet . In certain scenarios, it offers a viable alternative to existing approaches.
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
This state-of-practice investigation was performed using a literature review followed by a multiple-case study approach and presents how inconsistency between managerial strategies and execution can lead to failure by means of a behavioral framework.
Carmine Giardino, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 175 citations· ⚡19
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
It is found that roles of MVPs in startups were not fully aware by entrepreneurs, and entrepreneurs should consider a systematic approach to fully explore the value of MVP, as a multiple facet product (MFP).
Anh Nguyen-Duc, P. Abrahamsson· International Conference on...· 93 citations· ⚡9
It is found that what perceived as biggest challenges by software startups do vary across different life cycle stages, even though its significance decreases when the learning focuses of the startups move from problem to solution and their products mature.
Xiaofeng Wang, Henry Edison, Sohaib Shahid Bajwa et al.· International Conference on...· 62 citations· ⚡6
Agile methods continue to gain popularity. In particular, the Scrum method appears to be on the verge of becoming a de-facto standard in the industry, leading the so called Agile movement. While there are success stories and recommendations, there is little scientifically valid evidence of the challenges in the adoptio...
A. Marchenko, P. Abrahamsson· Agile Conference· 59 citations· ⚡11
Related blog posts
Microsoft Research Blog· microsoft.comJul 13, 2026
Cryptographic code supports vital protections in modern computing systems. Learn how a new method helps verify code as developers write it while preserving speed and adaptability as it gets implemented and evolves. The post Verifying Rust cryptography in SymCrypt, from standards to code appeared first on Microsoft Research.
Students in MIT’s Concourse program delve deeply into the human condition, debate challenging questions, and learn to develop judgment about issues that can’t be quantified.
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first on Microsoft Research.
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.