Skip to content

Category

climate science

381 papers

#machine learning Open access Jun 2025

Engineering RAG Systems for Real-World Applications: Design, Development, and Evaluation

Five domain-specific RAG applications developed for real-world scenarios across governance, cybersecurity, agriculture, industrial research, and medical diagnostics are presented, highlighting technical, operational, and ethical challenges affecting the reliability and usability of RAG systems in practice.

M. Hasan, Muhammad Waseem, Kai-Kristian Kemell et al. · 10 citations · ⚡1
#computer vision Open access Jun 2025

LLM-based Multi-Agent System for Intelligent Refactoring of Haskell Code

Results highlight the ability of LLM-based multi-agent in managing refactoring tasks targeted toward functional programming paradigms and hint that LLM-based multi-agent systems integration into the refactoring of functional programming languages can enhance maintainability and support automated development workflows.

Shahbaz Siddeeq, Muhammad Waseem, Z. Rasheed et al. · 5 citations
#machine learning Open access Feb 2025

Anomaly detection in smart power grids with graph-regularized MS-SVDD: a multimodal subspace learning approach

A generalized Multimodal Subspace Support Vector Data Description model with graph-embedded regularization is proposed, illustrating how relational and structural information can be systematically embedded into one-class models, enabling robust learning under complex, high-dimensional, and multimodal conditions.

Thomas Debelle, F. Sohrab, Pekka Abrahamsson et al. · 1 citation
#artificial intelligence Open access 2026

Bridging Humans and LLMs: Investigating Human-AI Collaboration in Multi-agent Requirements Analysis for Organizational AI Adoption

LLM-based multi-agent systems can support strategic AI planning by enabling iterative refinement with human experts by supporting structured and collaborative Requirements Engineering processes for AI adoption planning.

Malik Abdul Sami, Zheying Zhang, Muhammad Waseem et al. · 6 citations
#computer vision Review Aug 2026

REFINE: A Multi-Agent LLM Approach for Evidence-Guided Code Refactoring

This work introduces REFINE (Refactoring with Evidence-aware Flow for Integrated ageNtic Execution), a tool-agnostic, evidence-aware multi-agent approach for generating Java file-level refactoring candidates that achieves a higher median code-smell reduction with smaller edits and fewer public-method removals.

Muhammad Waseem, Aakash Ahmad, Pekka Abrahamsson · 0 citations
#computer vision Preprint Aug 2026

Trustworthy RAG: An Evaluation Agent for Detecting Misinformation and Knowledge Poisoning in Generative AI Systems

An Evaluation Agent, middleware that combines Natural Language Inference factual verification, a five-signal poison detector with relevance-weighted aggregation, and a Trust Index is proposed, which reliably blocks instruction injection of unsafe advice while contradiction and subtle semantic weakening remain hard.

Balkrishna Giri, M. Hasan, Jussi Rasku et al. · 0 citations
#computer vision Preprint Aug 2026

CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation

These findings show that reliable evaluation of LLM-generated code requires validated ground truth, protected tests, and multiple explicitly interpreted measures, and that CodeAssay provides a reproducible basis for evidence-based model evaluation in AI-augmented software development.

Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi et al. · 0 citations

Constitutional Midtraining: Content Presence Drives Alignment Gains

Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtr...

D. Cho, Cameron Tice, Bernie Hogan et al. · 1 citation

From tech blogs

See all →
Microsoft Research Blog Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville did not want to go into computer science—but that’s exactly where she landed. Neville discusses the starts and stops that led to her professional sweet spot and her work identifying “surprising failures” making it hard for AI to handle complexity.  The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.