Skip to content
Book Open access

MAPPE: Rethinking and Improving Fairness in LLMs for Medicine via Minimax Preference-based Prompt Evolution

Aug 2026 · Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2 · pp. 12832-12843 · 0 citations · 83 references

TL;DR

This paper introduces universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities, and proposes MAPPE, a training-free minimax prompt optimization framework that theoretically promotes universal fairness.

Abstract

Large Language Models (LLMs) have shown strong potential in medical applications such as question answering and clinical prediction. % Despite their growing adoption, fairness in LLMs for medicine remains underexplored, largely due to the mismatch between conventional fairness constraints and the clinically meaningful role of sensitive attributes. Existing approaches often enforce attribute-invariant constraints, leading to substantial performance degradation that is unacceptable in high-stakes healthcare settings. Moreover, fairness evaluation for medical LLMs is hindered by the lack of dedicated benchmarks. In this paper, we first rethink fairness in LLMs for medicine from a clinically grounded, utility-based perspective. Inspired by principles of health equity in medicine, we introduce universal fairness, a clinically grounded definition that reframes fairness as maximizing subgroup-aware diagnostic performance under attribute-conditioned health disparities. To achieve this objective in practice, we propose MAPPE, a training-free minimax prompt optimization framework. % MAPPE theoretically promotes universal fairness, while directly applicable to both closed-source and open-source LLMs. To systematically evaluate fairness, we construct FairMed, the first attribute-annotated benchmark for medical LLMs covering medical question answering and clinical prediction. % Experiments on both closed-source and open-source LLMs reveal demographic disparities, while MAPPE consistently improves worst-group and overall performance, outperforming existing fairness-oriented and prompt-based methods. The dataset and code are available at https://github.com/xiye7lai/FairMed.

Read PDF

Similar papers

Sep 2026

FairMoE-Health: Fairness-Aware Mixture of Experts for Equitable Multimodal Clinical Prediction.

Mixture-of-Experts (MoE) models for multi modal clinical prediction route patients to specialized ex pert networks based on input modalities, but we show that this routing mechanism introduces a previously un recognized source of demographic bias: because modality availability (e.g., whether a chest X-ray exists) corre...

Xiaoyang Wang, Christopher C. Yang · 0 citations
Review Open access Aug 2026

Navigating fairness in artificial intelligence-based prediction models: theoretical constructs and practical applications.

It is argued that clinical utility, performance-based metrics, calibration, and statistical parity are the most relevant group-based metrics for medical applications and that different metrics might be applicable depending on the intended use and ethical framework.

S. L. van der Meijden, Yuqing Wang, Madelena Y. Ng et al. · 1 citation
Open access Aug 2026

MaternaAI: Enhancing Equitable Maternal Healthcare in Kerala with Fairness-Aware and Explainable Learning Models

This paper introduces MaternaAI, a fairness-aware and explainable learning framework designed to enhance maternal healthcare predictions in Kerala, India, and proposes Adaptive Equity Score Optimization (AESO), a novel optimization algorithm that dynamically integrates fairness constraints into model training.

A. A. Jo · 0 citations
#software testing Open access Aug 2026

FIFT: Feature Importance-Guided Fairness Testing for machine learning software

Results indicate that global feature importance, used as an active search signal rather than a post-hoc diagnostic, improves both the effectiveness and the efficiency of individual fairness testing.

H. Mamman, Abdullateef Oluwagbemiga Balogun, Mustapha Maidawa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.