Fairness assessment is essential to align AI systems with normative and societal values. Because fairness is an inherently contested ethical concept, it is difficult to measure. Although most fairness metrics were developed for binary classification settings, real-world deployments require attention to intersectional social categories [10]. In these settings, an additional and often overlooked modeling step arises: aggregating subgroup-level measures and multiple pairwise group disparities into a single summary score. This step requires selecting a meta-metric, and different choices can produce divergent or even conflicting fairness assessments. We characterize the space of such choices through a three-step taxonomy that makes explicit the normative and technical assumptions embedded in meta-metric design. We show that different design decisions produce substantially different fairness evaluations, with important ethical, methodological, and practical consequences. We therefore argue that aggregation choices should be explicitly justified and carefully aligned with the underlying conception of fairness. To support this process, we provide a framework of criteria and practical guidelines for ethically grounded fairness assessment, with the aim of fostering more transparent and accountable evaluation of intersectional bias in AI systems.
Beyond base metric: The critical and overlooked role of meta-metrics in intersectional fairness
FAccT 2026, ACM Conference on Fairness, Accountability, and Transparency, 25-28 June 2026, Montreal, QC, Canada
Type:
Conference
City:
Montreal
Date:
2026-06-25
Department:
Communication systems
Eurecom Ref:
8905
Copyright:
Creative Commons Attribution 4.0 License (CC-BY)
See also:
PERMALINK : https://www.eurecom.fr/publication/8905