01085nas a2200109 4500000000100000008004100001260001200042100001800054700002300072245007500095520080500170 2026 d c04/20261 aAnna Zawadzka1 aPrzemysław Głomb00aArchitecture Matters: Gender Disparities in Automated Image Moderation3 a
Automated image moderation systems shape online visibility and dataset curation, yet prior work has identified demographic disparities in similarity-based NSFW classifiers. We compare such systems with instruction-tuned vision-language models (VLMs) that generate structured moderation decisions with textual rationales. Using the PHASE-annotated subset of the GCC dataset, we compute false positive rates overall and by gender.
Results show substantial variation in moderation strictness and in bias direction: two models show more pronounced strictness in moderation of male images, while other two exhibit higher removal rates for female images.
Reasoning-based models do not consistently mitigate bias but, in the proposed pipeline, they offer a transparency layer.