DIAG-Sapienza at GSI:detect: Joint Detection and Classification of Gender Stereotypes with Structured Prompting and Fine-Tuning
Unmasking Bias: How We Jointly Detect and Classify Gender Stereotypes in NLP
In the rapidly advancing field of Natural Language Processing (NLP), AI models are becoming increasingly powerful, but they often inherit the biases present in their training data. One of the most persistent issues is the perpetuation of harmful gender stereotypes.
Our latest work, DIAG-Sapienza at GSI:detect, tackles this critical problem head-on. We introduce a novel framework that doesn’t just detect potential bias; it jointly detects and accurately classifies the nature of gender stereotypes within text. Think of it as giving NLP models a precise diagnostic tool to pinpoint exactly where, how, and why their language is biased.
π‘ The Challenge: Bias Detection Isn’t Enough
Traditional methods often treat bias detection as a binary switchβeither bias exists or it doesn’t. However, real-world linguistic bias is complex and multi-layered. Our research refines this approach by using advanced techniques like Structured Prompting and dedicated fine-tuning, allowing the model to understand not just that there is bias, but the specific mechanisms and categories of stereotypes present (e.g., occupational, social roles).
βοΈ How Does DIAG-Sapienza Work?
At its core, our system leverages sophisticated prompt engineering combined with targeted fine-tuning. This structured approach guides the large language model to break down biased text into measurable components. By structuring the input and guiding the output, we achieve a robust joint detection and classification capability.
The result? A highly accurate, interpretable tool that academic researchers and industry developers can use to audit NLP systems for gender bias, ensuring that the next generation of AI is fairer and more equitable.
π Read the full technical details and methodology here: [DIAG-Sapienza at GSI:detect] (https://aclanthology.org/2026.evalita-1.12/)