DCAI
← 返回全部动态
arXiv 人工智能规则精选09月24日 12:00

Queer inclusion in speech datasets: An audit and taxonomy of practical tensions

arXiv:2609.25491v1 Announce Type: new Abstract: In this paper, we examine speech datasets for their inclusion of LGBTQIA+, or queer, voices and provide a taxonomy of tensions to better understand why there is a lack of such voices in current speech technology datasets. Through an audit of six diverse speech datasets, we find that measurable queer representation is low (0-1.4% of speakers) - insufficient for robust disparity measurement. We take this community as a case study to consider what challenges and tensions are associated with collecting speech data from marginalized communities. For comparison, we audit an additional two datasets fro

阅读 arXiv 人工智能 原文 ↗