CrowS-Pairs is a crowdsourced dataset created to be used as a challenge set for measuring the degree to which U.S. stereotypical biases are present in large pretrained masked language models such as
BERT. The dataset consists of 1,508 examples that cover stereotypes dealing with nine type of social bias. Each example consists of a pair of sentences, where one sentence is always about a historically disadvantaged group in the United States and the other sentence is about a contrasting advantaged group. The sentence about a historically disadvantaged group can demonstrate or violate a stereotype. The paired sentence is a minimal edit of the first sentence: The only words that change between them are those that identify the group.
We collected this data through Amazon Mechanical Turk, where each example was written by a crowdworker and then validated by five other crowdworkers. We required all workers to be in the United States, to have completed at least 5,000 HITs, and to have greater than 98% acceptance rate. We use the Fair Work tool \citep{fairwork} to ensure a minimum of $15 hourly wage.