Building a Classifier for Rating Condescension in NYT Headlines

Members: Shreyans Sethi, Ayush Sehgal, Gayatri Babel


Project Overview:

For our Annotation Project, we will evaluate how condescending New York Times article headlines and summaries are when comparing articles written about the Global North and Global South. For the purpose of this project, Global North will encompass the US, Europe, and Australia while the Global South will refer to regions of Latin America, Asia, Africa, and Oceania. Thus, we will only be using articles that clearly indicate relation to countries within the Global North or the Global South (ex. generic articles like “Nintendo Switch Sales Decline” will not be considered, but “America Eases Tensions in Africa” will). We plan on annotating documents using a 1 - 5 scale: 1 being not condescending at all and 5 being very condescending. On this scale, articles with a score of “1” would be unbiased and would not hint at any superiority or prejudice over the country in question. Articles with a “5” score could use language that indicates patronization, superiority, and/or dislike in an obvious and clear manner (as opposed to being subtle). The guidelines used to evaluate condescending tones are outlined here:

Guidelines for Annotating Condescension Scores

We are using the publicly available archive of all New York Times articles since 1981. This data is publicly available and under the NYT’s own copyright laws, we “are permitted to [...] use New York Times headlines with links back to the articles located on NYTimes.com”. We only plan to use the article headlines and descriptions on the archive page, we will create a works cited section with the links of all articles in our dataset.


GitHub Repository containing Data Splits, Code for the Classifier, and Analysis of Performance.