Taking Turns: Natural Language Inference on Interview Data

Date:

Poster presentation at Text as Data 2026, Berkeley, CA, USA

Abstract:

Natural language inference (NLI) is emerging as a promising text-classification method for the social sciences, as it enables researchers to efficiently classify text using a small set of short, declarative hypotheses. Its performance, however, depends heavily on corpus type. However, while the method has mostly been applied to ordered long-form text on one hand, or disordered short text (primarily social media text) on the other, interview transcripts do not fall into any of these types of text. Like social media text, they are colloquial, disfluent, and unedited. However, they are also composed of long, continuous speech, structurally similar to long-form edited text. Interview transcripts thus combine the extended premises of edited prose with the disorder of unstructured text. This paper reports on deploying NLI across 44 in-depth interviews with LGBTQ activists evaluating the marriage equality campaign a decade after Obergefell v. Hodges. Substantively, applying an independently-scored hypothesis set makes the multi-thematic texture of interviews visible and measurable. Across roughly two thousand coded premises, evaluative claims co-occur in patterned groups rather than incidentally, a structure a single-label approach such as topic modeling would flatten. Methodologically, that same density is the central constraint on classification, diluting entailment signal within long, multi-theme premises, alongside recurring patterns in how the model handles affective versus propositional claims and speaker attribution. The paper offers both a substantive account of retrospective activist evaluation and a transferable protocol for applying NLI to interviews, oral histories, and other naturalistic prose.