Research

Observational Equivalence of LLM and Human Annotation Paper

(with Kentaro Nakamura and Jing Ling Tan)

We show that LLM and human coding are observationally equivalent in annotation quality: this equivalence results from both agreement and disagrement. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. Thus, there is little empirical basis for preferring human coding on annotation quality alone