Tutorial 9: Aggregated Classification Pipelines: Propagating Probabilistic Assumptions

Publicado em: 30 Março 2022
no canal de: NLP and CSS 201: Beyond the Basics
344
6

Description:

NLP has helped massively scale-up previously small-scale content analyses. Many social scientists train NLP classifiers and then measure social constructs (e.g sentiment) for millions of unlabeled documents which are then used as variables in downstream causal analyses. However, there are many points when one can make hard (non-probabilistic) or soft (probabilistic) assumptions in pipelines that use text classifiers: (a) adjudicating training labels from multiple annotators, (b) training supervised classifiers, and (c) aggregating individual-level classifications at inference time. In practice, propagating these hard versus soft choices down the pipeline can dramatically change the values of final social measurements. In this tutorial, we will walk through data and Python code of a real-world social science research pipeline that uses NLP classifiers to infer many users’ aggregate “moral outrage” expression on Twitter. Along the way, we will quantify the sensitivity of our pipeline to these hard versus soft choices.

Tutorial host: Katherine Keith

Slides: https://docs.google.com/presentation/...

Code: https://colab.research.google.com/dri...

Github: https://github.com/kakeith/tutorial20...

This is part of a larger tutorial series, NLP+CSS 201: Beyond the basics, which is organized by Ian Stewart and Katherine Keith. Website: https://nlp-css-201-tutorials.github....


Nesta página do site você pode assistir ao vídeo on-line Tutorial 9: Aggregated Classification Pipelines: Propagating Probabilistic Assumptions duração hora minuto segundo em boa qualidade , que foi baixado pelo usuário NLP and CSS 201: Beyond the Basics 30 Março 2022, compartilhe o link com seus amigos e conhecidos, no youtube este vídeo já foi visto 344 vezes e gostou 6 espectadores. Boa visualização!