Description:
NLP has helped massively scale-up previously small-scale content analyses. Many social scientists train NLP classifiers and then measure social constructs (e.g sentiment) for millions of unlabeled documents which are then used as variables in downstream causal analyses. However, there are many points when one can make hard (non-probabilistic) or soft (probabilistic) assumptions in pipelines that use text classifiers: (a) adjudicating training labels from multiple annotators, (b) training supervised classifiers, and (c) aggregating individual-level classifications at inference time. In practice, propagating these hard versus soft choices down the pipeline can dramatically change the values of final social measurements. In this tutorial, we will walk through data and Python code of a real-world social science research pipeline that uses NLP classifiers to infer many users’ aggregate “moral outrage” expression on Twitter. Along the way, we will quantify the sensitivity of our pipeline to these hard versus soft choices.
Tutorial host: Katherine Keith
Slides: https://docs.google.com/presentation/...
Code: https://colab.research.google.com/dri...
Github: https://github.com/kakeith/tutorial20...
This is part of a larger tutorial series, NLP+CSS 201: Beyond the basics, which is organized by Ian Stewart and Katherine Keith. Website: https://nlp-css-201-tutorials.github....
En esta página del sitio puede ver el video en línea Tutorial 9: Aggregated Classification Pipelines: Propagating Probabilistic Assumptions de Duración hora minuto segunda en buena calidad , que subió el usuario NLP and CSS 201: Beyond the Basics 30 marzo 2022, comparta el enlace con amigos y conocidos, en youtube este video ya ha sido visto 344 veces y le gustó 6 a los espectadores. Disfruta viendo!