Description:
NLP has helped massively scale-up previously small-scale content analyses. Many social scientists train NLP classifiers and then measure social constructs (e.g sentiment) for millions of unlabeled documents which are then used as variables in downstream causal analyses. However, there are many points when one can make hard (non-probabilistic) or soft (probabilistic) assumptions in pipelines that use text classifiers: (a) adjudicating training labels from multiple annotators, (b) training supervised classifiers, and (c) aggregating individual-level classifications at inference time. In practice, propagating these hard versus soft choices down the pipeline can dramatically change the values of final social measurements. In this tutorial, we will walk through data and Python code of a real-world social science research pipeline that uses NLP classifiers to infer many users’ aggregate “moral outrage” expression on Twitter. Along the way, we will quantify the sensitivity of our pipeline to these hard versus soft choices.
Tutorial host: Katherine Keith
Slides: https://docs.google.com/presentation/...
Code: https://colab.research.google.com/dri...
Github: https://github.com/kakeith/tutorial20...
This is part of a larger tutorial series, NLP+CSS 201: Beyond the basics, which is organized by Ian Stewart and Katherine Keith. Website: https://nlp-css-201-tutorials.github....
Auf dieser Seite können Sie das Online-Video Tutorial 9: Aggregated Classification Pipelines: Propagating Probabilistic Assumptions mit der Dauer stunde minuten sekunde in guter Qualität ansehen, das der Benutzer NLP and CSS 201: Beyond the Basics 30 März 2022 hochgeladen hat, den Link mit Freunden und Bekannten teilen, dieses Video wurde auf Youtube bereits 344 Mal angesehen und es wurde von 6 den Zuschauern gefallen. Viel Spaß beim Betrachtenden Zuschauern gefallen!