Web scraping and text content extraction - Beginner tutorial for Python and the command-line

Publié le: 25 février 2021
sur la chaîne: Trafilatura Web Scraping
4,263
43

Learn how to easily download web pages and extract main text and metadata using a free #Python scraping library, on the shell or with a custom script. Web feeds discovery included.

This introductory level tutorial starts with the installation on Windows and MacOS, demonstrates the command-line interface, and shows how the #scraping functions are run in Python: main text extraction, metadata extraction, language filtering, and link discovery with RSS feeds.

--
Table of Contents:
0:00 Installation for Mac
0:50 Installation for Windows
1:36 Getting started on the command-line
2:50 Usage with Python
5:50 Find ATOM/RSS feeds
6:55 Extract metadata (title,, date, author, etc.)
7:49 Baseline extraction
--

Trafilatura is a freely available scraping tool which seamlessly downloads, parses, and scrapes web page data while preserving parts of the text formatting and page structure. It features parallel online and offline processing, robust and efficient extraction, link discovery in feeds and sitemaps, as well as language detection.
The output can be converted to different formats: TXT, CSV, JSON, XML and XML-TEI.

▶️ More information
Documentation: https://trafilatura.readthedocs.org
Blog posts: https://adrien.barbaresi.eu/blog/tag/...
Code: https://github.com/adbar/trafilatura

#webscraping​​ #pythonTutorial #pythonBeginners


Sur cette page du site, vous pouvez voir la vidéo en ligne Web scraping and text content extraction - Beginner tutorial for Python and the command-line durée heure minute seconde en bonne qualité , qui a été Téléchargé par l'utilisateur Trafilatura Web Scraping 25 février 2021, Partagez le lien avec vos amis et connaissances, sur youtube cette vidéo a déjà été regardée 4,263 fois et il a aimé 43 téléspectateurs. Bon visionnage!