Web scraping and text content extraction - Beginner tutorial for Python and the command-line

Published: 25 February 2021
on channel: Trafilatura Web Scraping
4,263
43

Learn how to easily download web pages and extract main text and metadata using a free #Python scraping library, on the shell or with a custom script. Web feeds discovery included.

This introductory level tutorial starts with the installation on Windows and MacOS, demonstrates the command-line interface, and shows how the #scraping functions are run in Python: main text extraction, metadata extraction, language filtering, and link discovery with RSS feeds.

--
Table of Contents:
0:00 Installation for Mac
0:50 Installation for Windows
1:36 Getting started on the command-line
2:50 Usage with Python
5:50 Find ATOM/RSS feeds
6:55 Extract metadata (title,, date, author, etc.)
7:49 Baseline extraction
--

Trafilatura is a freely available scraping tool which seamlessly downloads, parses, and scrapes web page data while preserving parts of the text formatting and page structure. It features parallel online and offline processing, robust and efficient extraction, link discovery in feeds and sitemaps, as well as language detection.
The output can be converted to different formats: TXT, CSV, JSON, XML and XML-TEI.

▶️ More information
Documentation: https://trafilatura.readthedocs.org
Blog posts: https://adrien.barbaresi.eu/blog/tag/...
Code: https://github.com/adbar/trafilatura

#webscraping​​ #pythonTutorial #pythonBeginners


On this page of the site you can watch the video online Web scraping and text content extraction - Beginner tutorial for Python and the command-line with a duration of hours minute second in good quality, which was uploaded by the user Trafilatura Web Scraping 25 February 2021, share the link with friends and acquaintances, this video has already been watched 4,263 times on youtube and it was liked by 43 viewers. Enjoy your viewing!