Web scraping and text content extraction - Beginner tutorial for Python and the command-line

Опубликовано: 25 Февраль 2021
на канале: Trafilatura Web Scraping
4,263
43

Learn how to easily download web pages and extract main text and metadata using a free #Python scraping library, on the shell or with a custom script. Web feeds discovery included.

This introductory level tutorial starts with the installation on Windows and MacOS, demonstrates the command-line interface, and shows how the #scraping functions are run in Python: main text extraction, metadata extraction, language filtering, and link discovery with RSS feeds.

--
Table of Contents:
0:00 Installation for Mac
0:50 Installation for Windows
1:36 Getting started on the command-line
2:50 Usage with Python
5:50 Find ATOM/RSS feeds
6:55 Extract metadata (title,, date, author, etc.)
7:49 Baseline extraction
--

Trafilatura is a freely available scraping tool which seamlessly downloads, parses, and scrapes web page data while preserving parts of the text formatting and page structure. It features parallel online and offline processing, robust and efficient extraction, link discovery in feeds and sitemaps, as well as language detection.
The output can be converted to different formats: TXT, CSV, JSON, XML and XML-TEI.

▶️ More information
Documentation: https://trafilatura.readthedocs.org
Blog posts: https://adrien.barbaresi.eu/blog/tag/...
Code: https://github.com/adbar/trafilatura

#webscraping​​ #pythonTutorial #pythonBeginners


На этой странице сайта вы можете посмотреть видео онлайн Web scraping and text content extraction - Beginner tutorial for Python and the command-line длительностью часов минут секунд в хорошем качестве, которое загрузил пользователь Trafilatura Web Scraping 25 Февраль 2021, поделитесь ссылкой с друзьями и знакомыми, на youtube это видео уже посмотрели 4,263 раз и оно понравилось 43 зрителям. Приятного просмотра!