Web scraping and text content extraction - Beginner tutorial for Python and the command-line

Publicado em: 25 Fevereiro 2021
no canal de: Trafilatura Web Scraping
4,263
43

Learn how to easily download web pages and extract main text and metadata using a free #Python scraping library, on the shell or with a custom script. Web feeds discovery included.

This introductory level tutorial starts with the installation on Windows and MacOS, demonstrates the command-line interface, and shows how the #scraping functions are run in Python: main text extraction, metadata extraction, language filtering, and link discovery with RSS feeds.

--
Table of Contents:
0:00 Installation for Mac
0:50 Installation for Windows
1:36 Getting started on the command-line
2:50 Usage with Python
5:50 Find ATOM/RSS feeds
6:55 Extract metadata (title,, date, author, etc.)
7:49 Baseline extraction
--

Trafilatura is a freely available scraping tool which seamlessly downloads, parses, and scrapes web page data while preserving parts of the text formatting and page structure. It features parallel online and offline processing, robust and efficient extraction, link discovery in feeds and sitemaps, as well as language detection.
The output can be converted to different formats: TXT, CSV, JSON, XML and XML-TEI.

▶️ More information
Documentation: https://trafilatura.readthedocs.org
Blog posts: https://adrien.barbaresi.eu/blog/tag/...
Code: https://github.com/adbar/trafilatura

#webscraping​​ #pythonTutorial #pythonBeginners


Nesta página do site você pode assistir ao vídeo on-line Web scraping and text content extraction - Beginner tutorial for Python and the command-line duração hora minuto segundo em boa qualidade , que foi baixado pelo usuário Trafilatura Web Scraping 25 Fevereiro 2021, compartilhe o link com seus amigos e conhecidos, no youtube este vídeo já foi visto 4,263 vezes e gostou 43 espectadores. Boa visualização!