Approaching the largest ‘API’: extracting information from the Internet with Python

Jonathan E. Germann

Code4Lib Journal (Feb 2018)

Approaching the largest ‘API’: extracting information from the Internet with Python

Jonathan E. Germann

Affiliations

Jonathan E. Germann

Journal volume & issue: no. 39

Abstract

Read online

This article explores the need for libraries to algorithmically access and manipulate the world’s largest API: the Internet. The billions of pages on the ‘Internet API’ (HTTP, HTML, CSS, XPath, DOM, etc.) are easily accessible and manipulable. Libraries can assist in creating meaning through the datafication of information on the world wide web. Because most information is created for human consumption, some programming is required for automated extraction. Python is an easy-to-learn programming language with extensive packages and community support for web page automation. Four packages (Urllib, Selenium, BeautifulSoup, Scrapy) in Python can automate almost any web page for all sized projects. An example warrant data project is explained to illustrate how well Python packages can manipulate web pages to create meaning through assembling custom datasets.

Published in Code4Lib Journal

ISSN: 1940-5758 (Online)
Publisher: Code4Lib
Country of publisher: United States
LCC subjects: Bibliography. Library science. Information resources
Website: https://journal.code4lib.org

About the journal