*Web Crawling: A Practical Introduction in Python*
Presentation as part of training on April 27th, 2021
by Jaren Haber, PhD, Postdoctoral Fellow
Massive Data Institute, Georgetown University
*Goals of this Presentation:*
- Understand how web-crawling and -scraping are useful for digital data collection
- Build intuitions around the uses and limits of:
- APIs (Application Programming Interfaces)
- Exploiting website structure (HTML/CSS)
- Scalable crawling
- Be familiar with common problems in web-crawling and their fixes, like:
- Nested websites -- vertical crawling (link extraction)
- Getting blocked -- polite pauses
- Gain practice with:
- Collecting domains to scrape
- Scalable and non-scalable website scraping
- Parsing website text (with BeautifulSoup)
wget, Requests, and Scrapy
*Acknowledgments:*
- D-Lab at the University of California, Berkeley
- Summer Institute in Computational Social Science