๐ท๏ธ
STUDENT NOTES ยท DAY 21 OF 24
The Friendly Spider ๐ท๏ธ
Every web page is data wearing a costume called HTML. Today you learn to unwrap it โ politely.
1HTML is a tree of tags
๐
Py: Tags inside tags inside tags โ a family tree. BeautifulSoup climbs it for you.
2Soup time ๐
๐ค
Bittu: requests DOWNLOADS the page (Day 15 skill). BeautifulSoup PARSES it. Two tools, one heist โ a legal one.
3find_all โ harvest everything
4Scrape POLITELY ๐
๐ฆ
Professor Hoot: Check robots.txt (the site's rules for bots), add delays with time.sleep, and never hammer a server. Good spiders are welcome; rude ones get blocked.
๐
Bugly: Scrape a site 1000 times a second and they'll block your IP. Then who's laughing? โฆme. Always me.
5Quick quiz! ๐ง
The right order for scraping isโฆ
๐ก Download first, parse second. Soup can't parse what you haven't fetched.
robots.txt tells youโฆ
๐ก It's the site's rulebook for crawlers โ read it, respect it.
6Your mission ๐
Mission 1: Paste the class demo HTML into a string and extract the h1 with soup.find.
Mission 2: Print every link's text AND its href from the demo page.
Mission 3: Visit any site's /robots.txt in your browser and read what it allows. Detective work!
Day 21 done โ friendly spider certified! ๐ท๏ธ
Next: automation โ code that does your chores while you sleep.
Go to Day 22: Automation โ