Skip to main content

Section 14.13 Code explaining activity

Subsection 14.13.1 Relevant tags

Here’s the relevant tag from https://www.si.umich.edu/people/barbara-ericson:
Figure 14.13.1. News article links on Dr. Ericson’s page

Activity 14.13.1.

Write down your best guess of what the code does.

Activity 14.13.2.

You can run the code below and see what happens.
#Get the webpage
# Load libraries for web scraping
from bs4 import BeautifulSoup
import requests
# Get a soup from a URL
url = 'https://www.si.umich.edu/people/barbara-ericson'
r = requests.get(url)
soup = BeautifulSoup(r.content, 'html.parser')

#Extract info from the webpage
# Get all tags of a certain type from the soup
tags = soup.find_all('a', class_='item-teaser--heading-link')
# Collect info from the tags
collect_info = []
for tag in tags:
  # Get link from tag
  info = tag.get('href')
  collect_info.append(info)

#Do something with the info
# Get a soup from multiple URLs
base_url = 'https://www.si.umich.edu/'
endings = collect_info
for ending in endings:
    url = base_url + ending
    r = requests.get(url)
    soup = BeautifulSoup(r.content, 'html.parser')

    # Get all tags of a certain type from the soup
    tags = soup.find_all('p')
    # Collect info from the tags
    collect_info = []
    for tag in tags:
        # Get text from tag
        info = tag.text
        collect_info.append(info)

    # Print the info
    print(collect_info)
You have attempted of activities on this page.