前幾天已經可以把網站上的資料抓下來了。
但是實際做資料分析時,我發現「抓到資料」其實只是第一步。
例如網站上的價格可能是:
£51.77
Availability 可能是:
In stock (22 available)
Rating 可能是:
Three
如果直接把這些文字存進 CSV,之後要用 Pandas 做統計就會比較麻煩。
所以今天想開始處理一個很重要的問題:
如何把爬蟲抓到的網頁文字,整理成真正可以分析的資料。
一、為什麼需要資料清理?
假設爬蟲抓到:
Price
£51.77
£35.02
£17.46
£44.18
這些看起來都是價格,但對 Python 來說它們其實是:
字串 String
如果直接拿來計算:
sum(prices)
就會發生問題。
因為 Python 不知道:
£51.77
是一個可以計算的數字。
所以我們需要先把:
£51.77
變成:
51.77
也就是:
網頁資料
↓
文字
↓
清理
↓
數字
↓
可以分析
二、先看看原始價格
今天一樣使用前幾天的 Books to Scrape。
import requests
from bs4 import BeautifulSoup
url = "https://books.toscrape.com/"
response = requests.get(url)
soup = BeautifulSoup(response.text, "html.parser")
books = soup.select("article.product_pod")
for book in books[:5]:
price = book.select_one(".price_color").text.strip()
print(price)
可能會看到:
£51.77
£53.74
£50.10
£47.82
£54.23
現在的 price 是字串。
可以確認:
print(type(price))
結果:
<class 'str'>
三、把價格轉成數字
最簡單的方法就是先把 £ 移除。
price_text = "£51.77"
price = price_text.replace("£", "")
print(price)
結果:
51.77
但現在還是字串。
print(type(price))
結果:
<class 'str'>
所以還需要:
price = float(price)
完整寫法:
price_text = "£51.77"
price = price_text.replace("£", "")
price = float(price)
print(price)
print(type(price))
結果:
51.77
<class 'float'>
現在就可以進行數學運算了。
例如:
print(price * 2)
結果:
103.54
四、把清理價格寫成 Function
如果之後有 1000 筆資料,總不可能每一筆都自己寫一次。
所以可以把這個功能做成 Function。
def clean_price(price_text):
price_text = price_text.replace("£", "")
return float(price_text)
測試:
print(clean_price("£51.77"))
print(clean_price("£35.02"))
結果:
51.77
35.02
這樣之後只需要:
price = clean_price(price_text)
就可以了。
五、處理 Availability
價格整理好了,再來看看庫存資料。
availability = book.select_one(
".availability"
).text.strip()
print(availability)
可能會看到:
In stock (22 available)
如果我要分析庫存數量,其實真正需要的是:
22
所以也需要清理。
六、把庫存數量取出來
可以使用 replace()。
text = "In stock (22 available)"
text = text.replace("In stock", "")
text = text.replace("(", "")
text = text.replace(")", "")
text = text.replace("available", "")
text = text.strip()
print(text)
結果:
22
然後:
stock = int(text)
就可以轉成整數。
七、把庫存清理也做成 Function
def clean_stock(stock_text):
stock_text = stock_text.replace(
"In stock", ""
)
stock_text = stock_text.replace(
"(", ""
)
stock_text = stock_text.replace(
")", ""
)
stock_text = stock_text.replace(
"available", ""
)
stock_text = stock_text.strip()
return int(stock_text)
測試:
print(clean_stock("In stock (22 available)"))
結果:
22
這樣之後就可以直接拿來做庫存分析。
八、Rating 也可以整理
Books to Scrape 的 Rating 是:
One
Two
Three
Four
Five
但如果要分析平均評分,我們希望它變成:
1
2
3
4
5
所以可以建立一個 Dictionary。
rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}
例如:
rating_text = "Three"
rating = rating_map[rating_text]
print(rating)
結果:
3
九、把三種資料一起整理
現在我們有三個 Function:
def clean_price(price_text):
price_text = price_text.replace("£", "")
return float(price_text)
def clean_stock(stock_text):
stock_text = stock_text.replace(
"In stock", ""
)
stock_text = stock_text.replace(
"(", ""
)
stock_text = stock_text.replace(
")", ""
)
stock_text = stock_text.replace(
"available", ""
)
stock_text = stock_text.strip()
return int(stock_text)
Rating:
rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}
接著就可以放進爬蟲。
十、完整實作
今天我把前幾天的爬蟲和今天的資料清理結合起來。
import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import time
session = requests.Session()
session.headers.update({
"User-Agent": "Mozilla/5.0"
})
def clean_price(price_text):
price_text = price_text.replace("£", "")
return float(price_text)
def clean_stock(stock_text):
stock_text = stock_text.replace(
"In stock", ""
)
stock_text = stock_text.replace(
"(", ""
)
stock_text = stock_text.replace(
")", ""
)
stock_text = stock_text.replace(
"available", ""
)
stock_text = stock_text.strip()
return int(stock_text)
rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}
url = "https://books.toscrape.com/"
all_books = []
while url:
print("正在抓取:", url)
try:
response = session.get(
url,
timeout=10
)
response.raise_for_status()
soup = BeautifulSoup(
response.text,
"html.parser"
)
books = soup.select(
"article.product_pod"
)
for book in books:
title = book.h3.a["title"]
price_text = book.select_one(
".price_color"
).text.strip()
availability_text = book.select_one(
".availability"
).text.strip()
rating_text = book.select_one(
"p.star-rating"
)["class"][1]
price = clean_price(
price_text
)
stock = clean_stock(
availability_text
)
rating = rating_map[
rating_text
]
all_books.append({
"title": title,
"price": price,
"stock": stock,
"rating": rating
})
next_button = soup.select_one(
"li.next a"
)
if next_button:
next_url = next_button.get("href")
url = urljoin(
url,
next_url
)
else:
url = None
time.sleep(1)
except requests.RequestException as e:
print("Request 發生錯誤:", e)
break
print()
print("爬蟲完成")
print("總共取得:", len(all_books), "筆資料")
print()
print(all_books[:3])
現在資料就不再只是網站上的原始文字。
例如原本:
price = "£51.77"
stock = "In stock (22 available)"
rating = "Three"
會變成:
price = 51.77
stock = 22
rating = 3
這樣就開始有「資料分析用資料」的樣子了。
十一、現在就可以直接做一些分析
例如找出價格最高的書:
most_expensive = max(
all_books,
key=lambda x: x["price"]
)
print(most_expensive)
也可以找價格最低:
cheapest = min(
all_books,
key=lambda x: x["price"]
)
print(cheapest)
計算平均價格:
average_price = sum(
book["price"]
for book in all_books
) / len(all_books)
print("平均價格:", average_price)
計算平均評分:
average_rating = sum(
book["rating"]
for book in all_books
) / len(all_books)
print("平均評分:", average_rating)
這時候就可以看到一個很明顯的差別:
以前我們只是:
Scraping → 把資料抓下來
現在開始變成:
Scraping → Cleaning → Analysis
十二、今天的實作流程
今天實際完成的流程:
網站
↓
Requests
↓
BeautifulSoup
↓
抓取原始文字
↓
清理資料
↓
String → Float / Integer
↓
建立結構化 Dictionary
↓
進行基本分析
這其實就是之後資料分析會一直使用的流程。
尤其是這幾個步驟:
£51.77
↓
51.77
In stock (22 available)
↓
22
Three
↓
3
看起來只是把文字改掉,但這一步非常重要。
因為如果資料型態不正確,之後使用 Pandas、SQLite 或 Matplotlib 做分析時,都可能遇到問題。