iT邦幫忙

2026 iThome 鐵人賽

DAY 3
0
自我挑戰組

Python 爬蟲與資料分析實戰:從網路資料到資訊視覺化系列 第 15 篇

Day 15|爬蟲資料的進一步處理:清理文字、價格與欄位

  • 分享至 

  • xImage
  •  

前幾天已經可以把網站上的資料抓下來了。

但是實際做資料分析時,我發現「抓到資料」其實只是第一步。

例如網站上的價格可能是:

£51.77

Availability 可能是:

In stock (22 available)

Rating 可能是:

Three

如果直接把這些文字存進 CSV,之後要用 Pandas 做統計就會比較麻煩。

所以今天想開始處理一個很重要的問題:

如何把爬蟲抓到的網頁文字,整理成真正可以分析的資料。

一、為什麼需要資料清理?

假設爬蟲抓到:

Price
£51.77
£35.02
£17.46
£44.18

這些看起來都是價格,但對 Python 來說它們其實是:

字串 String

如果直接拿來計算:

sum(prices)

就會發生問題。

因為 Python 不知道:

£51.77

是一個可以計算的數字。

所以我們需要先把:

£51.77

變成:

51.77

也就是:

網頁資料
↓
文字
↓
清理
↓
數字
↓
可以分析
二、先看看原始價格

今天一樣使用前幾天的 Books to Scrape。

import requests
from bs4 import BeautifulSoup

url = "https://books.toscrape.com/"

response = requests.get(url)

soup = BeautifulSoup(response.text, "html.parser")

books = soup.select("article.product_pod")

for book in books[:5]:

price = book.select_one(".price_color").text.strip()

print(price)

可能會看到:

£51.77
£53.74
£50.10
£47.82
£54.23

現在的 price 是字串。

可以確認:

print(type(price))

結果:

<class 'str'>
三、把價格轉成數字

最簡單的方法就是先把 £ 移除。

price_text = "£51.77"

price = price_text.replace("£", "")

print(price)

結果:

51.77

但現在還是字串。

print(type(price))

結果:

<class 'str'>

所以還需要:

price = float(price)

完整寫法:

price_text = "£51.77"

price = price_text.replace("£", "")

price = float(price)

print(price)
print(type(price))

結果:

51.77
<class 'float'>

現在就可以進行數學運算了。

例如:

print(price * 2)

結果:

103.54
四、把清理價格寫成 Function

如果之後有 1000 筆資料,總不可能每一筆都自己寫一次。

所以可以把這個功能做成 Function。

def clean_price(price_text):

price_text = price_text.replace("£", "")

return float(price_text)

測試:

print(clean_price("£51.77"))
print(clean_price("£35.02"))

結果:

51.77
35.02

這樣之後只需要:

price = clean_price(price_text)

就可以了。

五、處理 Availability

價格整理好了,再來看看庫存資料。

availability = book.select_one(
".availability"
).text.strip()

print(availability)

可能會看到:

In stock (22 available)

如果我要分析庫存數量,其實真正需要的是:

22

所以也需要清理。

六、把庫存數量取出來

可以使用 replace()。

text = "In stock (22 available)"

text = text.replace("In stock", "")
text = text.replace("(", "")
text = text.replace(")", "")
text = text.replace("available", "")

text = text.strip()

print(text)

結果:

22

然後:

stock = int(text)

就可以轉成整數。

七、把庫存清理也做成 Function
def clean_stock(stock_text):

stock_text = stock_text.replace(
    "In stock", ""
)

stock_text = stock_text.replace(
    "(", ""
)

stock_text = stock_text.replace(
    ")", ""
)

stock_text = stock_text.replace(
    "available", ""
)

stock_text = stock_text.strip()

return int(stock_text)

測試:

print(clean_stock("In stock (22 available)"))

結果:

22

這樣之後就可以直接拿來做庫存分析。

八、Rating 也可以整理

Books to Scrape 的 Rating 是:

One
Two
Three
Four
Five

但如果要分析平均評分,我們希望它變成:

1
2
3
4
5

所以可以建立一個 Dictionary。

rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}

例如:

rating_text = "Three"

rating = rating_map[rating_text]

print(rating)

結果:

3
九、把三種資料一起整理

現在我們有三個 Function:

def clean_price(price_text):

price_text = price_text.replace("£", "")

return float(price_text)

def clean_stock(stock_text):

stock_text = stock_text.replace(
    "In stock", ""
)

stock_text = stock_text.replace(
    "(", ""
)

stock_text = stock_text.replace(
    ")", ""
)

stock_text = stock_text.replace(
    "available", ""
)

stock_text = stock_text.strip()

return int(stock_text)

Rating:

rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}

接著就可以放進爬蟲。

十、完整實作

今天我把前幾天的爬蟲和今天的資料清理結合起來。

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import time

session = requests.Session()

session.headers.update({
"User-Agent": "Mozilla/5.0"
})

def clean_price(price_text):

price_text = price_text.replace("£", "")

return float(price_text)

def clean_stock(stock_text):

stock_text = stock_text.replace(
    "In stock", ""
)

stock_text = stock_text.replace(
    "(", ""
)

stock_text = stock_text.replace(
    ")", ""
)

stock_text = stock_text.replace(
    "available", ""
)

stock_text = stock_text.strip()

return int(stock_text)

rating_map = {
"One": 1,
"Two": 2,
"Three": 3,
"Four": 4,
"Five": 5
}

url = "https://books.toscrape.com/"

all_books = []

while url:

print("正在抓取:", url)

try:

    response = session.get(
        url,
        timeout=10
    )

    response.raise_for_status()

    soup = BeautifulSoup(
        response.text,
        "html.parser"
    )

    books = soup.select(
        "article.product_pod"
    )

    for book in books:

        title = book.h3.a["title"]

        price_text = book.select_one(
            ".price_color"
        ).text.strip()

        availability_text = book.select_one(
            ".availability"
        ).text.strip()

        rating_text = book.select_one(
            "p.star-rating"
        )["class"][1]

        price = clean_price(
            price_text
        )

        stock = clean_stock(
            availability_text
        )

        rating = rating_map[
            rating_text
        ]

        all_books.append({
            "title": title,
            "price": price,
            "stock": stock,
            "rating": rating
        })

    next_button = soup.select_one(
        "li.next a"
    )

    if next_button:

        next_url = next_button.get("href")

        url = urljoin(
            url,
            next_url
        )

    else:

        url = None

    time.sleep(1)

except requests.RequestException as e:

    print("Request 發生錯誤:", e)

    break

print()
print("爬蟲完成")
print("總共取得:", len(all_books), "筆資料")

print()
print(all_books[:3])

現在資料就不再只是網站上的原始文字。

例如原本:

price = "£51.77"
stock = "In stock (22 available)"
rating = "Three"

會變成:

price = 51.77
stock = 22
rating = 3

這樣就開始有「資料分析用資料」的樣子了。

十一、現在就可以直接做一些分析

例如找出價格最高的書:

most_expensive = max(
all_books,
key=lambda x: x["price"]
)

print(most_expensive)

也可以找價格最低:

cheapest = min(
all_books,
key=lambda x: x["price"]
)

print(cheapest)

計算平均價格:

average_price = sum(
book["price"]
for book in all_books
) / len(all_books)

print("平均價格:", average_price)

計算平均評分:

average_rating = sum(
book["rating"]
for book in all_books
) / len(all_books)

print("平均評分:", average_rating)

這時候就可以看到一個很明顯的差別:

以前我們只是:

Scraping → 把資料抓下來

現在開始變成:

Scraping → Cleaning → Analysis

十二、今天的實作流程

今天實際完成的流程:

網站
↓
Requests
↓
BeautifulSoup
↓
抓取原始文字
↓
清理資料
↓
String → Float / Integer
↓
建立結構化 Dictionary
↓
進行基本分析

這其實就是之後資料分析會一直使用的流程。

尤其是這幾個步驟:

£51.77
↓
51.77
In stock (22 available)
↓
22
Three
↓
3

看起來只是把文字改掉,但這一步非常重要。

因為如果資料型態不正確,之後使用 Pandas、SQLite 或 Matplotlib 做分析時,都可能遇到問題。


上一篇
Day 14|讓爬蟲更穩定:User-Agent、Session 與請求間隔
下一篇
Day 17|讓爬蟲更可靠:Retry 重試機制與錯誤處理
系列文
Python 爬蟲與資料分析實戰:從網路資料到資訊視覺化 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言