前面幾天我們拿到了不少資料:天氣預報、待辦清單。但拿到資料不等於知道答案。
想想這兩句話的差別:
A:「明天 06:00–18:00 多雲,25–30°C,降雨機率 60%;18:00–06:00 陰,23–26°C,降雨機率 30%。」
B:「明天白天有六成機率下雨,記得帶傘。」
A 是資料,B 是結論。使用者要的是 B。
而要從 A 變成 B,中間需要的就是今天的主題:整理、篩選、統計。
Python 的 sorted() 回傳一個新的排序過的 list(list.sort() 則是就地排序,會改掉原本的)。
numbers = [3, 1, 4, 1, 5, 9, 2, 6]
print(sorted(numbers)) # [1, 1, 2, 3, 4, 5, 6, 9]
print(sorted(numbers, reverse=True))# [9, 6, 5, 4, 3, 2, 1, 1]
print(numbers) # [3, 1, 4, 1, 5, 9, 2, 6] ← 原本的沒變
但我們的資料通常是 dict 組成的 list:
todos = [
{"title": "買牛奶", "priority": "中", "done": False, "days_left": 3},
{"title": "繳電費", "priority": "高", "done": False, "days_left": 1},
{"title": "整理房間", "priority": "低", "done": True, "days_left": 7},
{"title": "寫文章", "priority": "高", "done": False, "days_left": 0},
]
sorted(todos) # ❌ TypeError: '<' not supported between instances of 'dict'
dict 之間沒辦法直接比大小,所以要用 key 參數告訴 Python「依據什麼排序」:
# 依剩餘天數排序
by_deadline = sorted(todos, key=lambda t: t["days_left"])
for t in by_deadline:
print(f"{t['days_left']} 天:{t['title']}")
輸出:
0 天:寫文章
1 天:繳電費
3 天:買牛奶
7 天:整理房間
lambda t: t["days_left"] 是一個「匿名函式」,等同於:
def get_days_left(t):
return t["days_left"]
sorted(todos, key=get_days_left)
只是懶得另外定義一個函式,就直接寫在參數裡。lambda 參數: 回傳值,就這樣。
優先度是「高/中/低」,直接排會變成按筆劃或 Unicode 順序,不是我們要的。解法是先對應成數字:
PRIORITY_ORDER = {"高": 0, "中": 1, "低": 2}
by_priority = sorted(todos, key=lambda t: PRIORITY_ORDER[t["priority"]])
如果要「先照優先度,同優先度再照剩餘天數」,key 回傳一個 tuple:
smart_order = sorted(
todos,
key=lambda t: (
t["done"], # 未完成(False=0)排前面
PRIORITY_ORDER[t["priority"]], # 再照優先度
t["days_left"], # 再照剩餘天數
),
)
for t in smart_order:
mark = "✓" if t["done"] else "□"
print(f"{mark} [{t['priority']}] {t['title']}(剩 {t['days_left']} 天)")
輸出:
□ [高] 寫文章(剩 0 天)
□ [高] 繳電費(剩 1 天)
□ [中] 買牛奶(剩 3 天)
✓ [低] 整理房間(剩 7 天)
Python 比較 tuple 時是逐項比對:先比第一項,相同才比第二項,以此類推。這招讓多重排序變得非常簡潔。
(順帶一提,False < True,因為布林值本質上就是 0 和 1。)
Day 6 學過的 list comprehension 就是最好用的篩選工具:
# 未完成的
pending = [t for t in todos if not t["done"]]
# 高優先度且未完成的
urgent = [t for t in todos if t["priority"] == "高" and not t["done"]]
# 三天內到期的
soon = [t for t in todos if t["days_left"] <= 3]
比 filter(lambda t: ..., todos) 好讀太多,Python 社群也偏好這種寫法。
有時候我們只想知道「有沒有」、「是不是全部」:
has_urgent = any(t["priority"] == "高" and not t["done"] for t in todos)
all_done = all(t["done"] for t in todos)
print(f"有緊急事項嗎:{has_urgent}") # True
print(f"全部完成了嗎:{all_done}") # False
any() 只要有一個 True 就回傳 True,all() 要全部都 True。
而且它們有短路求值——any() 找到第一個 True 就停了,不會跑完整個清單。資料量大時差很多。
first_urgent = next(
(t for t in todos if t["priority"] == "高" and not t["done"]),
None, # 找不到時的預設值
)
if first_urgent:
print(f"最該先處理:{first_urgent['title']}")
記得加第二個參數當預設值,不然找不到時會拋 StopIteration。
temps = [26, 28, 31, 29, 25, 27, 30]
print(f"最低:{min(temps)}°C")
print(f"最高:{max(temps)}°C")
print(f"平均:{sum(temps) / len(temps):.1f}°C")
print(f"總筆數:{len(temps)}")
⚠️ sum(temps) / len(temps) 在空清單時會 ZeroDivisionError。實務上要先檢查:
average = sum(temps) / len(temps) if temps else 0
Python 內建 statistics,不用裝 numpy:
import statistics
temps = [26, 28, 31, 29, 25, 27, 30]
print(statistics.mean(temps)) # 28 平均數
print(statistics.median(temps)) # 28 中位數
print(statistics.stdev(temps)) # 2.16 標準差
median 在有極端值時比 mean 更能代表「典型情況」。
找「最…的那一筆資料」,而不只是最大的數字:
most_urgent = min(todos, key=lambda t: t["days_left"])
print(f"最急的是:{most_urgent['title']}") # 寫文章
from collections import Counter
priorities = [t["priority"] for t in todos]
counts = Counter(priorities)
print(counts) # Counter({'高': 2, '中': 1, '低': 1})
print(counts["高"]) # 2
print(counts.most_common(1)) # [('高', 2)]
Counter 是 dict 的子類別,專門用來數東西。比自己寫迴圈 counts[x] = counts.get(x, 0) + 1 乾淨多了。
如果要「照某個欄位分組」,我自己習慣用 defaultdict:
from collections import defaultdict
by_priority = defaultdict(list)
for t in todos:
by_priority[t["priority"]].append(t["title"])
for priority in ("高", "中", "低"):
items = by_priority[priority]
if items:
print(f"{priority}:{', '.join(items)}")
輸出:
高:繳電費, 寫文章
中:買牛奶
低:整理房間
defaultdict(list) 的意思是「取一個不存在的 key 時,自動給我一個空 list」,省掉了 if key not in d: d[key] = [] 這行。
現在把今天學的東西,套用到 Day 14 拿到的天氣資料上:
"""analysis/weather_insight.py"""
def analyze_weather(data):
"""把天氣預報分析成結構化的洞察。"""
periods = data["periods"]
if not periods:
return None
# 統計
rain_chances = [p["rain_probability"] for p in periods
if p["rain_probability"] is not None]
min_temps = [p["min_temp"] for p in periods if p["min_temp"] is not None]
max_temps = [p["max_temp"] for p in periods if p["max_temp"] is not None]
max_rain = max(rain_chances, default=0)
lowest = min(min_temps, default=None)
highest = max(max_temps, default=None)
# 找出最可能下雨的時段
rainiest = max(periods, key=lambda p: p["rain_probability"] or 0)
# 溫差
temp_range = (highest - lowest) if (highest and lowest) else 0
return {
"city": data["city"],
"max_rain": max_rain,
"lowest": lowest,
"highest": highest,
"temp_range": temp_range,
"rainiest_period": rainiest,
"will_rain": max_rain >= 40,
"big_temp_swing": temp_range >= 8,
"is_cold": lowest is not None and lowest <= 18,
"is_hot": highest is not None and highest >= 32,
}
def make_advice(insight):
"""把洞察轉成給人看的建議。"""
if insight is None:
return "目前沒有可用的天氣資料。"
advice = [f"【{insight['city']}】"]
# 主要結論
if insight["will_rain"]:
period = insight["rainiest_period"]
when = period["start"][5:16]
advice.append(
f"⚠️ {when} 前後降雨機率達 {period['rain_probability']}%,記得帶傘。"
)
else:
advice.append(f"☀️ 降雨機率最高只有 {insight['max_rain']}%,應該不用帶傘。")
# 溫度
advice.append(f"氣溫 {insight['lowest']}–{insight['highest']}°C。")
if insight["big_temp_swing"]:
advice.append(f"日夜溫差達 {insight['temp_range']} 度,建議洋蔥式穿搭。")
if insight["is_cold"]:
advice.append("早晚偏涼,記得加件外套。")
if insight["is_hot"]:
advice.append("白天偏熱,注意補充水分、避免長時間曝曬。")
return "\n".join(advice)
試跑:
import config
from tools.weather import get_weather
data = get_weather("臺北市", config.CWA_API_KEY)
insight = analyze_weather(data)
print(make_advice(insight))
輸出:
【臺北市】
⚠️ 09-30 06:00 前後降雨機率達 60%,記得帶傘。
氣溫 23–31°C。
日夜溫差達 8 度,建議洋蔥式穿搭。
**從一堆數字,變成三句可以行動的建議。**這就是分析的價值。
"""analysis/todo_insight.py"""
from collections import Counter, defaultdict
PRIORITY_ORDER = {"高": 0, "中": 1, "低": 2}
def analyze_todos(todos):
pending = [t for t in todos if not t["done"]]
done = [t for t in todos if t["done"]]
if not todos:
return {"total": 0, "message": "目前沒有任何待辦事項。"}
by_priority = Counter(t["priority"] for t in pending)
overdue = [t for t in pending if t.get("days_left", 99) < 0]
today = [t for t in pending if t.get("days_left") == 0]
return {
"total": len(todos),
"pending": len(pending),
"done": len(done),
"completion_rate": len(done) / len(todos),
"high_priority": by_priority["高"],
"overdue": overdue,
"due_today": today,
"next_action": min(
pending,
key=lambda t: (PRIORITY_ORDER[t["priority"]], t.get("days_left", 99)),
default=None,
),
}
def make_todo_summary(insight):
if insight["total"] == 0:
return insight["message"]
lines = [
f"待辦:{insight['pending']} 筆未完成 / 共 {insight['total']} 筆"
f"(完成率 {insight['completion_rate']:.0%})"
]
if insight["overdue"]:
titles = "、".join(t["title"] for t in insight["overdue"])
lines.append(f"🔴 已逾期 {len(insight['overdue'])} 筆:{titles}")
if insight["due_today"]:
titles = "、".join(t["title"] for t in insight["due_today"])
lines.append(f"📌 今天到期:{titles}")
if insight["next_action"]:
lines.append(f"👉 建議先處理:{insight['next_action']['title']}")
return "\n".join(lines)
{insight['completion_rate']:.0%} 這個格式化語法很好用——把 0.25 直接顯示成 25%。
這是我覺得今天最值得講的部分。
一個常見的誤解是:「反正 AI 很聰明,我把原始資料丟給它,它自己會分析啊。」
某種程度上是對的。但有三個理由讓我們還是要自己做分析:
原始的天氣 JSON 可能有兩千個 token。分析後的三句話只要五十個 token。如果 Agent 要同時看天氣、行事曆、待辦清單、交通狀況——不做壓縮的話 context 很快就爆了,而且每次呼叫都在燒錢。
分析的本質是壓縮:把資訊量保留下來,把資料量降下去。
「降雨機率 60% 算不算高」這種判斷,如果交給 AI,今天可能說「要帶傘」,明天可能說「應該還好」。但如果我們在程式裡寫死 will_rain = max_rain >= 40,行為就是可預測的。
Agent 系統裡有些判斷應該交給 AI(模糊的、需要理解語意的),有些應該用程式(明確的、有標準的)。**閾值判斷屬於後者。**這個區分在 Day 25 會再深入討論。
def test_will_rain():
insight = analyze_weather(fake_data_with_60_percent_rain)
assert insight["will_rain"] is True
程式邏輯可以寫測試,AI 的判斷很難。把能確定的部分抽出來用程式做,Agent 的可靠度就上升了。
| 工作 | 誰來做 |
|---|---|
| 取得原始資料 | 程式(Day 13–15 的工具) |
| 計算、統計、閾值判斷 | 程式(今天) |
| 理解使用者意圖、組合多種資訊、產生自然的回應 | AI(Day 18 之後) |
| 決定要用哪些工具、下一步做什麼 | AI(Day 21 之後) |
**AI 負責它擅長的模糊判斷,程式負責它擅長的精確計算。**這不是誰取代誰,是分工。
sorted(data, key=lambda x: ...) 是處理 dict list 的基本功key 回傳 tuple 就能做多重排序,Python 會逐項比對any / all,找第一個用 next
Counter 數次數、defaultdict(list) 分組,比手寫迴圈乾淨max(data, key=...) 找的是「最…的那一筆資料」,不只是數字明天講 logging——讓 Agent 的每一步都留下紀錄。當它做了一個你看不懂的決定時,你要有辦法查。