iT邦幫忙

2026 iThome 鐵人賽

DAY 16
0
AI Engineering

從 Prompt 到自主決策:用 Python × Agentic Workflow 實作生活助理系列 第 16 篇

從一堆資料到一句結論:排序、篩選與統計

  • 分享至 

  • xImage
  •  

前面幾天我們拿到了不少資料:天氣預報、待辦清單。但拿到資料不等於知道答案。

想想這兩句話的差別:

A:「明天 06:00–18:00 多雲,25–30°C,降雨機率 60%;18:00–06:00 陰,23–26°C,降雨機率 30%。」

B:「明天白天有六成機率下雨,記得帶傘。」

A 是資料,B 是結論。使用者要的是 B。

而要從 A 變成 B,中間需要的就是今天的主題:整理、篩選、統計。


一、排序:sorted 與 key

Python 的 sorted() 回傳一個新的排序過的 list(list.sort() 則是就地排序,會改掉原本的)。

numbers = [3, 1, 4, 1, 5, 9, 2, 6]

print(sorted(numbers))              # [1, 1, 2, 3, 4, 5, 6, 9]
print(sorted(numbers, reverse=True))# [9, 6, 5, 4, 3, 2, 1, 1]
print(numbers)                      # [3, 1, 4, 1, 5, 9, 2, 6] ← 原本的沒變

但我們的資料通常是 dict 組成的 list:

todos = [
    {"title": "買牛奶", "priority": "中", "done": False, "days_left": 3},
    {"title": "繳電費", "priority": "高", "done": False, "days_left": 1},
    {"title": "整理房間", "priority": "低", "done": True, "days_left": 7},
    {"title": "寫文章", "priority": "高", "done": False, "days_left": 0},
]

sorted(todos)   # ❌ TypeError: '<' not supported between instances of 'dict'

dict 之間沒辦法直接比大小,所以要用 key 參數告訴 Python「依據什麼排序」:

# 依剩餘天數排序
by_deadline = sorted(todos, key=lambda t: t["days_left"])
for t in by_deadline:
    print(f"{t['days_left']} 天:{t['title']}")

輸出:

0 天:寫文章
1 天:繳電費
3 天:買牛奶
7 天:整理房間

lambda 是什麼

lambda t: t["days_left"] 是一個「匿名函式」,等同於:

def get_days_left(t):
    return t["days_left"]

sorted(todos, key=get_days_left)

只是懶得另外定義一個函式,就直接寫在參數裡。lambda 參數: 回傳值,就這樣。

自訂排序順序

優先度是「高/中/低」,直接排會變成按筆劃或 Unicode 順序,不是我們要的。解法是先對應成數字:

PRIORITY_ORDER = {"高": 0, "中": 1, "低": 2}

by_priority = sorted(todos, key=lambda t: PRIORITY_ORDER[t["priority"]])

多重排序

如果要「先照優先度,同優先度再照剩餘天數」,key 回傳一個 tuple:

smart_order = sorted(
    todos,
    key=lambda t: (
        t["done"],                        # 未完成(False=0)排前面
        PRIORITY_ORDER[t["priority"]],    # 再照優先度
        t["days_left"],                   # 再照剩餘天數
    ),
)

for t in smart_order:
    mark = "✓" if t["done"] else "□"
    print(f"{mark} [{t['priority']}] {t['title']}(剩 {t['days_left']} 天)")

輸出:

□ [高] 寫文章(剩 0 天)
□ [高] 繳電費(剩 1 天)
□ [中] 買牛奶(剩 3 天)
✓ [低] 整理房間(剩 7 天)

Python 比較 tuple 時是逐項比對:先比第一項,相同才比第二項,以此類推。這招讓多重排序變得非常簡潔。

(順帶一提,False < True,因為布林值本質上就是 0 和 1。)


二、篩選:filter 與綜合表達式

Day 6 學過的 list comprehension 就是最好用的篩選工具:

# 未完成的
pending = [t for t in todos if not t["done"]]

# 高優先度且未完成的
urgent = [t for t in todos if t["priority"] == "高" and not t["done"]]

# 三天內到期的
soon = [t for t in todos if t["days_left"] <= 3]

比 filter(lambda t: ..., todos) 好讀太多,Python 社群也偏好這種寫法。

any 與 all

有時候我們只想知道「有沒有」、「是不是全部」:

has_urgent = any(t["priority"] == "高" and not t["done"] for t in todos)
all_done = all(t["done"] for t in todos)

print(f"有緊急事項嗎:{has_urgent}")   # True
print(f"全部完成了嗎:{all_done}")     # False

any() 只要有一個 True 就回傳 True,all() 要全部都 True。

而且它們有短路求值——any() 找到第一個 True 就停了,不會跑完整個清單。資料量大時差很多。

next:找第一個符合的

first_urgent = next(
    (t for t in todos if t["priority"] == "高" and not t["done"]),
    None,      # 找不到時的預設值
)

if first_urgent:
    print(f"最該先處理:{first_urgent['title']}")

記得加第二個參數當預設值,不然找不到時會拋 StopIteration。


三、統計:從數字到洞察

內建函式

temps = [26, 28, 31, 29, 25, 27, 30]

print(f"最低:{min(temps)}°C")
print(f"最高:{max(temps)}°C")
print(f"平均:{sum(temps) / len(temps):.1f}°C")
print(f"總筆數:{len(temps)}")

⚠️ sum(temps) / len(temps) 在空清單時會 ZeroDivisionError。實務上要先檢查:

average = sum(temps) / len(temps) if temps else 0

statistics 模組

Python 內建 statistics,不用裝 numpy:

import statistics

temps = [26, 28, 31, 29, 25, 27, 30]

print(statistics.mean(temps))     # 28 平均數
print(statistics.median(temps))   # 28 中位數
print(statistics.stdev(temps))    # 2.16 標準差

median 在有極端值時比 mean 更能代表「典型情況」。

max / min 搭配 key

找「最…的那一筆資料」,而不只是最大的數字:

most_urgent = min(todos, key=lambda t: t["days_left"])
print(f"最急的是:{most_urgent['title']}")   # 寫文章

Counter:算次數

from collections import Counter

priorities = [t["priority"] for t in todos]
counts = Counter(priorities)

print(counts)                    # Counter({'高': 2, '中': 1, '低': 1})
print(counts["高"])              # 2
print(counts.most_common(1))     # [('高', 2)]

Counter 是 dict 的子類別,專門用來數東西。比自己寫迴圈 counts[x] = counts.get(x, 0) + 1 乾淨多了。

groupby:分組

如果要「照某個欄位分組」,我自己習慣用 defaultdict:

from collections import defaultdict

by_priority = defaultdict(list)
for t in todos:
    by_priority[t["priority"]].append(t["title"])

for priority in ("高", "中", "低"):
    items = by_priority[priority]
    if items:
        print(f"{priority}:{', '.join(items)}")

輸出:

高:繳電費, 寫文章
中:買牛奶
低:整理房間

defaultdict(list) 的意思是「取一個不存在的 key 時,自動給我一個空 list」,省掉了 if key not in d: d[key] = [] 這行。


四、實戰:把天氣資料變成一句建議

現在把今天學的東西,套用到 Day 14 拿到的天氣資料上:

"""analysis/weather_insight.py"""


def analyze_weather(data):
    """把天氣預報分析成結構化的洞察。"""
    periods = data["periods"]
    if not periods:
        return None

    # 統計
    rain_chances = [p["rain_probability"] for p in periods
                    if p["rain_probability"] is not None]
    min_temps = [p["min_temp"] for p in periods if p["min_temp"] is not None]
    max_temps = [p["max_temp"] for p in periods if p["max_temp"] is not None]

    max_rain = max(rain_chances, default=0)
    lowest = min(min_temps, default=None)
    highest = max(max_temps, default=None)

    # 找出最可能下雨的時段
    rainiest = max(periods, key=lambda p: p["rain_probability"] or 0)

    # 溫差
    temp_range = (highest - lowest) if (highest and lowest) else 0

    return {
        "city": data["city"],
        "max_rain": max_rain,
        "lowest": lowest,
        "highest": highest,
        "temp_range": temp_range,
        "rainiest_period": rainiest,
        "will_rain": max_rain >= 40,
        "big_temp_swing": temp_range >= 8,
        "is_cold": lowest is not None and lowest <= 18,
        "is_hot": highest is not None and highest >= 32,
    }


def make_advice(insight):
    """把洞察轉成給人看的建議。"""
    if insight is None:
        return "目前沒有可用的天氣資料。"

    advice = [f"【{insight['city']}】"]

    # 主要結論
    if insight["will_rain"]:
        period = insight["rainiest_period"]
        when = period["start"][5:16]
        advice.append(
            f"⚠️ {when} 前後降雨機率達 {period['rain_probability']}%,記得帶傘。"
        )
    else:
        advice.append(f"☀️ 降雨機率最高只有 {insight['max_rain']}%,應該不用帶傘。")

    # 溫度
    advice.append(f"氣溫 {insight['lowest']}–{insight['highest']}°C。")

    if insight["big_temp_swing"]:
        advice.append(f"日夜溫差達 {insight['temp_range']} 度,建議洋蔥式穿搭。")
    if insight["is_cold"]:
        advice.append("早晚偏涼,記得加件外套。")
    if insight["is_hot"]:
        advice.append("白天偏熱,注意補充水分、避免長時間曝曬。")

    return "\n".join(advice)

試跑:

import config
from tools.weather import get_weather

data = get_weather("臺北市", config.CWA_API_KEY)
insight = analyze_weather(data)
print(make_advice(insight))

輸出:

【臺北市】
⚠️ 09-30 06:00 前後降雨機率達 60%,記得帶傘。
氣溫 23–31°C。
日夜溫差達 8 度,建議洋蔥式穿搭。

**從一堆數字,變成三句可以行動的建議。**這就是分析的價值。


五、再做一個:待辦清單的分析

"""analysis/todo_insight.py"""

from collections import Counter, defaultdict

PRIORITY_ORDER = {"高": 0, "中": 1, "低": 2}


def analyze_todos(todos):
    pending = [t for t in todos if not t["done"]]
    done = [t for t in todos if t["done"]]

    if not todos:
        return {"total": 0, "message": "目前沒有任何待辦事項。"}

    by_priority = Counter(t["priority"] for t in pending)
    overdue = [t for t in pending if t.get("days_left", 99) < 0]
    today = [t for t in pending if t.get("days_left") == 0]

    return {
        "total": len(todos),
        "pending": len(pending),
        "done": len(done),
        "completion_rate": len(done) / len(todos),
        "high_priority": by_priority["高"],
        "overdue": overdue,
        "due_today": today,
        "next_action": min(
            pending,
            key=lambda t: (PRIORITY_ORDER[t["priority"]], t.get("days_left", 99)),
            default=None,
        ),
    }


def make_todo_summary(insight):
    if insight["total"] == 0:
        return insight["message"]

    lines = [
        f"待辦:{insight['pending']} 筆未完成 / 共 {insight['total']} 筆"
        f"(完成率 {insight['completion_rate']:.0%})"
    ]

    if insight["overdue"]:
        titles = "、".join(t["title"] for t in insight["overdue"])
        lines.append(f"🔴 已逾期 {len(insight['overdue'])} 筆:{titles}")

    if insight["due_today"]:
        titles = "、".join(t["title"] for t in insight["due_today"])
        lines.append(f"📌 今天到期:{titles}")

    if insight["next_action"]:
        lines.append(f"👉 建議先處理:{insight['next_action']['title']}")

    return "\n".join(lines)

{insight['completion_rate']:.0%} 這個格式化語法很好用——把 0.25 直接顯示成 25%。


六、為什麼 Agent 需要這一層

這是我覺得今天最值得講的部分。

一個常見的誤解是:「反正 AI 很聰明,我把原始資料丟給它,它自己會分析啊。」

某種程度上是對的。但有三個理由讓我們還是要自己做分析:

1. Token 成本

原始的天氣 JSON 可能有兩千個 token。分析後的三句話只要五十個 token。如果 Agent 要同時看天氣、行事曆、待辦清單、交通狀況——不做壓縮的話 context 很快就爆了,而且每次呼叫都在燒錢。

分析的本質是壓縮:把資訊量保留下來,把資料量降下去。

2. 確定性

「降雨機率 60% 算不算高」這種判斷,如果交給 AI,今天可能說「要帶傘」,明天可能說「應該還好」。但如果我們在程式裡寫死 will_rain = max_rain >= 40,行為就是可預測的。

Agent 系統裡有些判斷應該交給 AI(模糊的、需要理解語意的),有些應該用程式(明確的、有標準的)。**閾值判斷屬於後者。**這個區分在 Day 25 會再深入討論。

3. 可測試

def test_will_rain():
    insight = analyze_weather(fake_data_with_60_percent_rain)
    assert insight["will_rain"] is True

程式邏輯可以寫測試,AI 的判斷很難。把能確定的部分抽出來用程式做,Agent 的可靠度就上升了。

所以分工是這樣的

工作 誰來做
取得原始資料 程式(Day 13–15 的工具)
計算、統計、閾值判斷 程式(今天)
理解使用者意圖、組合多種資訊、產生自然的回應 AI(Day 18 之後)
決定要用哪些工具、下一步做什麼 AI(Day 21 之後)

**AI 負責它擅長的模糊判斷,程式負責它擅長的精確計算。**這不是誰取代誰,是分工。


小結

  • sorted(data, key=lambda x: ...) 是處理 dict list 的基本功
  • key 回傳 tuple 就能做多重排序,Python 會逐項比對
  • 篩選用 list comprehension,判斷存在用 any / all,找第一個用 next
  • Counter 數次數、defaultdict(list) 分組,比手寫迴圈乾淨
  • max(data, key=...) 找的是「最…的那一筆資料」,不只是數字
  • 分析的本質是壓縮:把大量資料變成少量結論
  • Agent 需要這一層,因為它省 token、行為可預測、而且可以寫測試

明天講 logging——讓 Agent 的每一步都留下紀錄。當它做了一個你看不懂的決定時,你要有辦法查。


上一篇
什麼才算一個「工具」?Tool 的介面設計
下一篇
讓程式記住它做過什麼:logging 與 Agent 的可觀測性
系列文
從 Prompt 到自主決策:用 Python × Agentic Workflow 實作生活助理 共 18 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言