前面十七天,我們蓋好了地基:
今天 AI 終於要進場了。
先把神祕感拿掉。從寫程式的角度,大型語言模型(LLM)就是一個函式:
回應文字 = 模型(對話訊息串)
輸入是一串對話,輸出是下一段文字。就這樣。
它不是資料庫(不會精確記住事實),不是搜尋引擎(不會即時查網路),也不是計算機(算數會錯)。它是一個根據上下文預測合理接續文字的函式。
這個認知很重要。後面我們之所以要給它「工具」,就是因為這個函式本身只會產生文字——要讓它真的做事,得由我們來接手。
這個系列我會用 Anthropic 的 Claude API 當例子。概念在各家服務上大同小異,換成別家也只是改幾行程式碼。
pip install anthropic
照 Day 12 的做法,寫進 .env:
ANTHROPIC_API_KEY=sk-ant-api03-xxxxxxxxxxxx
到 https://console.anthropic.com 申請。
import anthropic
import config # Day 12 寫的設定模組
client = anthropic.Anthropic(api_key=config.ANTHROPIC_API_KEY)
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=[
{"role": "user", "content": "用一句話解釋什麼是 Python 的 list comprehension"}
],
)
print(response.content[0].text)
輸出大概會像:
List comprehension 是一種用一行程式碼從現有序列建立新 list 的簡潔語法,
例如 [x*2 for x in numbers] 會把每個元素乘以二。
**成功了。**AI 現在是我們程式的一部分了。
💡
client = anthropic.Anthropic()不帶參數也可以——SDK 會自己去環境變數找ANTHROPIC_API_KEY。但我習慣明確從config傳進去,這樣哪裡出問題比較好查。
response 不只有文字:
print(response.id) # msg_01XyZ...
print(response.model) # claude-opus-5
print(response.stop_reason) # end_turn
print(response.usage) # Usage(input_tokens=25, output_tokens=58)
print(response.content) # [TextBlock(text='...', type='text')]
幾個要記住的:
content 是一個 list不是字串!因為一次回應可能包含多個區塊(文字、工具呼叫、思考過程)。所以取文字要:
text = response.content[0].text
更保險的寫法(Day 21 開始會用到工具,content 裡就不只文字了):
def extract_text(response):
"""從回應中取出所有文字區塊並串接。"""
parts = [block.text for block in response.content if block.type == "text"]
return "\n".join(parts)
stop_reason:為什麼停下來| 值 | 意思 | 要怎麼處理 |
|---|---|---|
end_turn |
正常講完了 | 直接用 |
max_tokens |
被截斷了 | 提高 max_tokens 或縮短要求 |
tool_use |
它想呼叫工具 | Day 21 的主題 |
refusal |
它拒絕回答 | 檢查請求內容 |
**一定要檢查 stop_reason。**最常見的意外就是 max_tokens 把回答砍掉一半,而你以為模型就是這樣回答的。
usage:這次花了多少print(f"輸入 {response.usage.input_tokens} tokens")
print(f"輸出 {response.usage.output_tokens} tokens")
**Token 是計價單位。**輸入和輸出的單價不同(輸出通常貴 5 倍左右)。
粗略換算:英文大約 1 token ≈ 4 個字元,中文大約 1 個字 ≈ 1~2 tokens。
這個數字要一直放在心上。Day 17 的 trace 我把 usage 記下來,就是為了事後能分析成本。
messages 是一個 list,每個元素有 role 和 content:
messages = [
{"role": "user", "content": "我叫 pzhiqi"},
{"role": "assistant", "content": "你好 pzhiqi!有什麼可以幫你的嗎?"},
{"role": "user", "content": "我剛剛說我叫什麼?"},
]
role 只有兩種:user(使用者)和 assistant(AI)。而且必須交替出現,從 user 開始。
這點非常關鍵,新手常常搞錯:
**模型完全不記得上一次的對話。**每次呼叫都是全新的。
它之所以「記得」你叫 pzhiqi,只是因為你把前面的對話又傳了一次。
所以多輪對話要自己維護歷史:
messages = []
def chat(user_input):
messages.append({"role": "user", "content": user_input})
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
messages=messages,
)
reply = extract_text(response)
messages.append({"role": "assistant", "content": reply}) # ← 別忘了這行
return reply
print(chat("我叫 pzhiqi"))
print(chat("我剛剛說我叫什麼?")) # 它會回答 pzhiqi
**忘記把 AI 的回覆加回 messages,是最常見的 bug。**症狀是:AI 完全不記得自己剛剛說過什麼,一直重複自我介紹。
每一輪都要把整串歷史重新傳一次。第 10 輪對話的輸入 token 是第 1 輪的好幾倍。
這就是為什麼 Day 26 要專門講記憶管理——不能無限制地累積。
system 是一個獨立的參數,用來設定 AI 的身分、風格和規則:
SYSTEM_PROMPT = """你是一個台灣使用者的生活助理,名字叫「小幫」。
你的任務是協助使用者管理日常行程、待辦事項與天氣資訊。
回答時請遵守:
- 使用繁體中文與台灣用語
- 簡潔具體,不要長篇大論
- 給建議時說明理由
- 不確定的事情直接說不知道,不要猜測
"""
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
system=SYSTEM_PROMPT, # ← 獨立參數,不在 messages 裡
messages=[{"role": "user", "content": "今天適合出門嗎?"}],
)
注意 system 不放在 messages 裡面,它是 create() 的獨立參數。
system prompt 的內容在每次呼叫時都會被送出,所以:
這個區分還有一個好處:穩定的前綴可以被快取,降低成本。這算進階議題,但原則先記著。
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024, # 最多生成多少 token
temperature=1.0, # 隨機程度
system=SYSTEM_PROMPT,
messages=messages,
)
max_tokens(必填)
輸出的上限。設太小會被截斷(stop_reason 會是 max_tokens)。一般對話設 1024~4096 就夠。
temperature(0~1)
控制隨機性。低的話回答比較穩定、可預測;高的話比較有變化。
實務建議:
| 用途 | temperature |
|---|---|
| 抽取結構化資料、分類 | 0 |
| 一般對話、助理 | 0.7~1.0 |
| 創意寫作 | 1.0 |
⚠️ 但要注意:新一代的推理模型(thinking models)不支援 temperature,傳了會直接報錯。像 Claude Opus 5 就屬於這類。所以上面那個範例其實會出錯——正確寫法是不要傳 temperature:
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
system=SYSTEM_PROMPT,
messages=messages,
)
這是很常見的坑:照著舊教學寫,結果收到 400 錯誤。看官方文件比看舊部落格可靠。
把 Day 8(例外)、Day 13(重試)、Day 17(logging)學的東西全用上:
"""llm_client.py — 包裝 LLM 呼叫。"""
import logging
import time
import anthropic
import config
logger = logging.getLogger(__name__)
MODEL = "claude-opus-5"
MAX_TOKENS = 4096
class LLMError(Exception):
"""呼叫模型失敗。"""
class LLMClient:
def __init__(self, model=MODEL, system=None):
self.client = anthropic.Anthropic(api_key=config.ANTHROPIC_API_KEY)
self.model = model
self.system = system
self.total_input_tokens = 0
self.total_output_tokens = 0
def call(self, messages, tools=None, max_tokens=MAX_TOKENS, retries=3):
"""呼叫模型,自動重試可恢復的錯誤。"""
kwargs = {
"model": self.model,
"max_tokens": max_tokens,
"messages": messages,
}
if self.system:
kwargs["system"] = self.system
if tools:
kwargs["tools"] = tools
last_error = None
for attempt in range(retries):
try:
start = time.time()
response = self.client.messages.create(**kwargs)
elapsed = time.time() - start
self.total_input_tokens += response.usage.input_tokens
self.total_output_tokens += response.usage.output_tokens
logger.info(
"模型回應:stop=%s,in=%d out=%d tokens,耗時 %.2fs",
response.stop_reason,
response.usage.input_tokens,
response.usage.output_tokens,
elapsed,
)
if response.stop_reason == "max_tokens":
logger.warning("回應被 max_tokens 截斷了,考慮調高上限")
return response
except anthropic.RateLimitError as e:
# 請求太頻繁,等久一點
last_error = e
wait = 2 ** (attempt + 2) # 4, 8, 16 秒
logger.warning("觸發速率限制,%d 秒後重試", wait)
time.sleep(wait)
except anthropic.APIStatusError as e:
if e.status_code >= 500:
last_error = e
wait = 2 ** attempt
logger.warning("伺服器錯誤 %d,%d 秒後重試", e.status_code, wait)
time.sleep(wait)
else:
# 4xx 是我們的問題,重試沒用
logger.error("請求錯誤 %d:%s", e.status_code, e)
raise LLMError(f"請求錯誤({e.status_code}):{e}") from e
except anthropic.APIConnectionError as e:
last_error = e
wait = 2 ** attempt
logger.warning("連線失敗,%d 秒後重試", wait)
time.sleep(wait)
raise LLMError(f"重試 {retries} 次後仍失敗:{last_error}")
@staticmethod
def extract_text(response):
"""取出回應中的所有文字。"""
parts = [b.text for b in response.content if b.type == "text"]
return "\n".join(parts).strip()
def report_usage(self):
return (
f"累計用量:輸入 {self.total_input_tokens:,} tokens,"
f"輸出 {self.total_output_tokens:,} tokens"
)
用起來:
from logger_setup import setup_logging
from llm_client import LLMClient, LLMError
setup_logging()
llm = LLMClient(system="你是一個簡潔的台灣生活助理。")
messages = [{"role": "user", "content": "明天台北 25 度會下雨,我該穿什麼?"}]
try:
response = llm.call(messages)
print(LLMClient.extract_text(response))
except LLMError as e:
print(f"呼叫失敗:{e}")
print(llm.report_usage())
這個 LLMClient 會一路用到 Day 30,之後只會往上加功能。
把今天學的東西組成一個真的能聊天的程式:
"""chat.py — 簡單的多輪對話。"""
from logger_setup import setup_logging
from llm_client import LLMClient, LLMError
SYSTEM_PROMPT = """你是一個台灣使用者的生活助理,名字叫「小幫」。
使用繁體中文與台灣用語,回答簡潔具體。
你目前還沒有查詢天氣或行事曆的能力,
如果使用者問到這類問題,請誠實說明你需要這些資訊才能回答。"""
MAX_HISTORY = 20 # 最多保留幾則訊息
def main():
setup_logging()
llm = LLMClient(system=SYSTEM_PROMPT)
messages = []
print("=== 小幫上線(輸入 quit 離開)===\n")
while True:
user_input = input("你:").strip()
if user_input.lower() in ("quit", "exit", "掰掰"):
break
if not user_input:
continue
messages.append({"role": "user", "content": user_input})
try:
response = llm.call(messages)
except LLMError as e:
print(f"小幫:抱歉,我現在有點問題({e})\n")
messages.pop() # 把失敗的這則拿掉,不要污染歷史
continue
reply = LLMClient.extract_text(response)
messages.append({"role": "assistant", "content": reply})
print(f"小幫:{reply}\n")
# 控制歷史長度,避免越來越貴
if len(messages) > MAX_HISTORY:
messages = messages[-MAX_HISTORY:]
# 確保第一則是 user(API 要求)
while messages and messages[0]["role"] != "user":
messages.pop(0)
print(f"\n{llm.report_usage()}")
if __name__ == "__main__":
main()
跑跑看:
=== 小幫上線(輸入 quit 離開)===
你:我叫 pzhiqi,正在參加鐵人賽
小幫:你好 pzhiqi!鐵人賽 30 天不簡單,加油。有什麼需要幫忙的嗎?
你:我剛剛說我在幹嘛?
小幫:你說你正在參加鐵人賽。
你:明天台北天氣如何?
小幫:抱歉,我目前還沒有查詢天氣的能力。如果你告訴我天氣資訊,
我可以幫你判斷要不要帶傘或穿什麼。
你:quit
累計用量:輸入 1,847 tokens,輸出 236 tokens
注意兩件事:
messages 裡第二點正是今天結束時的狀態:AI 有腦,但沒有手。
回頭看整個系列的地圖:
Day 2–7 Python 基礎 ✅
Day 8–11 例外 / 檔案 / JSON / Class ✅
Day 12–15 設定 / HTTP / API / Tool ✅ ← 手(能查天氣、能記待辦)
Day 16–17 分析 / 日誌 ✅
Day 18 第一次呼叫 LLM ✅ ← 腦(能理解、能對話)
Day 19–20 Prompt / 結構化輸出 ← 讓腦更可靠
Day 21–22 Tool Use ← 把腦跟手接起來 ⭐
Day 23–26 Workflow / Agent / 記憶
Day 27–30 整合專案
**Day 21 是整個系列的轉捩點。**在那之前,工具是我們呼叫的;在那之後,是 AI 自己決定要呼叫什麼。
但在那之前,我們得先確保跟 AI 的溝通是可靠的——這就是接下來兩天的主題。
response.content 是一個 list,要用 content[0].text 或過濾 type == "text" 取文字stop_reason,max_tokens 代表回答被截斷了usage 記錄 token 用量,這是計價單位,要持續關注messages 是最常見的 bugsystem 是獨立參數,放穩定的角色設定;動態資料放 messagestemperature,傳了會報錯明天講 Prompt 設計——同一件事,換個講法,結果可能天差地遠。