昨天我們讓 Agent 成功呼叫 LLM 分析規格書。但試著丟一份 100 頁的規格書進去,你就會碰到這個:
Mistral 7B 的上下文窗口是 32k token。一份 100 頁 SRS?大約 50k+ token。LLM 會截斷,或者到中間就開始胡說八道。
簡單的解法是按字符數切:「每 5000 字符一個 chunk」。這樣能進 LLM 的上下文。但問題是,規格書不是雜文。切到一半的時候,第一個 chunk 結尾是「系統支持多用戶並行」,第二個 chunk 開頭是「...編輯、分享筆記」。LLM 看不到這兩句是同一個需求,而且它忘記了第一個 chunk 說的內容。
如果規格書的邏輯矛盾跨越了 chunk 邊界呢?LLM 看不到。
所以需要按 Markdown 層級分塊,不是按字符。
我們用 # 級標題作為主要章節邊界。每個 # 開始一個新 chunk,## 和 ### 跟著父 chunk。這樣:
# 第 1 章:系統概述
## 1.1 功能範圍
### 1.1.1 筆記編輯器
## 1.2 用戶角色
# 第 2 章:功能需求
## 2.1 筆記管理
- REQ-2.1.1:系統支持多用戶
切割成:
LLM 看到 REQ-2.1.1 時,知道它在「第 2 章 - 功能需求 - 筆記管理」的上下文裡。
還要防止單個 chunk 超限。我們估算 token:1 token ≈ 4 字符。簡單,夠用。
src/chunking.py:
import re
from dataclasses import dataclass
from typing import List, Optional
@dataclass
class Chunk:
"""文檔分塊"""
content: str
title: str
level: int # 標題層級(1=#, 2=##)
start_line: int
end_line: int
token_estimate: int
def __str__(self) -> str:
indent = " " * (self.level - 1)
return (
f"{indent}{'#' * self.level} {self.title}\n"
f"{indent} (行 {self.start_line}-{self.end_line}, "
f"~{self.token_estimate} tokens)"
)
class MarkdownChunker:
"""Markdown 感知的分塊器"""
DEFAULT_TOKEN_LIMIT = 4000
CHARS_PER_TOKEN = 4 # 1 token ≈ 4 字符
def __init__(self, token_limit: int = DEFAULT_TOKEN_LIMIT):
self.token_limit = token_limit
def _estimate_tokens(self, text: str) -> int:
"""估算 token 數"""
return max(1, len(text) // self.CHARS_PER_TOKEN)
def _get_heading_level(self, line: str) -> Optional[int]:
"""判斷是否是標題,返回層級或 None"""
match = re.match(r'^(#+)\s', line)
return len(match.group(1)) if match else None
def chunk(self, text: str) -> List[Chunk]:
"""分塊"""
lines = text.split('\n')
chunks = []
current_chunk_lines = []
current_title = "未分類"
current_level = 0
current_start_line = 0
for line_num, line in enumerate(lines):
heading_level = self._get_heading_level(line)
# 遇到 # 級標題,就開始新 chunk
if heading_level == 1 and current_chunk_lines:
chunk_text = '\n'.join(current_chunk_lines)
chunks.append(Chunk(
content=chunk_text,
title=current_title,
level=current_level,
start_line=current_start_line,
end_line=line_num - 1,
token_estimate=self._estimate_tokens(chunk_text)
))
current_chunk_lines = []
current_start_line = line_num
# 更新當前標題
if heading_level:
current_title = re.sub(r'^#+\s+', '', line).strip()
current_level = heading_level
current_chunk_lines.append(line)
# 最後一個 chunk
if current_chunk_lines:
chunk_text = '\n'.join(current_chunk_lines)
chunks.append(Chunk(
content=chunk_text,
title=current_title,
level=current_level,
start_line=current_start_line,
end_line=len(lines) - 1,
token_estimate=self._estimate_tokens(chunk_text)
))
return chunks
def validate_chunks(self, chunks: List[Chunk]) -> List[str]:
"""檢查超限"""
warnings = []
for i, chunk in enumerate(chunks):
if chunk.token_estimate > self.token_limit:
warnings.append(
f"⚠️ Chunk {i+1} ('{chunk.title}') "
f"超過限制:{chunk.token_estimate} > {self.token_limit} tokens"
)
return warnings
def print_chunks(chunks: List[Chunk], show_content: bool = False):
"""顯示分塊結果"""
print(f"\n{'='*70}")
print(f"文檔分塊結果(共 {len(chunks)} 個 chunks)")
print(f"{'='*70}\n")
total_tokens = 0
for i, chunk in enumerate(chunks, 1):
print(f"{i}. {chunk}")
total_tokens += chunk.token_estimate
if show_content:
preview = chunk.content[:200].replace('\n', '\n ')
print(f" 預覽:{preview}...\n")
else:
print()
print(f"{'='*70}")
print(f"總計:{total_tokens} tokens")
print(f"{'='*70}\n")
在 src/main.py 加上這個:
from src.chunking import MarkdownChunker, print_chunks
# OllamaAgent 加新方法
def chunk_srs(self, srs_text: str, token_limit: int = 4000):
"""分塊 SRS"""
chunker = MarkdownChunker(token_limit=token_limit)
chunks = chunker.chunk(srs_text)
warnings = chunker.validate_chunks(chunks)
if warnings:
print("\n⚠️ 警告:")
for warning in warnings:
print(f" {warning}")
return chunks
# 新的 CLI 命令
@app.command()
def chunk(
filepath: str = typer.Argument(..., help="SRS 檔案路徑"),
token_limit: int = typer.Option(
4000,
"--limit",
"-l",
help="單個 chunk 的 token 限制"),
show_content: bool = typer.Option(
False,
"--content",
"-c",
help="是否顯示內容預覽")):
"""分塊 SRS 文檔
使用方式:
python -m src.main chunk tests/fixtures/srs_small.md
python -m src.main chunk tests/fixtures/srs_small.md -c
python -m src.main chunk tests/fixtures/srs_small.md -l 3000
"""
try:
agent = OllamaAgent()
print(f"📖 讀取: {filepath}")
srs_text = agent.read_srs(filepath)
print(f"✓ {len(srs_text)} 字符")
print(f"\n🔨 分塊中(限制:{token_limit} tokens/chunk)...")
chunks = agent.chunk_srs(srs_text, token_limit=token_limit)
print_chunks(chunks, show_content=show_content)
except Exception as e:
print(f"❌ 錯誤: {e}")
raise typer.Exit(1)
# 基本分塊
python -m src.main chunk tests/fixtures/srs_small.md
# 顯示預覽
python -m src.main chunk tests/fixtures/srs_small.md -c
# 降低 token 限制,看看會不會分得更多
python -m src.main chunk tests/fixtures/srs_small.md -l 2000
預期:
📖 讀取: tests/fixtures/srs_small.md
✓ 847 字符
🔨 分塊中(限制:4000 tokens/chunk)...
======================================================================
文檔分塊結果(共 4 個 chunks)
======================================================================
1. # 線上筆記系統 - 需求規格書
(行 0-8, ~212 tokens)
2. ## 功能需求
(行 9-16, ~280 tokens)
3. ## 性能需求
(行 17-20, ~150 tokens)
4. ## 安全需求
(行 21-27, ~120 tokens)
======================================================================
總計:762 tokens
======================================================================
按 # 級邊界,不按 ##:因為 # 是主章節,每個主章節代表一個完整的邏輯單元。如果按 ## 分,chunks 會很碎。
Token 估算用簡單公式:中文、英文都大致是 1 token ≈ 4 字符。如果想精確,可以用 tiktoken,但這裡簡單估算夠了。
明天(Day 4)開始,我們加入「全局約束追蹤」。現在 chunks 只是切割,它們還是獨立的。明天會讓 Agent 在分析每個 chunk 時,記住所有之前 chunks 裡提出過的約束,這樣才能檢測到「第 2 章說多用戶,第 8 章說單用戶」的矛盾。
第 1 週:透明的 SRS 分析