iT邦幫忙

2026 iThome 鐵人賽

DAY 18
0
AI Engineering

30 天打造 Codebase Intelligence Agent:從程式碼檢索、結構化索引到變更影響分析實戰系列 第 18 篇

Day 18:讓本地 LLM 幫候選結果重新排序:LLM Reranker 與防呆降級機制

  • 分享至 

  • xImage
  •  

昨天我們透過自動化 Eval Runner 拿到了第一份客觀成績單。

數據清晰指出了系統的短板:在純粹由自然語言提問的 semantic_intent 分類中,MRR@5 只有 0.50。雖然正確的檔案有被抓進 Top 5 候選集裡(Hit@5 = 1.00),但排在第 1 名的常常是無關的雜訊片段或粗粒度的檔案宣告。

如果只靠 Day 14 的規則式重排(Deterministic Reranker),我們只能比對字串和語法節點類型,無法真正理解「這個自然語言提問」與「這段 Python 邏輯」在語意深度上的契合度。

今天我們引入進階的 LLM Reranker:讓輕量本機模型(gemma4:e4b)擔任第二階段排序的評審,重新洗牌候選 Evidence!

兩階段檢索架構(Two-Stage Retrieval Pipeline)

在程式碼庫檢索中,絕對不能一開始就把所有程式碼都丟給 LLM 判斷,那樣耗時會以分鐘計算。工業界標準做法是:

使用者問題
   │
   ▼
【第一階段:粗篩 (High Recall)】
Hybrid Search (Lexical + Vector) 快速召回 Top 10~20 候選
   │
   ▼
【第二階段:精排 (High Precision)】
本地 LLM (gemma4:e4b) 逐一精讀候選片段,輸出最佳排序
   │
   ▼
最終 Top 3~5 乾淨 Evidence

本地模型最現實的工程問題:輸出格式失控

本地小模型(如 4B / 9B 等級)在做 Rerank 時最容易遇到的現實痛點是:

  1. 胡言亂語:叫它排序,它卻開始親切地「解釋程式碼」、「回答問題」或「補寫範例」。
  2. JSON 損毀:輸出未閉合的括號、多餘的 Markdown 引號或雜訊字元。
  3. 模型服務超時 / 當機:本地顯卡記憶體不足時回應失敗。

如果 Reranker 失敗就導致整個 CLI 噴錯中斷,這套工具就失去了可靠性。

因此,本篇的靈魂是:「Fallback 不是多餘的保險,而是不可或缺的防線。」

一旦 Ollama 逾時、未回應或產生的 JSON 編號不合法,系統直接安全降級回原本的第一階段排序,確保檢索管線永遠不會掛掉!

核心實作:app/llm_reranker.py

在 app/llm_reranker.py 中實作調用邏輯、Prompt 限制與嚴格的驗證過濾:

# app/llm_reranker.py
import json
import re
import urllib.request
import urllib.error
from typing import List, Dict, Any

class LLMReranker:
    def __init__(self, model: str = "gemma4:e4b", host: str = "http://localhost:11434", timeout: int = 15):
        self.model = model
        self.url = f"{host}/api/chat"
        self.timeout = timeout

    def build_prompt(self, query: str, candidates: List[Dict[str, Any]]) -> str:
        """將候選片段加上編號,限制模型只輸出索引陣列"""
        items_text = []
        for idx, item in enumerate(candidates):
            content_snippet = item.get("content", "").strip()[:300]
            items_text.append(
                f"[{idx}] File: {item.get('path')}:{item.get('start_line')}\nSnippet:\n{content_snippet}"
            )

        prompt = (
            f"You are a code search ranker. Given the user query, rank the following code candidates by relevance.\n"
            f"Query: \"{query}\"\n\n"
            f"Candidates:\n"
            f"{chr(10).join(items_text)}\n\n"
            f"Instructions:\n"
            f"1. Evaluate how well each candidate answers the query.\n"
            f"2. Output ONLY a valid JSON list of integers representing candidate indices from most to least relevant.\n"
            f"3. Do not include any explanations, markdown, or code blocks.\n"
            f"Example output format: [2, 0, 1]"
        )
        return prompt

    def rerank(self, query: str, candidates: List[Dict[str, Any]], limit: int = 5) -> List[Dict[str, Any]]:
        if not candidates:
            return []

        # 候選過少時不需增加額外延遲
        if len(candidates) <= 1:
            return candidates[:limit]

        prompt = self.build_prompt(query, candidates)
        payload = {
            "model": self.model,
            "messages": [{"role": "user", "content": prompt}],
            "stream": False,
            "options": {"temperature": 0.0} # 零溫度確保確定性
        }

        try:
            req = urllib.request.Request(
                self.url,
                data=json.dumps(payload).encode("utf-8"),
                headers={"Content-Type": "application/json"}
            )
            with urllib.request.urlopen(req, timeout=self.timeout) as resp:
                data = json.loads(resp.read().decode("utf-8"))
                raw_reply = data["message"]["content"].strip()

            # 解析與驗證模型輸出的排序清單
            order = self._parse_json_indices(raw_reply, len(candidates))
            if not order:
                # 格式損壞時觸發 Fallback
                return candidates[:limit]

            # 根據模型排序重新洗牌
            reranked = [candidates[i] for i in order]
            
            # 若模型漏掉某些索引,將漏掉的按原順序補在最後
            seen = set(order)
            for idx, item in enumerate(candidates):
                if idx not in seen:
                    reranked.append(item)

            return reranked[:limit]

        except (urllib.error.URLError, TimeoutError, json.JSONDecodeError, Exception):
            # 關鍵防護:任何網路超時或模型異常,一律安全回退原始排序
            return candidates[:limit]

    def _parse_json_indices(self, reply: str, max_len: int) -> List[int]:
        """從模型文字中萃取整數陣列,並嚴格過濾非法索引"""
        try:
            # 嘗試直接解析
            parsed = json.loads(reply)
            if isinstance(parsed, list):
                return self._validate_indices(parsed, max_len)
        except json.JSONDecodeError:
            pass

        # 若帶有 Markdown 或雜訊,嘗試用正則抓取 [ ... ]
        match = re.search(r"\[[\d\s,]+\]", reply)
        if match:
            try:
                parsed = json.loads(match.group(0))
                if isinstance(parsed, list):
                    return self._validate_indices(parsed, max_len)
            except json.JSONDecodeError:
                pass

        return []

    def _validate_indices(self, indices: List[Any], max_len: int) -> List[int]:
        valid = []
        for x in indices:
            if isinstance(x, int) and 0 <= x < max_len and x not in valid:
                valid.append(x)
        return valid

實作單元測試:tests/unit/test_llm_reranker.py

測試重點不在調用真實模型,而是驗證各種損壞輸出下的容錯防禦機制:

# tests/unit/test_llm_reranker.py
from app.llm_reranker import LLMReranker

def test_parse_json_indices_clean():
    reranker = LLMReranker()
    # 正常 JSON
    res = reranker._parse_json_indices("[2, 0, 1]", max_len=3)
    assert res == [2, 0, 1]

def test_parse_json_indices_with_markdown_noise():
    reranker = LLMReranker()
    # 夾帶額外文字的常見錯誤輸出
    noisy_reply = "Here is the ranked order:\n```json\n[1, 0, 2]\n```\nHope it helps!"
    res = reranker._parse_json_indices(noisy_reply, max_len=3)
    assert res == [1, 0, 2]

def test_parse_json_indices_out_of_bounds_filtering():
    reranker = LLMReranker()
    # 索引越界防護(如長度只有 3,但模型噴出 99)
    res = reranker._parse_json_indices("[1, 99, 0, -5]", max_len=3)
    assert res == [1, 0]

def test_fallback_on_corrupted_response():
    reranker = LLMReranker()
    candidates = [{"path": "a.py"}, {"path": "b.py"}]
    # 傳入完全無法解析的文字
    invalid_reply = "I cannot rank these files."
    indices = reranker._parse_json_indices(invalid_reply, max_len=2)
    assert indices == []

執行測試確認全部通過:

uv run pytest tests/unit/test_llm_reranker.py -v

tests/unit/test_llm_reranker.py::test_parse_json_indices_clean PASSED    [ 25%]
tests/unit/test_llm_reranker.py::test_parse_json_indices_with_markdown_noise PASSED [ 50%]
tests/unit/test_llm_reranker.py::test_parse_json_indices_out_of_bounds_filtering PASSED [ 75%]
tests/unit/test_llm_reranker.py::test_fallback_on_corrupted_response PASSED [100%]
============================== 4 passed in 0.04s ==============================

實際執行 Rerank 驗證:查詢 mobileai-local-rag

確認本地模型 gemma4:e4b 就緒後,我們針對昨天 MRR 偏低的自然語言題目進行重排測試:

uv run python -m app.cli search-rerank "對話問答的進入點在哪個檔案?"

第一階段 Hybrid 輸出:

  1. src/rag_common.py:1-20(設定與通用常數宣告)
  2. src/build_index.py:42(建立索引邏輯)
  3. src/rag_chat.py:65(def chat_loop(): 真正的對話迴圈)

第二階段 LLM Reranker 洗牌後:

[LLM Reranked Evidence]
1. [Rank 1 (was 3)] src/rag_chat.py:65
   def chat_loop(): 處理使用者對話輸入與 RAG 回應的主迴圈

2. [Rank 2 (was 1)] src/rag_common.py:1
   共用設定與模型設定

3. [Rank 3 (was 2)] src/build_index.py:42
   向量索引建立實作

本地 LLM 精準辨認出「對話進入點」指的是聊天主迴圈 chat_loop,成功將原本排在第 3 名的真實答案升到了第 1 名!

總結與下一步

今天我們完成了檢索系統的最後一塊拼圖:

  1. 雙階段精排:第一階段負責速度與覆蓋,第二階段交給本機模型精細比對。
  2. 零崩潰 Fallback:透過嚴格的正規表示式與索引邊界過濾,模型格式損壞時自動無痛降級。

但 LLM Reranker 不是免費的午餐——它帶來了顯著的模型推理延遲。

明天,我們將深入打磨 Reranker 的 Prompt 與防呆設定;並在 Day 20 重新執行自動化評估,拿出數據對比:加上 LLM Reranker 後,MRR 到底提升了多少?這個延遲成本到底值不值得?


上一篇
Day 17:把評估變成一條可以重跑的指令:打造自動化 Eval Runner
系列文
30 天打造 Codebase Intelligence Agent:從程式碼檢索、結構化索引到變更影響分析實戰 共 18 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言