前幾天我們已經完成:
但如果每次都要打開 JSON 檔案,再手動比較數據,當實驗數量增加後,很快就會變得難以管理
所以今天要把前幾天累積的資料進一步視覺化,建立一個簡單的 RAG Evaluation Dashboard
目標很簡單:
讓 Evaluation 與 Experiment Tracking 從「資料」變成「可以快速判讀的資訊」
一個實用的 RAG Evaluation Dashboard,不需要一開始就做得非常複雜
今天不需要重新執行 RAG
如果前面的環境已經建立好,只需要新增:
pip install streamlit pandas plotly
先建立一個簡單的資料載入函式
from pathlib import Path
import json
import pandas as pd
def load_experiments(directory="experiments"):
rows = []
for experiment_dir in Path(directory).iterdir():
if not experiment_dir.is_dir():
continue
config_path = experiment_dir / "config.json"
metrics_path = experiment_dir / "metrics.json"
if not config_path.exists() or not metrics_path.exists():
continue
with open(config_path, "r", encoding="utf-8") as f:
config = json.load(f)
with open(metrics_path, "r", encoding="utf-8") as f:
metrics = json.load(f)
rows.append({
"experiment_id": config.get("experiment_id"),
"model": config.get("model"),
"dataset_version": config.get("dataset_version"),
"prompt_version": config.get("prompt_version"),
"chunk_size": config.get("chunk_size"),
"chunk_overlap": config.get("chunk_overlap"),
"top_k": config.get("top_k"),
"context_precision": metrics.get("context_precision"),
"answer_relevance": metrics.get("answer_relevance"),
"faithfulness": metrics.get("faithfulness"),
"latency": metrics.get("latency"),
})
return pd.DataFrame(rows)
這裡最重要的概念是:
把原本分散在不同 JSON 的資訊,整理成可以分析的 DataFrame
Dashboard 負責呈現結果,不代表單一數值就能證明某個參數一定比較適合所有資料集
因此仍然要搭配實驗數據比較
而統計的價值不只是找到數值高的項目
真正重要的是:
它可以幫助我們決定下一輪實驗應該從哪個問題開始
讓 Dashboard 從「查看資料」變成「探索實驗」
讓 Dashboard 可以從「全部實驗」切換到「特定實驗群組」
今天真正完成的不是一個漂亮的網頁,而是把前面累積的 RAG Evaluation 資料轉換成可以快速理解的資訊
到目前為止,我們已經不只是「做出一個可以回答問題的 RAG」
而是開始建立一套可評估、可追蹤的 RAG 系統
Dashboard 解決的是:
「我們如何看懂系統品質」
但當 RAG 真正部署到 Production 後,新的問題會出現:
讓我們開始思考一個真正上線的 AI Assistant,除了「回答得對不對」之外,還需要如何被監控與維運