前面我們已經加入 Skill,讓 Agent 在需要時才載入額外的 knowledge 或 task guidance。
User Request
↓
Agent
↓
Read relevant SKILL.md
↓
Continue the task
Skill 少的時候,Agent 自己從 description 判斷就夠了。
但當 Skill 慢慢變多:
5
20
50
100
問題就變成:
Agent 要怎麼知道這次應該先讀哪一個 Skill?
這其實是一個 routing problem。
前面做的 SkillEnforcerMiddleware 比較適合處理 hard requirement,例如某個 action 前一定要先讀某個 Skill。
但「這次 request 跟哪個 Skill 最相關?」比較像 judgment。
所以今天想加入 Jev,讓它負責這一層 Skill Routing。
Jev 是由 Typesafe.AI 推出,用 RLCD 訓練的模型。
Jev 跟一般 ChatModel 的使用方式不太一樣。
一般 LLM 通常是:
Input
→ Model
→ Generated Text
但在 Automation 裡,routing 通常不需要一大段自然語言。
我們真正需要的是:
要 / 不要
選哪一個
給分 / 排名
Jev 的 input 主要包含:
State
→ 要判斷的內容
Questions
→ 要回答的問題
例如:
{
"state": {
"request": "Write a report on which categories are winning"
},
"model": "jev-latest",
"questions": {
"which_skill": {
"type": "choice",
"instructions": "Which skill is the right one for this request?",
"criteria": {
"query_database": "Use when the request needs database analysis.",
"trend_report": "Use when the request asks for a report, summary, or write-up.",
"forecast": "Use when the request asks about future values or trends.",
"none": "No skill is needed."
}
}
}
}
它回的不是一般 generated text,而是 typed judgment。
而且回答範圍是我們先定義好的。
例如:
query_database 0.08
trend_report 0.87
forecast 0.03
none 0.02
所以 downstream 可以直接拿結果繼續執行,不需要再 parse 一段自然語言。
Jev 主要有三種 judgment type:
| Type | 問題類型 |
|---|---|
Choice |
哪一個 option 最符合? |
Noul |
Yes / No? |
Score |
程度有多高? |
這篇會用 Choice,因為我們真正想問的是:
Which skill best matches this request?
Jev 也可以在同一個 request 裡,對同一份 state 一次回答多個 questions。
為了不要讓 Agent code 到處直接處理 Jev API,我另外包了一層 JevJudge。
Method name 直接對應 Jev 的 judgment type:
class JevJudge:
async def choice(...):
...
async def noul(...):
...
async def score(...):
...
async def choice_then_noul(...):
...
用途大概是:
choice()
→ Choice
→ 從所有 labels 裡找最符合的
noul()
→ Noul
→ 每個 label 分別判斷 Yes / No
score()
→ Score
→ 每個 label 分別評分
choice_then_noul()
→ Choice + Noul
→ 先找候選,再確認其他候選是否也適用
目前 Skill Routing 的核心是 choice():
User Request
↓
choice()
↓
Choice over all Skills + none
↓
Best matching Skill
所有 Skill 一起競爭,所以很適合找出「最相關的那一個」。
top_n 和 multiSkill Router 還有兩個設定:
top_n: 1
multi: false
top_n 決定最多可以選幾個 Skill。
目前:
top_n: 1
所以只需要 choice() 找出最相關的一個。
如果之後一個 request 可能同時需要多個 Skill,例如:
Forecast next month's views
and show it as an interactive chart
可能得到:
forecast 0.70
chart-interactive 0.25
correlation 0.05
如果:
top_n: 2
multi: false
就直接拿前兩名:
forecast
chart-interactive
但這樣有個問題:如果 request 其實只需要一個 Skill,第二名也可能被一起選進來。
所以 multi: true 時,會改用 choice_then_noul():
Choice
→ 先找最可能的 Skill
Noul
→ 再確認 runners-up 是否也適用
只有通過確認的候選才會留下。
也就是:
multi: false
→ ranking
multi: true
→ ranking + verification
代價是 multi: true 會多一次 Jev request。
目前我的設定是:
top_n: 1
multi: false
所以實際上 Skill Routing 只會用 choice()。
choice_then_noul() 則留給之後真的需要 multi-skill routing 的情況。
Router 不會先讀完整的 SKILL.md。
它主要看每個 Skill 的:
name
description
所以 description 不只是 documentation,而是:
Skill discovery interface
如果 description 太模糊,或不同 Skill 之間 overlap 太多,routing 表現也會跟著下降。
最後 flow 變成:
User Request
↓
Jev Skill Router
↓
Relevant Skill Hint
↓
Agent
↓
Read SKILL.md
↓
Continue the task
Router 的結果只是 hint,不是 hard rule。
如果 Jev timeout 或 API error:
Router Failure
→ Agent continues normally
所以 routing 是 optimization,不是 single point of failure。
最後我會把 responsibility 分成:
Jev
→ Which Skill is relevant?
Code
→ Is this Skill absolutely required before this action?
也就是:
Soft relevance
→ Routing
Hard dependency
→ Enforcement
這樣可以保留 progressive disclosure:
Many Skills
↓
Small metadata
↓
Routing
↓
Read only relevant SKILL.md