iT邦幫忙

2026 iThome 鐵人賽

DAY 15
0
AI 自動化

AaaS from Scratch: 從一次性定義,到規模化分析系列 第 15 篇

[Day 15] 用 Jev 來做 Skill Routing

  • 分享至 

  • xImage
  •  

前面我們已經加入 Skill,讓 Agent 在需要時才載入額外的 knowledge 或 task guidance。

User Request
    ↓
Agent
    ↓
Read relevant SKILL.md
    ↓
Continue the task

Skill 少的時候,Agent 自己從 description 判斷就夠了。

但當 Skill 慢慢變多:

5
20
50
100

問題就變成:

Agent 要怎麼知道這次應該先讀哪一個 Skill?

這其實是一個 routing problem。

前面做的 SkillEnforcerMiddleware 比較適合處理 hard requirement,例如某個 action 前一定要先讀某個 Skill。

但「這次 request 跟哪個 Skill 最相關?」比較像 judgment。

所以今天想加入 Jev,讓它負責這一層 Skill Routing。

Jev 是什麼?

Jev 是由 Typesafe.AI 推出,用 RLCD 訓練的模型。

Jev 跟一般 ChatModel 的使用方式不太一樣。

一般 LLM 通常是:

Input
→ Model
→ Generated Text

但在 Automation 裡,routing 通常不需要一大段自然語言。

我們真正需要的是:

要 / 不要
選哪一個
給分 / 排名

Jev 的 input 主要包含:

State
→ 要判斷的內容

Questions
→ 要回答的問題

例如:

{
  "state": {
    "request": "Write a report on which categories are winning"
  },
  "model": "jev-latest",
  "questions": {
    "which_skill": {
      "type": "choice",
      "instructions": "Which skill is the right one for this request?",
      "criteria": {
        "query_database": "Use when the request needs database analysis.",
        "trend_report": "Use when the request asks for a report, summary, or write-up.",
        "forecast": "Use when the request asks about future values or trends.",
        "none": "No skill is needed."
      }
    }
  }
}

它回的不是一般 generated text,而是 typed judgment。

而且回答範圍是我們先定義好的。

例如:

query_database   0.08
trend_report     0.87
forecast         0.03
none             0.02

所以 downstream 可以直接拿結果繼續執行,不需要再 parse 一段自然語言。

三種 Judgment Type

Jev 主要有三種 judgment type:

Type 問題類型
Choice 哪一個 option 最符合?
Noul Yes / No?
Score 程度有多高?

這篇會用 Choice,因為我們真正想問的是:

Which skill best matches this request?

Jev 也可以在同一個 request 裡,對同一份 state 一次回答多個 questions。

包一層 JevJudge

為了不要讓 Agent code 到處直接處理 Jev API,我另外包了一層 JevJudge。

Method name 直接對應 Jev 的 judgment type:

class JevJudge:
    async def choice(...):
        ...

    async def noul(...):
        ...

    async def score(...):
        ...

    async def choice_then_noul(...):
        ...

用途大概是:

choice()
→ Choice
→ 從所有 labels 裡找最符合的

noul()
→ Noul
→ 每個 label 分別判斷 Yes / No

score()
→ Score
→ 每個 label 分別評分

choice_then_noul()
→ Choice + Noul
→ 先找候選,再確認其他候選是否也適用

目前 Skill Routing 的核心是 choice():

User Request
    ↓
choice()
    ↓
Choice over all Skills + none
    ↓
Best matching Skill

所有 Skill 一起競爭,所以很適合找出「最相關的那一個」。

top_n 和 multi

Skill Router 還有兩個設定:

top_n: 1
multi: false

top_n 決定最多可以選幾個 Skill。

目前:

top_n: 1

所以只需要 choice() 找出最相關的一個。

如果之後一個 request 可能同時需要多個 Skill,例如:

Forecast next month's views
and show it as an interactive chart

可能得到:

forecast            0.70
chart-interactive   0.25
correlation         0.05

如果:

top_n: 2
multi: false

就直接拿前兩名:

forecast
chart-interactive

但這樣有個問題:如果 request 其實只需要一個 Skill,第二名也可能被一起選進來。

所以 multi: true 時,會改用 choice_then_noul():

Choice
→ 先找最可能的 Skill

Noul
→ 再確認 runners-up 是否也適用

只有通過確認的候選才會留下。

也就是:

multi: false
→ ranking

multi: true
→ ranking + verification

代價是 multi: true 會多一次 Jev request。

目前我的設定是:

top_n: 1
multi: false

所以實際上 Skill Routing 只會用 choice()。

choice_then_noul() 則留給之後真的需要 multi-skill routing 的情況。

Skill Description 是 Interface

Router 不會先讀完整的 SKILL.md。

它主要看每個 Skill 的:

name
description

所以 description 不只是 documentation,而是:

Skill discovery interface

如果 description 太模糊,或不同 Skill 之間 overlap 太多,routing 表現也會跟著下降。

放進 Agent Pipeline

最後 flow 變成:

User Request
      ↓
Jev Skill Router
      ↓
Relevant Skill Hint
      ↓
Agent
      ↓
Read SKILL.md
      ↓
Continue the task

Router 的結果只是 hint,不是 hard rule。

如果 Jev timeout 或 API error:

Router Failure
→ Agent continues normally

所以 routing 是 optimization,不是 single point of failure。

Routing 和 Enforcement

最後我會把 responsibility 分成:

Jev
→ Which Skill is relevant?

Code
→ Is this Skill absolutely required before this action?

也就是:

Soft relevance
→ Routing

Hard dependency
→ Enforcement

這樣可以保留 progressive disclosure:

Many Skills
    ↓
Small metadata
    ↓
Routing
    ↓
Read only relevant SKILL.md

上一篇
[Dat 14] Data Upload 與 HTML Report
下一篇
[Day 16] 用 Jev 做 Context Management(1)
系列文
AaaS from Scratch: 從一次性定義,到規模化分析 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言