到目前為止,這個系列每次執行 Agent,基本上都是從 terminal 開始:
uv run python scripts/ask.py "..."
看 trace、等 execution 完成,再自己去 reports/ 打開 chart。
拿來開發 Agent 沒問題,但這不會是 user 最後使用它的方式。
所以今天要幫 Agent 加上一個 frontend。
Frontend 用的是 Vite + React + TypeScript,backend 則是 FastAPI。
這篇不會特別深入 UI design,而是先處理 frontend 接上 Agent 之後最先遇到的兩個問題:
:chatgpt-content-reference{index="0"}
整體結構大概是:
Frontend Backend
Chat Page
│
├── message ───────────────► FastAPI
│ │
│ ▼
│ Agent + Daytona
│
◄──────────── SSE Stream
目前 frontend 主要有兩個 page:
Chat
Data
今天先聚焦在 Chat。

:chatgpt-content-reference{index="1"}
Agent execution 跟一般 API request 很不一樣。
一個 turn 可能需要:
10 sec
30 sec
60 sec
如果 frontend 只顯示:
Loading...
user 完全不知道 Agent 現在正在做什麼。
但前面我們已經知道,Agent 實際上一直在跑 ReAct loop:
Reason
↓
Tool Call
↓
Observe
↓
Reason
↓
...
所以 frontend 也可以把這個 execution process stream 出來。
目前 /api/chat 使用 Server-Sent Events(SSE)。
每個 Agent step 會被轉成一個 event:
{"type": "thread", "thread_id": "..."}
{"type": "thinking", "text": "Let me inspect the data first."}
{"type": "tool_call", "name": "query_database", "args": {"sql": "..."}}
{"type": "tool_result", "content": "[{...}]"}
{"type": "answer", "text": "..."}
{"type": "artifact", "name": "report.html", "url": "..."}
{"type": "done"}
Backend 主要就是把 Agent 的 astream() 轉成 frontend 比較容易使用的 event。
Conceptually:
async for update in agent.astream(...):
...
yield {
"type": "tool_call",
...
}
這樣 UI 就可以即時顯示:
Working...
✓ Read skill
✓ Query database
✓ Export data
✓ Execute Python
✓ Generate report
等 final answer 回來之後,再把這些 execution steps collapse 起來。

我覺得這裡很重要的一點是:
Frontend 不只是顯示 final answer,也可以把 Agent 的 execution process 變成 user 可以觀察的東西。
:chatgpt-content-reference{index="2"}
Browser 原生的 EventSource 很方便,但它只支援 GET。
Chat message 顯然比較適合:
POST /api/chat
因為 message、thread id 等資訊都要放在 body。
所以 frontend 不是直接使用 EventSource,而是:
fetch()
↓
read response stream
↓
parse SSE events
這裡比較麻煩的是 network chunk 不一定剛好跟 event boundary 對齊。
可能會收到:
chunk 1:
data: {"type":"tool_
chunk 2:
call","name":"query_database"}
所以 parser 需要先 buffer,直到遇到完整 event 再處理。
這種 implementation detail 平常不太會注意,但做 streaming UI 時很容易踩到。
:chatgpt-content-reference{index="3"}
有了 frontend 之後,就會開始有 follow-up question:
User:
Which category gets the most views?
Agent:
...
User:
Can you visualize the top 5?
第二個問題不應該是一個全新的 Agent run。
它應該延續前面的 conversation。
目前一個 conversation 大概包含:
Conversation
├── Agent
├── Memory
└── Sandbox
每個 conversation 都有自己的 thread_id。
第一次進來時:
Create Agent
Create Sandbox
Create Conversation State
之後 follow-up:
Reuse Agent
Reuse Memory
Reuse Sandbox
:chatgpt-content-reference{index="4"}
假設 turn 1 裡 Agent 已經:
export_query(...)
↓
data.parquet
↓
generated analysis.py
如果 turn 2 user 接著說:
Can you make another chart?
我們不希望重新建立一個全新的 sandbox,再把所有東西重新 export 一次。
比較合理的是:
Turn 1
→ Sandbox A
→ data.parquet
Turn 2
→ Same Sandbox A
→ reuse data.parquet
這樣 follow-up 會快很多,也比較符合:
conversation
=
working session
這個概念。
但 Daytona sandbox 留著就會持續產生 cost。
所以目前 conversation 有 idle timeout。
例如:
15 minutes idle
↓
Destroy Sandbox
不過 conversation memory 可以繼續保留。
所以下次 user 回來:
Conversation Memory
✓ still exists
Old Sandbox
✗ destroyed
New Sandbox
✓ created when needed
這裡把:
conversation lifetime
跟:
sandbox lifetime
拆開來。
我覺得這個 distinction 很重要。
Conversation 可以活得比較久,但 execution environment 不需要永遠存在。
:chatgpt-content-reference{index="5"}
同一個 conversation 裡,我目前只允許一次處理一個 turn。
如果 Agent 還在回答,user 又送第二個 message,不會同時讓兩個 run 共用同一個 sandbox。
Conceptually:
Turn 1
→ Running
Turn 2 arrives
→ Reject / wait
不然很容易出現 race condition。
例如:
Turn 1 writes file A
Turn 2 overwrites file A
Turn 1 reads changed file
所以 conversation 不只是保存 history。
它其實也是一個 execution boundary。
原本的 implementation 也是同一個方向:同一個 conversation 同時間只允許一個 turn,避免兩個 execution 同時操作同一個 sandbox。:chatgpt-content-reference{index="6"}
今天先把 Agent 從 terminal 拉到 frontend。
主要處理兩件事:
Streaming
→ user 看得到 Agent 正在做什麼
Conversation + Sandbox Reuse
→ follow-up 可以延續同一個 working session
現在 user 已經可以:
Ask Question
→ Watch Agent Execution
→ Get Answer
→ Continue Conversation
Frontend 不只是把 Agent 包上一層 UI。
它也讓我們開始需要處理:
long-running execution
conversation state
sandbox lifecycle
concurrency
下一篇再補上 frontend 的另外一半:
讓 user upload 自己的 data,以及安全地 render Agent 產生的 HTML report。