iT邦幫忙

2026 iThome 鐵人賽

DAY 18
0
AI 自動化

AaaS from Scratch: 從一次性定義,到規模化分析系列 第 18 篇

[Day 18] Sandbox Lifecycle:從 Create 到 Cleanup

  • 分享至 

  • xImage
  •  

前面我們已經讓 Agent 可以在 sandbox 裡執行 Python。

一個 conversation 第一次需要執行 code 時建立 sandbox,後續就繼續使用同一個 sandbox。

Conversation
    ↓
Sandbox

在細節上有這些問題要一一考慮:

Sandbox 掛掉怎麼辦?

Sandbox idle 太久被刪掉,
下一次還能繼續聊嗎?

Server restart 之後,
之前留下來的 sandbox 怎嗎處里?

如果未來要換 sandbox provider,
ConversationManager 也要一起重寫嗎?

所以這次先不處理「同時可以開多少 sandbox」。

先把 sandbox 在 system 裡的位置整理清楚。

Architecture

目前大概拆成:

Browser
   ↓
FastAPI
   ↓
ConversationManager
   ↓
LazySandbox
   ↓
SandboxProvider
   ↓
Daytona Sandbox

另外 conversation history 還是由 application 自己保存:

Conversation
├── HistoryStore
├── Context Filter
└── Sandbox Reference

這裡幾個 component 的責任不太一樣。

ConversationManager

ConversationManager 管的是:

哪個 conversation 有哪個 sandbox
Sandbox 現在是什麼狀態
什麼時候建立
什麼時候檢查
什麼時候刪掉

它知道:

conversation_id
→ sandbox

但它不應該知道 Daytona API 到底怎麼 create sandbox。

SandboxProvider

真正跟 provider 溝通的是:

SandboxProvider

大概提供:

create
prepare
destroy
is_alive
find

今天我們用 Daytona:

DaytonaProvider

未來如果真的換成其他 provider,上面的 lifecycle decision 不需要全部一起改。

所以 dependency 大概是:

ConversationManager
        ↓
SandboxProvider
        ↓
Daytona

Manager 決定:

什麼時候要一個 sandbox。

Provider 決定:

怎麼建立那個 sandbox。

LazySandbox

Agent 本身也不直接拿 Daytona,而是拿:

LazySandbox

可以先把它想成 Agent 和 real sandbox 中間的一層:

Agent
  ↓
LazySandbox
  ↓
Real Sandbox

這樣 Agent 不需要知道底下現在到底是哪一個 sandbox。

對 Agent 來說只是在做:

write_file(...)
execute(...)

至於 real sandbox 是剛建立的、重新建立的,或之後被換掉,都不是 Agent 本身需要處理的事情。

這一層後面還會有其他用途,先留到後面的 compute management 再講。

Conversation != Sandbox

這次最重要的改變其實是:

Conversation 和 Sandbox 有不同的 lifetime。

Conversation 比較像:

long-lived logical session

裡面有:

messages
history
context
analysis state

Sandbox 比較像:

temporary execution resource

裡面有:

Python process
temporary files
scripts
intermediate results

例如 user 一段時間沒有講話:

Conversation
→ still exists

Sandbox
→ deleted

下一次 user 回來:

Conversation
→ load history

Need Python?
→ create another sandbox

Conversation 不需要因為 sandbox 消失就一起消失。

Sandbox Lifecycle

既然 sandbox 是一個 temporary resource,就需要有自己的 lifecycle。

最基本的流程大概是:

new
 ↓
starting
 ↓
assigned
 ↓
busy
 ↓
idle
 ├──→ busy
 └──→ deleted

Starting

Conversation 第一次需要 sandbox:

Conversation
    ↓
starting
    ↓
Provider creates sandbox
    ↓
prepare environment
    ↓
assigned

prepare 可能包含:

install packages
prepare working directory
copy required files

完成之後,這個 sandbox 才正式屬於 conversation。

Busy

當 Agent 真的開始執行 code 時,sandbox 會進入 busy 狀態。

目前同一個 conversation 不會同時執行兩個 turn,主要是避免兩個 request 同時修改同一個 sandbox 裡的 files 或 state。

Turn A
→ modifying files

Turn B
→ modifying the same files

Turn 完成之後:

busy
 ↓
idle

Idle

Idle 不代表 sandbox 馬上被刪掉。

因為 user 很可能幾分鐘後接著問:

Can you chart that?

如果前面的:

analysis.py
result.parquet
chart data

還在 sandbox 裡,就可以直接繼續使用。

根據 Daytona 的官方文件, sandbox 閒置超過 15 分鐘才會被刪除:

Sandbox 掛掉怎麼辦?

另外一條 lifecycle 是 failure,例如 conversation 還在,但 sandbox 已經死掉:

Conversation
    ↓
 Sandbox
    ↓
  dead

這時不應該變成:

Sandbox dead
→ Conversation dead

而是:

Sandbox dead
→ remove sandbox reference
→ create another sandbox when needed

Conversation history 還在。

真正會失去的是 sandbox 裡還沒有另外保存的 temporary files。

Server Crash 也會留下 Sandbox

還有一個比較容易忽略的情況:

Backend process dies

但 Daytona sandbox 是另外一個 resource,所以 backend 掛掉不代表 sandbox 自動消失。

可能會變成:

Server
→ dead

Sandbox
→ still running

這些就是 orphaned sandboxes。

所以建立 sandbox 時會帶上可以識別它的 metadata。

下一次 server startup,可以找到:

上一個 server instance
留下來的 sandboxes

再把它們 cleanup。

正常 shutdown 也一樣:

Server shutdown
→ destroy managed sandboxes

Resource lifecycle 不能只處理:

create

還要確定最後一定有對應的:

destroy

Sandbox Event

這些 lifecycle transition 也會記錄下來。

例如:

sandbox_ready
sandbox_deleted
sandbox_replaced

跟 conversation history 放在一起。

所以之後如果遇到:

為什麼這一輪突然沒有之前的 file?

至少可以回頭看到:

Turn 12
→ sandbox died

Turn 13
→ new sandbox created

不然只看 Agent messages,很難知道 execution environment 中間發生了什麼。

到這裡還沒有處理 Capacity

現在 architecture 大概變成:

Conversation
│
├── HistoryStore
├── Context Filter
│
└── Sandbox Reference
        ↓
ConversationManager
        ↓
   LazySandbox
        ↓
  SandboxProvider
        ↓
   Real Sandbox

而 lifecycle 是:

create
→ assigned
→ busy
→ idle
→ reuse / replace / delete

這篇主要處理的是:

Sandbox 是什麼 resource,以及它的 lifetime 怎麼跟 Conversation 分開。

但現在還假設:

Conversation needs sandbox
→ create one

實際上當很多 user 同時進來,不可能無限 create:

10 conversations
100 conversations
1000 conversations

下一個問題就變成:

有限的 Sandbox capacity 到底要怎麼分?

也就是從 Sandbox Lifecycle 進到 Sandbox Scheduling & Capacity。


上一篇
[Day 17] 用 Jev 做 Context Management(2)
下一篇
[Day 19] Sandbox Scheduling:有限的 Capacity 怎麼分?
系列文
AaaS from Scratch: 從一次性定義,到規模化分析 共 21 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言