上一篇處理的是 Sandbox capacity:
但拿到 sandbox 之後,還有另一個問題:
這個 task 到底需要多少 compute?
假設現在每個 sandbox 預設是:
1 vCPU
1 GB Memory
Agent 跑了一段 Python,結果 memory 爆掉。
這時可能有幾種情況:
所以這次先不急著 scale up,而是先把每個 execution step 量起來。
Agent 執行 Python 時,大概會經過:
write_file
→ analysis.py
execute
→ python3 analysis.py
而所有 execute 都會經過前面做的 LazySandbox。
所以 measurement 可以直接放在 execution layer:
Agent
↓
execute()
↓
LazySandbox
↓
Metering
↓
Real Sandbox
每次 command 執行時會記錄:
measurement 另外存進 conversation history,Agent 看到的還是原本 command 的 output。
也就是 Agent 不需要自己回報用了多少 CPU、多少 memory,這些由 execution layer 自己量。
先在 default sandbox 上跑幾個不同 workload:
| Step | Total | CPU | Peak Memory |
|---|---|---|---|
echo hello |
1.7 s | 0.00 s | 12 MB |
| Python + import Polars | 1.5 s | 0.14 s | 40 MB |
| Chart over 40K rows | 2.9 s | 1.0 s | 210 MB |
| 2M-row group-by | 2.3 s | 0.4 s | 192 MB |
| Hold 600 MB | 1.7 s | 0.2 s | 609 MB |
| Hold 1.5 GB | killed | - | 973 MB |
這邊先看到兩件事。
第一個是 Daytona 每次 command 本身大概就有 1.5 秒左右的 overhead。
第二個是目前這些 analysis workload 裡,CPU 看起來不是主要問題,memory 反而比較容易先碰到限制。
所以 execution 慢或失敗時,不能直接猜是 CPU 不夠。
先量才知道 bottleneck 到底在哪裡。
原本 command 如果因為 memory 被 kill,Agent 可能只看到一個 exit code。
這其實沒什麼幫助。
Agent 不知道是:
所以現在如果 measurement 發現 process 被 kill,而且 peak memory 已經很接近 sandbox limit,就會把原因一起回給 Agent。
例如:
Killed: command ran out of memory.
Peak memory: 973 MB
Sandbox limit: 1024 MB
重點不是只告訴 Agent「失敗了」,而是讓它知道:
為什麼失敗,以及下一步可以怎麼改。
前面收集的 measurement 最後也做成一個 Load page。
可以看到:

除了看最後有沒有成功之外,也可以直接看到不同 conversation 的 execution pattern。
先看一個正常的 case。

User 問:
Compute the median and 90th percentile of peak views per video for each category in the sandbox with Python, and make an interactive bar chart of the medians.
資料本來在 Postgres。
Agent 先用 export_query 算出每支影片的 peak views,再把結果輸出成 Parquet 到 sandbox,最後大約是 6,351 rows。
接著 Agent:
整個 code execution 只有一個 step,peak memory 大約 208 MB / 1 GB,run time 約 1.2 秒。

接下來換成比較大的資料。

User 想分析一份大約 20 million events 的 event log,並計算不同 device 的 typical session length。
Agent 一開始先檢查資料,接著用 Polars 計算 median。
但真正執行 analysis 時,memory 很快碰到 sandbox limit,process 被 OOM kill。

這時 system 沒有直接換成 bigger sandbox,而是先把 memory problem 告訴 Agent,讓它換 execution strategy。
Agent 接著把原本的 Polars code rewrite 成 DuckDB,再重新執行。

這次同一份大資料就可以完成。
這個 case 比較重要的是:
OOM
→ 不一定要 Bigger Sandbox
→ Rewrite
→ Retry
透過rewrite execution strategy 來降低使用的記憶體量。
例如原本可能是:
df = pl.read_parquet(path)
result = (
df.group_by("device")
.agg(pl.col("session_length").median())
)
如果這段 code 已經因為 memory 不夠失敗,再跑一次通常不會突然成功。
所以 Agent 可以改成 DuckDB:
SELECT
device,
median(session_length)
FROM 'events.parquet'
GROUP BY device;
整個 recovery path 變成:
OOM
→ Agent sees memory problem
→ Rewrite with DuckDB
→ Retry
目前給 Agent 的方向也跟著調整:
這裡不是說 DuckDB 比 Polars 好。
比較像是用在不同的情況:
重點是:
先調整 execution strategy,再考慮加 resource。
做到這裡,其實已經有兩種不同情況:
所以這次想做的不是:
OOM
→ Scale Up
而是:
Measure
→ Understand the bottleneck
→ Rewrite / Adjust Execution Strategy
→ Retry
先知道 task 到底用了多少 resource,再決定下一步。