
Relix 已經具備一個小型 API 框架該有的東西,但工程上一定會被問,效能如何 ?
這篇不是要把 Relix 變成世界最快的框架,目標是用工程化的方式回答四個問題,怎麼量測、瓶頸在哪、接下來可以怎麼改、什麼時候該直接用成熟框架
下面的 HttpURLConnection 程式是循序送 request,只適合確認暖機後的單一路徑 latency 是否出現明顯退化,它沒有並行控制,也混入 client 與連線成本,不能拿來宣稱 server 吞吐量或高併發 p99
import java.net.HttpURLConnection
import java.net.URI
fun latencySmokeTest(
url: String,
warmup: Int = 1000,
iterations: Int = 10_000,
) {
// Warm up:讓 JVM JIT 編譯器暖機
repeat(warmup) {
val conn = URI(url).toURL().openConnection() as HttpURLConnection
conn.inputStream.readBytes()
conn.disconnect()
}
// 量測
val latencies = mutableListOf<Long>()
val start = System.nanoTime()
repeat(iterations) {
val reqStart = System.nanoTime()
val conn = URI(url).toURL().openConnection() as HttpURLConnection
conn.inputStream.readBytes()
conn.disconnect()
latencies.add(System.nanoTime() - reqStart)
}
val totalMs = (System.nanoTime() - start) / 1_000_000
latencies.sort()
val p50 = latencies[latencies.size / 2] / 1_000_000.0
val p95 = latencies[(latencies.size * 0.95).toInt()] / 1_000_000.0
val p99 = latencies[(latencies.size * 0.99).toInt()] / 1_000_000.0
println("=== Latency Smoke Test ===")
println("Requests: $iterations")
println("Total: ${totalMs}ms")
println("p50: ${"%.2f".format(p50)}ms")
println("p95: ${"%.2f".format(p95)}ms")
println("p99: ${"%.2f".format(p99)}ms")
}
URI(url).toURL() 不用 URL(url),是因為 JDK 20 之後那個建構式被標為 deprecated,照舊寫法編譯會噴兩個 warning
使用方式
fun main() {
val app = Relix { port = 8080 }
app.routing {
get("/hello") { ok("Hello, World!") }
}
app.start()
latencySmokeTest("http://localhost:8080/hello")
app.stop()
}
跑起來長這樣
Relix started on 0.0.0.0:8080
=== Latency Smoke Test ===
Requests: 10000
Total: 1338ms
p50: 0.12ms
p95: 0.20ms
p99: 0.35ms
Relix shutting down...
Relix stopped
這是單機跑一次的結果,循序送、沒有並行,只能用來確認暖機後單一路徑沒有明顯退化,不是 benchmark,換一台機器、換一個 JDK 版本、背景多開幾個程式,數字都會不一樣,也不能拿它推論 Relix 的吞吐量或高併發下的 p99
Relix stopped 印兩次不是 bug,stop() 有兩個呼叫來源,一個是 main() 裡的 app.stop(),一個是 start() 註冊的 shutdown hook,第 25 篇就講過這個坑,這裡沒有擋,它就老實地跑第二次,onStopping 的 hook 也會跟著跑兩次,這個範例沒註冊 hook 所以無所謂,真的有連線池要關,就照第 25 篇說的加一個 AtomicBoolean 讓 stop() 只生效一次
port 這裡跟著 application.properties 寫 8080,不要自己改成別的數字,第 25 篇定出來的優先順序是 DSL → 設定檔 → 環境變數,後者蓋前者,Relix { port = 9999 } 會被設定檔裡的 port=8080 蓋掉,server 起在 8080,smoke test 卻打 9999,跑起來就是一行 Relix started on 0.0.0.0:8080 之後接一個 java.net.ConnectException: Connection refused,真的要換 port,改 application.properties 或設 RELIX_PORT,兩邊的數字要一致
還有一點,main() 裡例外飛出來的時候 app.stop() 不會執行,JDK HttpServer 的 thread 不是 daemon thread,JVM 不會跟著結束,8080 會一直被佔住,下一次跑就變成 Address already in use,遇到的話先確認上一輪的 process 還在不在
為什麼要 warm up ? JVM 剛啟動時用 interpreter 跑位元碼,效能差,跑幾千次之後 JIT 編譯器會把 hot path 編譯成 native code,warm up 讓量測結果反映的是 JIT 後的效能,而不是 interpreter 的效能
為什麼記錄 percentile 而不是平均值 ? 平均值會被極端值拉偏,p95 告訴你「95% 的 request 在幾毫秒以內完成」,p99 告訴你 tail latency,對 API server 來說,p99 比平均值重要得多,一百個 request 裡有一個慢 2 秒,使用者體驗就差了
要量 throughput 與 tail latency,請用 wrk、hey 或 vegeta 之類能控制連線與並行數的工具,例如固定 30 秒、20 條連線,分別測空 handler 與完整 middleware 組合,每次記錄 CPU 型號、JDK 版本、heap、executor、工具版本、命令、commit 與原始輸出
wrk -t4 -c20 -d30s http://127.0.0.1:8080/hello
先做多次 warm-up,再至少跑三輪正式測試,比較 median 與分布,不只挑最好的一次,這是我這台機器上跑的三輪,打的是前面那個 /hello 空 handler,沒有裝任何 middleware
Running 30s test @ http://127.0.0.1:8080/hello
4 threads and 20 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 468.19us 1.13ms 75.38ms 99.47%
Req/Sec 11.75k 1.31k 13.28k 92.28%
1408312 requests in 30.10s, 174.60MB read
Requests/sec: 46787.28
Latency 425.97us 193.36us 12.16ms 73.47%
Req/Sec 11.86k 769.81 13.13k 82.72%
1421265 requests in 30.10s, 176.21MB read
Requests/sec: 47217.43
Latency 424.85us 184.46us 10.71ms 71.02%
Req/Sec 11.86k 686.05 12.99k 80.56%
1421308 requests in 30.10s, 176.21MB read
Requests/sec: 47219.63
median 是 47217,三輪之間差不到 1%,這個部分算穩定,分布就沒那麼漂亮,第一輪的 Max 是 75.38ms、標準差 1.13ms,後兩輪的 Max 掉到 10 到 12ms、標準差不到 200us,第一輪明顯還在暖機,JIT 還沒把 hot path 編完,這就是為什麼只跑一輪、只看 Requests/sec 會騙自己,平均值三輪都差不多,尾巴差了六倍
環境要一起記下來,數字才有意義
wrk -t4 -c20 -d30s http://127.0.0.1:8080/hello
換句話說,這組數字只在這個組合下成立,別拿去跟別人貼的 Ktor 或 Spring Boot 數字比,那些多半是不同機器、不同 JDK、不同 client 工具、甚至 client 跑在另一台機器上量出來的,要比就自己在同一台機器上把兩邊都跑一次
還有一件事這三輪沒有做到,前面說「分別測空 handler 與完整 middleware 組合」,上面只有空 handler 那一半,所以這裡不推論 middleware 的固定成本,真要知道 CORS、StatusPages、Authentication 疊上去各要花多少,得照同樣的流程再測一組帶 middleware 的路由,兩組的差值才是那一層的成本
Relix 用 com.sun.net.httpserver.HttpServer 當底層引擎,好處是零依賴、API 簡單,限制也很明顯
Executor 決定並行模型
JDK HttpServer 用 Executor 分派 request。若沒有明確設定,會使用實作提供的預設 executor,不能直接描述成固定的 thread-per-request,應該可以顯式設定 executor,讓 thread 數、queue 與拒絕政策可觀察,handler 若執行阻塞 I/O,執行它的 thread 仍會被占住
Request → Thread Pool (200 threads) → Handler (blocking I/O) → 阻塞等待
↓
thread 佔住 50ms
↓
200 concurrent requests 就滿了
流量控制能力有限
相關的 adapter 沒有定義 queue 上限、拒絕政策、request body streaming 或負載保護,實際行為取決於 executor 與底層實作,因此要靠 bounded queue、timeout、body limit、readiness 與外層 proxy 一起控制
HTTP/2 不支援
com.sun.net.httpserver 只支援 HTTP/1.1,如果你需要 multiplexing、server push、header 壓縮,得換底層
哪一層是瓶頸必須由 profiler 與分層 benchmark 證明,不要在沒有量測資料時先假設 router、pipeline 或底層 I/O 一定占多少時間
如果你想用 profiler (VisualVM、async-profiler、IntelliJ Profiler) 分析 Relix 的效能,可以看這幾個地方
CPU hotspot
Router.match(),路由比對,路由少 (< 100) 時不會是瓶頸,路由多時可以考慮 trie 取代線性掃描Pipeline.execute(),middleware 鏈的執行。middleware 數量通常是個位數,不會是問題Json.encodeToString() / Json.decodeFromString(),JSON 序列化反序列化,大 payload 會吃 CPU記憶體
ByteArray 分配,每個 response body 都建一個 ByteArray,高吞吐時 GC 壓力會上來Map<String, List<String>> headers,header 的分配和複製。每個 request/response 一組Thread 利用率
當框架的上層抽象 (RelixCall、Router、Pipeline、Plugin) 穩定之後,換底層引擎是合理的下一步
目前
Request → JDK HttpServer → RelixRequest → RelixCall → Pipeline → Handler → RelixResponse
替換後
Request → Netty → RelixRequest → RelixCall → Pipeline → Handler → RelixResponse
只有最左邊那一段需要改,RelixRequest 以右的東西完全不動
需要實作的介面
interface RelixEngine {
fun start(host: String, port: Int, handler: suspend (RelixRequest) -> RelixResponse)
fun stop(gracePeriodMs: Long)
}
class JdkEngine : RelixEngine { /* 目前的實作 */ }
class NettyEngine : RelixEngine { /* Netty 版 */ }
把 engine 做成可替換的介面。應用層選擇用哪個 engine
val app = Relix {
engine = NettyEngine()
port = 8080
}
Netty 的好處
這裡不做 Netty 的實作,但可以知道介面邊界在哪,你就知道該怎麼切,實作骨架大概是這樣
class NettyEngine : RelixEngine {
private lateinit var channel: Channel
private val bossGroup = NioEventLoopGroup(1) // 接 connection
private val workerGroup = NioEventLoopGroup() // 處理 I/O
override fun start(host: String, port: Int, handler: suspend (RelixRequest) -> RelixResponse) {
val bootstrap = ServerBootstrap()
.group(bossGroup, workerGroup)
.channel(NioServerSocketChannel::class.java)
.childHandler(object : ChannelInitializer<SocketChannel>() {
override fun initChannel(ch: SocketChannel) {
ch.pipeline()
.addLast(HttpServerCodec()) // HTTP 編解碼
.addLast(HttpObjectAggregator(1 shl 20)) // 把分塊請求合成一個 FullHttpRequest
.addLast(RelixHandlerAdapter(handler)) // 橋接到 RelixRequest/RelixResponse
}
})
channel = bootstrap.bind(host, port).sync().channel()
}
// stop() 略
}
event loop 能用少量 thread 管理大量連線,但不保證換成 Netty 就一定更快,handler 是否阻塞、buffer 策略、TLS、payload 與 executor 設定都會影響結果,是否替換 engine,應以相同 workload 的量測與維運需求決定
目前 Relix 用 runBlocking 在 JDK HttpServer 的 thread 上執行 suspend handler,這可以動,但有改進空間
// 目前:直接在 HttpServer 的 thread 上 runBlocking
executor.execute {
val response = runBlocking {
pipeline.execute(call)
}
writeResponse(exchange, response)
}
把同一個 executor 轉成 dispatcher,再在 executor callback 裡 runBlocking(dispatcher) 只會多一次排程,並不會把 adapter 變成 non-blocking
// 不建議:同一批 thread 之間重複排程
val dispatcher = executor.asCoroutineDispatcher()
executor.execute {
val response = runBlocking(dispatcher) {
pipeline.execute(call)
}
writeResponse(exchange, response)
}
若保留同步 JDK adapter,就接受 callback thread 會等到 response 完成,並使用 bounded executor 控制資源,若要真正 non-blocking,engine API 必須以 coroutine 啟動 request、在完成時非阻塞地寫回 response,不能只替換 dispatcher,阻塞式資料庫 driver 仍可明確移到 Dispatchers.IO
// handler 裡面
get("/users") {
val users = withContext(Dispatchers.IO) {
// 資料庫查詢(阻塞 I/O)
userRepository.findAll()
}
ok(users)
}
Dispatchers.IO 讓阻塞 I/O 跑在另一組並行配額上,不佔用 Dispatchers.Default 的額度 (兩者共用同一個 thread pool,各自有上限,不是兩個獨立的池子)。這在高併發場景下會有明顯改善
第 25 篇做了基本的 graceful shutdown,addShutdownHook + server.stop(5),更完善的策略分三階段
class GracefulShutdown(
private val server: HttpServer,
private val drainTimeMs: Long = 5_000,
) {
@Volatile
private var acceptingRequests = true
// 掛在 pipeline 最外層。drain 開始後,新進來的 request 一律回 503
fun drainMiddleware(): RelixMiddleware = { next ->
if (acceptingRequests) {
next()
} else {
RelixResponse(
503,
mapOf(
"Content-Type" to listOf("text/plain; charset=utf-8"),
"Connection" to listOf("close"),
),
"Service Unavailable".toByteArray(),
)
}
}
// readiness probe 讀這個值,k8s 才知道要把 Pod 從 endpoint 移掉
fun isReady(): Boolean = acceptingRequests
fun shutdown() {
// Phase 1: Drain — readiness 轉 false,新 request 由 middleware 回 503
println("Phase 1: Draining new connections...")
acceptingRequests = false
// Phase 2: Timeout — 等待進行中的 request 完成
println("Phase 2: Waiting for in-flight requests (${drainTimeMs}ms)...")
Thread.sleep(drainTimeMs)
// Phase 3: Force — 強制關閉
println("Phase 3: Force stopping...")
server.stop(0)
println("Shutdown complete")
}
}
acceptingRequests 標成 @Volatile,因為 shutdown() 跑在 shutdown hook 的 thread 上,讀它的 middleware 跑在 executor 的 thread 上。沒有 @Volatile 的話,工作 thread 可能一直看到快取的舊值
三個階段
drainMiddleware() 對新 request 回 503,不要用 removeContext() 代替,它只是把 context 拿掉,client 收到的會是 JDK HttpServer 預設的 404,不是 503,負載平衡器沒辦法從中判斷「這台正在下線」在 Kubernetes 環境下,drain phase 很重要,Pod 收到 SIGTERM 後,k8s 會從 service endpoint 移除這個 Pod,但已經路由過來的 request 還需要處理完,drain phase 確保這些 request 不會被砍掉
| 功能 | Relix | Ktor |
|---|---|---|
| HTTP Server | JDK HttpServer | Netty/CIO/Jetty/Tomcat |
| Routing | DSL + path params | DSL + path params + type-safe |
| Pipeline | Middleware chain | Phase-based pipeline |
| Content Negotiation | kotlinx.serialization | 多格式 (Jackson/Gson/kotlinx) |
| Authentication | Bearer | Bearer/JWT/OAuth/Session/多策略 |
| Validation | 自製 DSL | RequestValidation plugin + required parameter helpers |
| Static Files | 基本 | 完整 (range request/caching) |
| WebSocket | 無 | 有 |
| HTTP Client | 無 | 有 |
| HTTP/2 | 無 | 有 (Netty engine) |
| Testing | relixTest DSL | testApplication DSL |
Relix 刻意不做的東西
Relix 跟 Ktor 取捨不同的地方
RequestValidation,但偏向 request body 驗證 hookbenchmark 用 HttpURLConnection 準嗎 ?
基本上不夠準,HttpURLConnection 本身的開銷 (DNS lookup、TCP handshake、HTTP parsing) 會混進量測結果,更精確的做法是用專門的 HTTP benchmark 工具 (wrk、hey、vegeta),它們會重複使用連線、控制並行數。前面用 HttpURLConnection 是為了「零依賴、一個檔案就能跑」,重點在示範 benchmark 的結構 (warm up → measure → report),不在絕對精確的數字
什麼時候該用成熟框架 ?
要快速上線、要完整 feature、要穩定生態 → 直接用 Ktor 或 Spring,要深入理解框架設計、要客製自己的抽象、要做研究或教學 → Relix 很適合,「造輪子」的價值不在於取代成熟框架,而在於讓你做更好的工程判斷
Relix 的 JDK HttpServer adapter 適合研究和理解,但實際瓶頸仍要用明確的 executor 設定、負載工具與 profiler 驗證,循序 smoke test 不能當吞吐量 benchmark,可重現的量測要固定環境、並行度與原始輸出,後續方向包括替換 engine、隔離阻塞 I/O,以及用 readiness 配合 graceful shutdown,Relix 的定位是理解框架設計,不是用未驗證的數字取代成熟框架
最後一篇用心得收尾,聊聊哪一篇真的寫到卡住、手刻完再回頭翻 Ktor 原始碼是什麼感受,還有 32 篇下來回頭看 TDD 呈現方式的取捨
同步刊登於 Blog
圖片來源:AI 產生