iT邦幫忙

2026 iThome 鐵人賽

0
Software Development

Kotlin 手刻 Ktor 從零開始系列 第 31 篇

Kotlin 手刻 Ktor 從零開始 Day 31 效能與改進方向,Relix 能走多遠 ?

  • 分享至 

  • xImage
  •  

https://ithelp.ithome.com.tw/upload/images/20260822/20121948NIqCtgaIUJ.png

Relix 已經具備一個小型 API 框架該有的東西,但工程上一定會被問,效能如何 ?

這篇不是要把 Relix 變成世界最快的框架,目標是用工程化的方式回答四個問題,怎麼量測、瓶頸在哪、接下來可以怎麼改、什麼時候該直接用成熟框架

量測方法,先做 latency smoke test

下面的 HttpURLConnection 程式是循序送 request,只適合確認暖機後的單一路徑 latency 是否出現明顯退化,它沒有並行控制,也混入 client 與連線成本,不能拿來宣稱 server 吞吐量或高併發 p99

import java.net.HttpURLConnection
import java.net.URI

fun latencySmokeTest(
    url: String,
    warmup: Int = 1000,
    iterations: Int = 10_000,
) {
    // Warm up:讓 JVM JIT 編譯器暖機
    repeat(warmup) {
        val conn = URI(url).toURL().openConnection() as HttpURLConnection
        conn.inputStream.readBytes()
        conn.disconnect()
    }

    // 量測
    val latencies = mutableListOf<Long>()
    val start = System.nanoTime()

    repeat(iterations) {
        val reqStart = System.nanoTime()
        val conn = URI(url).toURL().openConnection() as HttpURLConnection
        conn.inputStream.readBytes()
        conn.disconnect()
        latencies.add(System.nanoTime() - reqStart)
    }

    val totalMs = (System.nanoTime() - start) / 1_000_000
    latencies.sort()
    val p50 = latencies[latencies.size / 2] / 1_000_000.0
    val p95 = latencies[(latencies.size * 0.95).toInt()] / 1_000_000.0
    val p99 = latencies[(latencies.size * 0.99).toInt()] / 1_000_000.0

    println("=== Latency Smoke Test ===")
    println("Requests:  $iterations")
    println("Total:     ${totalMs}ms")
    println("p50:       ${"%.2f".format(p50)}ms")
    println("p95:       ${"%.2f".format(p95)}ms")
    println("p99:       ${"%.2f".format(p99)}ms")
}

URI(url).toURL() 不用 URL(url),是因為 JDK 20 之後那個建構式被標為 deprecated,照舊寫法編譯會噴兩個 warning

使用方式

fun main() {
    val app = Relix { port = 8080 }
    app.routing {
        get("/hello") { ok("Hello, World!") }
    }
    app.start()

    latencySmokeTest("http://localhost:8080/hello")

    app.stop()
}

跑起來長這樣

Relix started on 0.0.0.0:8080
=== Latency Smoke Test ===
Requests:  10000
Total:     1338ms
p50:       0.12ms
p95:       0.20ms
p99:       0.35ms
Relix shutting down...
Relix stopped

這是單機跑一次的結果,循序送、沒有並行,只能用來確認暖機後單一路徑沒有明顯退化,不是 benchmark,換一台機器、換一個 JDK 版本、背景多開幾個程式,數字都會不一樣,也不能拿它推論 Relix 的吞吐量或高併發下的 p99

Relix stopped 印兩次不是 bug,stop() 有兩個呼叫來源,一個是 main() 裡的 app.stop(),一個是 start() 註冊的 shutdown hook,第 25 篇就講過這個坑,這裡沒有擋,它就老實地跑第二次,onStopping 的 hook 也會跟著跑兩次,這個範例沒註冊 hook 所以無所謂,真的有連線池要關,就照第 25 篇說的加一個 AtomicBoolean 讓 stop() 只生效一次

port 這裡跟著 application.properties 寫 8080,不要自己改成別的數字,第 25 篇定出來的優先順序是 DSL → 設定檔 → 環境變數,後者蓋前者,Relix { port = 9999 } 會被設定檔裡的 port=8080 蓋掉,server 起在 8080,smoke test 卻打 9999,跑起來就是一行 Relix started on 0.0.0.0:8080 之後接一個 java.net.ConnectException: Connection refused,真的要換 port,改 application.properties 或設 RELIX_PORT,兩邊的數字要一致

還有一點,main() 裡例外飛出來的時候 app.stop() 不會執行,JDK HttpServer 的 thread 不是 daemon thread,JVM 不會跟著結束,8080 會一直被佔住,下一次跑就變成 Address already in use,遇到的話先確認上一輪的 process 還在不在

為什麼要 warm up ? JVM 剛啟動時用 interpreter 跑位元碼,效能差,跑幾千次之後 JIT 編譯器會把 hot path 編譯成 native code,warm up 讓量測結果反映的是 JIT 後的效能,而不是 interpreter 的效能

為什麼記錄 percentile 而不是平均值 ? 平均值會被極端值拉偏,p95 告訴你「95% 的 request 在幾毫秒以內完成」,p99 告訴你 tail latency,對 API server 來說,p99 比平均值重要得多,一百個 request 裡有一個慢 2 秒,使用者體驗就差了

可重現的負載測試

要量 throughput 與 tail latency,請用 wrk、hey 或 vegeta 之類能控制連線與並行數的工具,例如固定 30 秒、20 條連線,分別測空 handler 與完整 middleware 組合,每次記錄 CPU 型號、JDK 版本、heap、executor、工具版本、命令、commit 與原始輸出

wrk -t4 -c20 -d30s http://127.0.0.1:8080/hello

先做多次 warm-up,再至少跑三輪正式測試,比較 median 與分布,不只挑最好的一次,這是我這台機器上跑的三輪,打的是前面那個 /hello 空 handler,沒有裝任何 middleware

Running 30s test @ http://127.0.0.1:8080/hello
  4 threads and 20 connections
  Thread Stats   Avg      Stdev     Max   +/- Stdev
    Latency   468.19us    1.13ms  75.38ms   99.47%
    Req/Sec    11.75k     1.31k   13.28k    92.28%
  1408312 requests in 30.10s, 174.60MB read
Requests/sec:  46787.28
    Latency   425.97us  193.36us  12.16ms   73.47%
    Req/Sec    11.86k   769.81    13.13k    82.72%
  1421265 requests in 30.10s, 176.21MB read
Requests/sec:  47217.43
    Latency   424.85us  184.46us  10.71ms   71.02%
    Req/Sec    11.86k   686.05    12.99k    80.56%
  1421308 requests in 30.10s, 176.21MB read
Requests/sec:  47219.63

median 是 47217,三輪之間差不到 1%,這個部分算穩定,分布就沒那麼漂亮,第一輪的 Max 是 75.38ms、標準差 1.13ms,後兩輪的 Max 掉到 10 到 12ms、標準差不到 200us,第一輪明顯還在暖機,JIT 還沒把 hot path 編完,這就是為什麼只跑一輪、只看 Requests/sec 會騙自己,平均值三輪都差不多,尾巴差了六倍

環境要一起記下來,數字才有意義

  • Apple M2 Max,12 核,macOS 26.5.2
  • Temurin JDK 21.0.12+8,heap 沒調過,ergonomic 給到 16GB
  • wrk 4.2.0,wrk -t4 -c20 -d30s http://127.0.0.1:8080/hello
  • Relix 沒有明確設定 executor,用的是 JDK HttpServer 的預設
  • client 跟 server 在同一台機器上,兩邊搶同一顆 CPU

換句話說,這組數字只在這個組合下成立,別拿去跟別人貼的 Ktor 或 Spring Boot 數字比,那些多半是不同機器、不同 JDK、不同 client 工具、甚至 client 跑在另一台機器上量出來的,要比就自己在同一台機器上把兩邊都跑一次

還有一件事這三輪沒有做到,前面說「分別測空 handler 與完整 middleware 組合」,上面只有空 handler 那一半,所以這裡不推論 middleware 的固定成本,真要知道 CORS、StatusPages、Authentication 疊上去各要花多少,得照同樣的流程再測一組帶 middleware 的路由,兩組的差值才是那一層的成本

瓶頸分析,JDK HttpServer 的限制

Relix 用 com.sun.net.httpserver.HttpServer 當底層引擎,好處是零依賴、API 簡單,限制也很明顯

Executor 決定並行模型

JDK HttpServer 用 Executor 分派 request。若沒有明確設定,會使用實作提供的預設 executor,不能直接描述成固定的 thread-per-request,應該可以顯式設定 executor,讓 thread 數、queue 與拒絕政策可觀察,handler 若執行阻塞 I/O,執行它的 thread 仍會被占住

Request → Thread Pool (200 threads) → Handler (blocking I/O) → 阻塞等待
                                        ↓
                                   thread 佔住 50ms
                                        ↓
                           200 concurrent requests 就滿了

流量控制能力有限

相關的 adapter 沒有定義 queue 上限、拒絕政策、request body streaming 或負載保護,實際行為取決於 executor 與底層實作,因此要靠 bounded queue、timeout、body limit、readiness 與外層 proxy 一起控制

HTTP/2 不支援

com.sun.net.httpserver 只支援 HTTP/1.1,如果你需要 multiplexing、server push、header 壓縮,得換底層

哪一層是瓶頸必須由 profiler 與分層 benchmark 證明,不要在沒有量測資料時先假設 router、pipeline 或底層 I/O 一定占多少時間

Profiling 觀察重點

如果你想用 profiler (VisualVM、async-profiler、IntelliJ Profiler) 分析 Relix 的效能,可以看這幾個地方

CPU hotspot

  • Router.match(),路由比對,路由少 (< 100) 時不會是瓶頸,路由多時可以考慮 trie 取代線性掃描
  • Pipeline.execute(),middleware 鏈的執行。middleware 數量通常是個位數,不會是問題
  • Json.encodeToString() / Json.decodeFromString(),JSON 序列化反序列化,大 payload 會吃 CPU

記憶體

  • ByteArray 分配,每個 response body 都建一個 ByteArray,高吞吐時 GC 壓力會上來
  • Map<String, List<String>> headers,header 的分配和複製。每個 request/response 一組

Thread 利用率

  • thread pool 的 active count vs total count
  • 如果 active count 長期等於 total count,表示 thread 都在忙 (或阻塞),需要加大 pool 或減少阻塞

改進方向,Netty 替換底層

當框架的上層抽象 (RelixCall、Router、Pipeline、Plugin) 穩定之後,換底層引擎是合理的下一步

目前
Request → JDK HttpServer → RelixRequest → RelixCall → Pipeline → Handler → RelixResponse

替換後
Request → Netty → RelixRequest → RelixCall → Pipeline → Handler → RelixResponse

只有最左邊那一段需要改,RelixRequest 以右的東西完全不動

需要實作的介面

interface RelixEngine {
    fun start(host: String, port: Int, handler: suspend (RelixRequest) -> RelixResponse)
    fun stop(gracePeriodMs: Long)
}

class JdkEngine : RelixEngine { /* 目前的實作 */ }
class NettyEngine : RelixEngine { /* Netty 版 */ }

把 engine 做成可替換的介面。應用層選擇用哪個 engine

val app = Relix {
    engine = NettyEngine()
    port = 8080
}

Netty 的好處

  • 非阻塞 I/O (event loop 模型),不需要 thread-per-request
  • 內建 HTTP/2 支援
  • 成熟的 backpressure 和 flow control
  • Ktor 的底層引擎之一就是 Netty

這裡不做 Netty 的實作,但可以知道介面邊界在哪,你就知道該怎麼切,實作骨架大概是這樣

class NettyEngine : RelixEngine {
    private lateinit var channel: Channel
    private val bossGroup = NioEventLoopGroup(1)            // 接 connection
    private val workerGroup = NioEventLoopGroup()           // 處理 I/O

    override fun start(host: String, port: Int, handler: suspend (RelixRequest) -> RelixResponse) {
        val bootstrap = ServerBootstrap()
            .group(bossGroup, workerGroup)
            .channel(NioServerSocketChannel::class.java)
            .childHandler(object : ChannelInitializer<SocketChannel>() {
                override fun initChannel(ch: SocketChannel) {
                    ch.pipeline()
                        .addLast(HttpServerCodec())                  // HTTP 編解碼
                        .addLast(HttpObjectAggregator(1 shl 20))     // 把分塊請求合成一個 FullHttpRequest
                        .addLast(RelixHandlerAdapter(handler))        // 橋接到 RelixRequest/RelixResponse
                }
            })
        channel = bootstrap.bind(host, port).sync().channel()
    }
    // stop() 略
}

event loop 能用少量 thread 管理大量連線,但不保證換成 Netty 就一定更快,handler 是否阻塞、buffer 策略、TLS、payload 與 executor 設定都會影響結果,是否替換 engine,應以相同 workload 的量測與維運需求決定

Coroutine Dispatcher 策略

目前 Relix 用 runBlocking 在 JDK HttpServer 的 thread 上執行 suspend handler,這可以動,但有改進空間

// 目前:直接在 HttpServer 的 thread 上 runBlocking
executor.execute {
    val response = runBlocking {
        pipeline.execute(call)
    }
    writeResponse(exchange, response)
}

把同一個 executor 轉成 dispatcher,再在 executor callback 裡 runBlocking(dispatcher) 只會多一次排程,並不會把 adapter 變成 non-blocking

// 不建議:同一批 thread 之間重複排程
val dispatcher = executor.asCoroutineDispatcher()

executor.execute {
    val response = runBlocking(dispatcher) {
        pipeline.execute(call)
    }
    writeResponse(exchange, response)
}

若保留同步 JDK adapter,就接受 callback thread 會等到 response 完成,並使用 bounded executor 控制資源,若要真正 non-blocking,engine API 必須以 coroutine 啟動 request、在完成時非阻塞地寫回 response,不能只替換 dispatcher,阻塞式資料庫 driver 仍可明確移到 Dispatchers.IO

// handler 裡面
get("/users") {
    val users = withContext(Dispatchers.IO) {
        // 資料庫查詢(阻塞 I/O)
        userRepository.findAll()
    }
    ok(users)
}

Dispatchers.IO 讓阻塞 I/O 跑在另一組並行配額上,不佔用 Dispatchers.Default 的額度 (兩者共用同一個 thread pool,各自有上限,不是兩個獨立的池子)。這在高併發場景下會有明顯改善

進階 Graceful Shutdown

第 25 篇做了基本的 graceful shutdown,addShutdownHook + server.stop(5),更完善的策略分三階段

class GracefulShutdown(
    private val server: HttpServer,
    private val drainTimeMs: Long = 5_000,
) {
    @Volatile
    private var acceptingRequests = true

    // 掛在 pipeline 最外層。drain 開始後,新進來的 request 一律回 503
    fun drainMiddleware(): RelixMiddleware = { next ->
        if (acceptingRequests) {
            next()
        } else {
            RelixResponse(
                503,
                mapOf(
                    "Content-Type" to listOf("text/plain; charset=utf-8"),
                    "Connection" to listOf("close"),
                ),
                "Service Unavailable".toByteArray(),
            )
        }
    }

    // readiness probe 讀這個值,k8s 才知道要把 Pod 從 endpoint 移掉
    fun isReady(): Boolean = acceptingRequests

    fun shutdown() {
        // Phase 1: Drain — readiness 轉 false,新 request 由 middleware 回 503
        println("Phase 1: Draining new connections...")
        acceptingRequests = false

        // Phase 2: Timeout — 等待進行中的 request 完成
        println("Phase 2: Waiting for in-flight requests (${drainTimeMs}ms)...")
        Thread.sleep(drainTimeMs)

        // Phase 3: Force — 強制關閉
        println("Phase 3: Force stopping...")
        server.stop(0)
        println("Shutdown complete")
    }
}

acceptingRequests 標成 @Volatile,因為 shutdown() 跑在 shutdown hook 的 thread 上,讀它的 middleware 跑在 executor 的 thread 上。沒有 @Volatile 的話,工作 thread 可能一直看到快取的舊值

三個階段

  1. Drain,先讓 readiness 失敗,並由最外層的 drainMiddleware() 對新 request 回 503,不要用 removeContext() 代替,它只是把 context 拿掉,client 收到的會是 JDK HttpServer 預設的 404,不是 503,負載平衡器沒辦法從中判斷「這台正在下線」
  2. Timeout,等一段時間讓進行中的 request 完成
  3. Force,超時就強制關閉

在 Kubernetes 環境下,drain phase 很重要,Pod 收到 SIGTERM 後,k8s 會從 service endpoint 移除這個 Pod,但已經路由過來的 request 還需要處理完,drain phase 確保這些 request 不會被砍掉

與 Ktor 的功能對照

功能 Relix Ktor
HTTP Server JDK HttpServer Netty/CIO/Jetty/Tomcat
Routing DSL + path params DSL + path params + type-safe
Pipeline Middleware chain Phase-based pipeline
Content Negotiation kotlinx.serialization 多格式 (Jackson/Gson/kotlinx)
Authentication Bearer Bearer/JWT/OAuth/Session/多策略
Validation 自製 DSL RequestValidation plugin + required parameter helpers
Static Files 基本 完整 (range request/caching)
WebSocket 無 有
HTTP Client 無 有
HTTP/2 無 有 (Netty engine)
Testing relixTest DSL testApplication DSL

Relix 刻意不做的東西

  • WebSocket,需要完全不同的 connection 管理模型
  • HTTP Client,這不是 server 框架的責任 (用 ktor-client 或 OkHttp)
  • 多 engine 支援,這裡只用簡單的 JDK HttpServer
  • Session management,這裡用簡單的 stateless token

Relix 跟 Ktor 取捨不同的地方

  • Relix 刻意說明 Validation DSL,Ktor 有 RequestValidation,但偏向 request body 驗證 hook
  • Relix 的 DI 容器刻意極簡,方便說明,Ktor 3.x 已有內建 DI plugin,實務上也常接 Koin / Dagger

常見陷阱與設計取捨

benchmark 用 HttpURLConnection 準嗎 ?

基本上不夠準,HttpURLConnection 本身的開銷 (DNS lookup、TCP handshake、HTTP parsing) 會混進量測結果,更精確的做法是用專門的 HTTP benchmark 工具 (wrk、hey、vegeta),它們會重複使用連線、控制並行數。前面用 HttpURLConnection 是為了「零依賴、一個檔案就能跑」,重點在示範 benchmark 的結構 (warm up → measure → report),不在絕對精確的數字

什麼時候該用成熟框架 ?

要快速上線、要完整 feature、要穩定生態 → 直接用 Ktor 或 Spring,要深入理解框架設計、要客製自己的抽象、要做研究或教學 → Relix 很適合,「造輪子」的價值不在於取代成熟框架,而在於讓你做更好的工程判斷


小結

Relix 的 JDK HttpServer adapter 適合研究和理解,但實際瓶頸仍要用明確的 executor 設定、負載工具與 profiler 驗證,循序 smoke test 不能當吞吐量 benchmark,可重現的量測要固定環境、並行度與原始輸出,後續方向包括替換 engine、隔離阻塞 I/O,以及用 readiness 配合 graceful shutdown,Relix 的定位是理解框架設計,不是用未驗證的數字取代成熟框架


下一篇

最後一篇用心得收尾,聊聊哪一篇真的寫到卡住、手刻完再回頭翻 Ktor 原始碼是什麼感受,還有 32 篇下來回頭看 TDD 呈現方式的取捨


參考資料


同步刊登於 Blog

圖片來源:AI 產生


上一篇
Kotlin 手刻 Ktor 從零開始 Day 30 綜合實戰 (下),加上驗證、錯誤處理與認證
下一篇
Kotlin 手刻 Ktor 從零開始 Day 32 旅程的終點
系列文
Kotlin 手刻 Ktor 從零開始 共 32 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言