iT邦幫忙

2026 iThome 鐵人賽

DAY 12
0
Kubernetes

凌晨四點,女友帶著 GPU 來我家學習 Kubernetes:打造 K8s AI Infra 的 30 夜系列 第 12 篇

【Day 12】KubeRay:當 Kubernetes 轉角遇見 Ray

  • 分享至 

  • xImage
  •  

過去幾天都在講 Kubernetes 怎麼把 GPU 交到 Pod 手上。

一部分是怎麼描述和切分:device plugin 讓 Kubernetes 認得 GPU,但只能整張給;DRA 讓你說得出要哪一張、要多少;HAMi 讓那個數字在容器裡真的被執行。

另一部分是卡不夠的時候誰先拿:Kueue 擋在門口管額度,Volcano 保證一群 Pod 一起上,KAI 用可搶佔性換超額使用,而 Kubernetes 自己也開始把 gang 做成內建 API。

這些都是把資源交出去之前的事。今天換個角度,看拿到資源之後,上面跑的東西是什麼。

訓練和推論一個大模型,本來就不是一台機器的事:參數要切到幾十張卡上、資料要並行處理、推論要同時服務大量請求。Ray 就是專門處理這類工作的框架,把一份 Python 程式碼攤到一整個叢集上跑。

今天的主角是 KubeRay,它把 Ray 這套分散式運算框架搬到 Kubernetes 上,變成 kubectl 管得動的物件。


Ray 是什麼

Ray 是一個把 AI 和 Python 應用從一台機器擴展到一整個叢集的框架,組成是一個分散式執行期,加上一組 AI 函式庫。

五個函式庫各自對應一種工作:

函式庫 做什麼
Data 可擴展的 ML 資料集
Train 分散式訓練
Tune 可擴展的超參數搜尋
RLlib 可擴展的強化學習
Serve 可擴展且可程式化的服務

底層的 Ray Core 提供三個原語:Task(無狀態的遠端函式)、Actor(有狀態的遠端物件)、Object Reference(遠端呼叫回傳的 future)。

用起來就是在函式上加一行裝飾器:

import ray
ray.init()

@ray.remote
def f(x):
    return x * x

futures = [f.remote(i) for i in range(4)]
print(ray.get(futures))   # [0, 1, 4, 9]

.remote() 一呼叫,這個工作就被丟進 Ray 的叢集,由 Ray 自己決定放到哪個節點執行。賣點是同一份程式碼不用改,筆電上怎麼跑,叢集上就怎麼跑。

有一件事跟今天的主題直接相關:Ray 有自己的一套資源系統。 @ray.remote(num_gpus=1) 這個數字是講給 Ray 的排程器聽的,跟 Kubernetes 的 nvidia.com/gpu 是兩回事。等一下的實驗會看到它們並排。


KubeRay 是做什麼的

Ray 的叢集有 head 和 worker。在 Kubernetes 上,這些節點就是 Pod,而 KubeRay 是一個 Kubernetes operator,負責把它們的部署和生命週期管起來。

它提供三個 CRD:

CRD 做什麼
RayCluster 完整管理一個 Ray 叢集的生命週期,包含建立、刪除、autoscaling 和容錯
RayJob 建立一個 RayCluster,等叢集就緒後把 job 送進去;可以設定跑完就把叢集刪掉
RayService 由 RayCluster 加上 Ray Serve 的 deployment graph 組成,提供零停機升級與高可用

三者的關係是疊的:RayJob 和 RayService 底下都會生一個 RayCluster。今天只用最基本的 RayCluster。

把前面幾天和今天擺在一起,會看到兩層排程:

排的單位 資源怎麼宣告 誰在排
Kubernetes Pod nvidia.com/gpu: 1 kube-scheduler(或 Volcano、KAI)
Ray task / actor num_gpus=1 Ray 自己

Kubernetes 把 Ray 的 head 和 worker 排成 Pod,工作就結束了。接下來那些 task 怎麼分到這些 Pod 上,是 Ray 在決定的。


實驗環境

沿用 fake-gpu-operator 的叢集,gpu-lab-worker 上有 8 張假卡。


裝 KubeRay operator

helm repo add kuberay https://ray-project.github.io/kuberay-helm/ && helm repo update
"kuberay" has been added to your repositories
Hang tight while we grab the latest from your chart repositories...
...Successfully got an update from the "kuberay" chart repository
Update Complete. ⎈Happy Helming!⎈
helm upgrade -i kuberay-operator kuberay/kuberay-operator \
  --version 1.7.1 -n kuberay --create-namespace --wait --timeout 5m
Release "kuberay-operator" does not exist. Installing it now.
NAME: kuberay-operator
LAST DEPLOYED: Fri Sep 25 17:24:09 2026
NAMESPACE: kuberay
STATUS: deployed
REVISION: 1
TEST SUITE: None
kubectl -n kuberay get pods
NAME                                READY   STATUS    RESTARTS   AGE
kuberay-operator-799d594664-5jddf   1/1     Running   0          37s

一個 Pod。 對照前幾天:Kueue 一個、Volcano 三個、KAI 七個、KubeRay 一個。它不是排程器,只是一個 operator。


實驗一:一份 yaml 長出一整個叢集

apiVersion: ray.io/v1
kind: RayCluster
metadata:
  name: demo
  namespace: kuberay-lab
spec:
  rayVersion: "2.52.0"
  headGroupSpec:
    rayStartParams: {}
    template:
      spec:
        containers:
          - name: ray-head
            image: rayproject/ray:2.52.0
            resources:
              limits:   { cpu: "2", memory: 4Gi }
              requests: { cpu: "2", memory: 4Gi }
  workerGroupSpecs:
    - groupName: workers
      replicas: 2
      rayStartParams: {}
      template:
        spec:
          containers:
            - name: ray-worker
              image: rayproject/ray:2.52.0
              resources:
                limits:   { cpu: "2", memory: 4Gi, nvidia.com/gpu: 1 }
                requests: { cpu: "2", memory: 4Gi, nvidia.com/gpu: 1 }

結構分成 headGroupSpec 和 workerGroupSpecs 兩塊,裡面各放一份標準的 Pod template。head 沒有要 GPU,worker 每個要一張。

kubectl create ns kuberay-lab
kubectl apply -f ~/kuberay-lab/raycluster.yaml
kubectl -n kuberay-lab get pods
NAME                        READY   STATUS    RESTARTS   AGE
demo-head-ds6ld             1/1     Running   0          74s
demo-workers-worker-qd892   1/1     Running   0          74s
demo-workers-worker-spmtj   1/1     Running   0          73s

一份 yaml,三個 Pod。 名字格式是 <cluster>-head-<亂碼> 和 <cluster>-<groupName>-worker-<亂碼>。

沒有 KubeRay 的話,這三個 Pod 要自己寫兩份 Deployment、一個 Service,還要手動把 worker 的啟動參數指向 head 的位址。

kubectl -n kuberay-lab get raycluster
NAME   DESIRED WORKERS   AVAILABLE WORKERS   CPUS   MEMORY   GPUS   STATUS   AGE
demo   2                 2                   6      12Gi     2      ready    100s

這一行是原生物件做不到的事。 CPUS 6 是 head 的 2 加上兩個 worker 的 2+2,MEMORY 12Gi、GPUS 2 同理。它把整個叢集的容量彙總成一列,而不是要你自己去加。

再看 GPU 落在誰身上:

kubectl -n kuberay-lab get pods -o custom-columns='NAME:.metadata.name,GPU:.spec.containers[0].resources.limits.nvidia\.com/gpu,NODE:.spec.nodeName'
NAME                        GPU      NODE
demo-head-ds6ld             <none>   gpu-lab-worker
demo-workers-worker-qd892   1        gpu-lab-worker
demo-workers-worker-spmtj   1        gpu-lab-worker

head 那欄是 <none>,兩個 worker 各一張,三個 Pod 都落在同一個節點上(叢集只有那台有卡)。

還有一個 Service:

kubectl -n kuberay-lab get svc
NAME            TYPE        CLUSTER-IP   EXTERNAL-IP   PORT(S)                                         AGE
demo-head-svc   ClusterIP   None         <none>        10001/TCP,8265/TCP,6379/TCP,8080/TCP,8000/TCP   2m3s

CLUSTER-IP 是 None,也就是 headless service。五個 port 分別給 Ray client、dashboard、GCS、metrics 和 Serve 用。worker 就是靠這個名字找到 head 的。


實驗二:它真的是一個 Ray 叢集

Pod 起來不代表 Ray 認得彼此。進 head 問它自己:

HEAD=$(kubectl -n kuberay-lab get pod -l ray.io/node-type=head -o name)
kubectl -n kuberay-lab exec $HEAD -- ray status
======== Autoscaler status: 2026-09-25 10:27:46.350306 ========
Node status
---------------------------------------------------------------
Active:
 (no active nodes)
Idle:
 2 workers
 1 headgroup
Pending:
 (no pending nodes)
Recent failures:
 (no failures)

Resources
---------------------------------------------------------------
Total Usage:
 0.0/6.0 CPU
 0.0/2.0 GPU
 0B/12.00GiB memory
 0B/3.29GiB object_store_memory

(ray.io/node-type=head 是 KubeRay 貼在 Pod 上的 label,用來分辨 head 和 worker。)

Ray 看到的和 Kubernetes 看到的對得起來:6 CPU、2 GPU、12GiB,跟 kubectl get raycluster 那一行一模一樣。因為 Ray 的數字是從容器的 limits 推出來的。

但 Ray 的帳上多了一欄 Kubernetes 沒有的:object_store_memory 3.29GiB,Ray 的物件儲存。

兩套帳的欄位不一樣,這是它們第一次分岔。


實驗三:6 個 task,2 張卡

送六個各要一張 GPU 的 task 進去,每個睡 2 秒,看它們落在哪、拿到什麼:

kubectl -n kuberay-lab exec $HEAD -- python -c "
import ray, socket, os, time
ray.init(address='auto')
@ray.remote(num_gpus=1)
def who():
    time.sleep(2)
    return socket.gethostname(), os.environ.get('CUDA_VISIBLE_DEVICES')
t = time.time()
for h, c in ray.get([who.remote() for _ in range(6)]):
    print(h, '| CUDA_VISIBLE_DEVICES =', c)
print('總共 %.1f 秒' % (time.time() - t))
"
2026-09-25 10:28:18,733 INFO worker.py:2014 -- Connected to Ray cluster. View the dashboard at 10.244.1.26:8265
(這裡省略一段關於 num_gpus=0 的 FutureWarning)
(raylet) There are tasks with infeasible resource requests that cannot be scheduled. See https://docs.ray.io/en/latest/ray-core/scheduling/index.html#ray-scheduling-resources for more details. Possible solutions: 1. Updating the ray cluster to include nodes with all required resources 2. To cause the tasks with infeasible requests to raise an error instead of hanging, set the 'RAY_enable_infeasible_task_early_exit=true'. This feature will be turned on by default in a future release of Ray.
(autoscaler +5s) Tip: use `ray status` to view detailed cluster status. To disable these messages, set RAY_SCHEDULER_EVENTS=0.
(autoscaler +5s) No available node types can fulfill resource requests {'CPU': 1.0, 'GPU': 1.0}*2. Add suitable node types to this cluster to resolve this issue.
demo-workers-worker-spmtj | CUDA_VISIBLE_DEVICES = 0
demo-workers-worker-qd892 | CUDA_VISIBLE_DEVICES = 0
demo-workers-worker-spmtj | CUDA_VISIBLE_DEVICES = 0
demo-workers-worker-qd892 | CUDA_VISIBLE_DEVICES = 0
demo-workers-worker-spmtj | CUDA_VISIBLE_DEVICES = 0
demo-workers-worker-qd892 | CUDA_VISIBLE_DEVICES = 0
總共 7.0 秒

六個 task 要六張卡,Ray 手上只有兩張,但沒有任何一個失敗。

總共 7 秒。每個 task 睡 2 秒,2 × 3 = 6 秒,加上啟動開銷:兩張卡各跑了三輪。

Kubernetes 在這段時間裡什麼都沒做。它早在三個 Pod 落到節點上的那一刻就完成任務了,這六個 task 要分幾輪、誰先誰後,是 Ray 在決定的。

兩層對「資源不夠」的反應也不一樣:Kubernetes 把 Pod 停在 Pending 等資源,Ray 則讓 task 排隊等前一個跑完。

節點上另外 6 張卡,Ray 碰不到

節點上有 8 張,這個 Ray 叢集只看得到 2 張,因為只開了兩個 worker Pod,每個要一張。其餘 6 張不在任何一個 worker 容器裡,所以 ray status 的帳就是 0.0/2.0 GPU。

輸出裡那行 autoscaler 訊息就是在講這件事:No available node types can fulfill resource requests {'CPU': 1.0, 'GPU': 1.0}*2。Ray 知道自己還有兩個 task 放不下,但它能做的只有開口要新節點,給不給是 Kubernetes 的事。這次沒開 autoscaling,所以它只能等前面的 task 跑完。

(raylet 那行用的字是 infeasible,但六個 task 最後都跑完了,所以這裡指的是當下排不下,不是永遠不可能。)

另外兩個觀察

六筆結果平均落在兩個 worker 上,head 一個都沒拿到,因為它沒有宣告 GPU。

CUDA_VISIBLE_DEVICES 全部都是 0。Ray 會替每個拿到 GPU 的 task 設定這個變數,但每個 worker 容器裡只有一張卡,所以容器內的索引永遠是 0。這是前面幾天那條線的延續:容器看到的裝置編號,和實體機器上的編號不是同一件事。


接上排程器

KubeRay 可以整合前面幾天看過的排程器,Volcano 和 KAI 都支援,而且兩邊的接法形狀幾乎一樣:

Volcano KAI
裝 operator 時 --set batchScheduler.name=volcano --set batchScheduler.name=kai-scheduler
RayCluster 用什麼入群 volcano.sh/queue-name label kai.scheduler/queue label

這個開關裝在 operator 上,不是裝在每個 RayCluster 上,跟前幾天「Pod 自己寫 schedulerName」是不同的形狀。

要解的問題兩邊一樣:head 和 worker 是一組的。一個 RayCluster 如果 head 要 1 張卡、四個 worker 各要 0.5 張,整組就是 3 張,五個 Pod 要嘛一起上、要嘛一個都不上。Volcano 的行為也是如此,當一個叢集要的資源超過佇列容量,所有 Pod 一起 pending,而不是先排一部分上去。

KAI 還支援讓兩個 worker 共用一張卡、各 0.5,用 time-slicing 做。但這個 time-slicing 只發生在 Kubernetes 這一層:它讓多個 Pod 落在同一張卡上,卻不強制記憶體隔離,要靠應用程式自己管好用量。

排程器只負責把兩個 Pod 放到同一張卡上,不管誰用了多少記憶體。這跟前面幾天的結論是同一句話。


小結

KubeRay 做到的是:一份 yaml 描述一整個 Ray 叢集,head、worker、service 和它們之間的連線都由 operator 產生和維持,kubectl get raycluster 一行就能看到整個叢集的容量和狀態。

而今天真正的重點是那六個 task:Kubernetes 排了 3 個 Pod 就下班了,Ray 在那些 Pod 裡面又排了三輪。 兩層排程、兩套資源記帳,上面那層只看得到 nvidia.com/gpu: 1,下面那層才知道有六個 task 在排隊。

而接上 Volcano 或 KAI 之後會變成三層:排程器決定幾個 Pod 能不能一起上,Kubernetes 決定放在哪,Ray 決定 task 怎麼輪。


參考資料

ray-project/ray
Ray — Getting Started
Ray Core — Key Concepts
Ray Core — Resources
Ray Core — Accelerator Support
ray-project/kuberay
Ray on Kubernetes
Getting Started with KubeRay
RayCluster Quick Start
RayCluster Configuration
KubeRay Ecosystem
KubeRay integration with Volcano
Gang scheduling, queue priority, and GPU sharing for RayClusters using KAI Scheduler
Gang scheduling, Priority scheduling, and Autoscaling for KubeRay CRDs with Kueue
KubeRay integration with scheduler plugins


上一篇
【Day 11】原生 gang scheduling:Kubernetes 自己的答案
下一篇
【Day 13】Slinky:當 HPC 的老將 Slurm 走進 Kubernetes
系列文
凌晨四點,女友帶著 GPU 來我家學習 Kubernetes:打造 K8s AI Infra 的 30 夜 共 17 篇
圖片
  熱門推薦
圖片
{{ item.channelVendor }} | {{ item.webinarstarted }} |
{{ formatDate(item.duration) }}
直播中

尚未有邦友留言

立即登入留言