콘텐츠로 이동
Study Note온프렘 GPU 플랫폼

6. GPU를 굶기지 않는 데이터와 연결

GPU가 빨라도 model이 늦게 도착하거나 요청 길이 막히면 비싼 장비가 기다린다

이 장에서 처음 나오는 말4개
cold start콜드 스타트
빈 node에서 image·model을 받고 초기화·compile·warm-up해 Ready가 되는 시간이다.
local cache로컬 캐시
다시 받을 수 있는 model·compile 결과를 node NVMe에 보관해 다음 시작을 줄이는 층이다.
RDMARemote Direct Memory Access
CPU copy 부담을 줄여 node 사이 memory 전송 지연을 낮추는 network 방식이다.
east-west동서 트래픽
cluster 안의 model Pod·cache·다중 노드 worker 사이를 흐르는 통신이다.

추론 Pod에는 두 가지가 들어온다

섹션 제목: “추론 Pod에는 두 가지가 들어온다”
사내 image registry와 model object store가 각각 node image cache와 node model cache를 거쳐 vLLM·FastAPI Pod에 닿고, LiteLLM·Gateway의 API 요청이 그 Pod로 들어가며 Pod가 GPU와 관측 신호로 이어지는 데이터 경로

왼쪽 두 줄은 서버를 준비하는 길이고, 아래 한 줄은 준비된 서버를 호출하는 길이다. model download가 느리면 Ready가 늦고, Gateway가 stream을 buffer하면 Ready인 서버도 사용자에게 느리게 보인다.

Ready 시간 = image pull + model fetch + model load + kernel compile + warm-up · health stabilization
줄일 대상수단함정
image pull사내 registry·node pre-pullARM64·amd64 image를 같은 tag로 덮음
model fetchobject mirror·local cachemutable model name·checksum 누락
model loadquantization·memory headroomload 성공 뒤 실제 traffic OOM을 놓침
compilepersistent compile cacheruntime·driver 변경 뒤 stale cache
readiness실제 warm-up request단순 TCP open을 Ready로 오인

Pod Ready 12분 하나만 기록하면 registry가 느린지 model load가 느린지 알 수 없다. 각 단계를 event·metric으로 분리하고 image digest·model revision을 함께 남긴다.

KServe의 LocalModelCache 지원 범위는 release와 serving API에 따라 달라질 수 있다. 선택한 version에서 Standard InferenceService와 LLMInferenceService의 지원을 따로 확인한다. 맞지 않으면 generic prefetch DaemonSet이나 node image를 이용한다.

이용자 요청

LiteLLM → Gateway → KServe Service.

TLS·streaming·timeout·connection cancel이 핵심이다.

다중 노드 추론

한 model replica의 leader·worker와 GPU 사이.

NIC·MTU·RDMA·topology·collective 통신이 핵심이다.

artifact 입출력

registry·model store → node cache.

throughput·checksum·cache hit·장애 격리가 핵심이다.

RDMA와 Network Operator는 모든 serving의 필수 구성요소가 아니다. single-node Standard InferenceService는 먼저 일반 Pod network로 합격시킨다. model이 한 node에 들어가지 않아 LeaderWorkerSet 다중 노드 추론을 쓸 때 NIC driver·RDMA device·network attachment와 topology를 별도 기준선으로 올린다.

다중 노드 추론은 node 수보다 연결이 중요하다

섹션 제목: “다중 노드 추론은 node 수보다 연결이 중요하다”
Inference route가 leader Pod로 요청을 보내고 leader와 worker Pod가 NCCL·RDMA로 오가며 각각 노드 A와 노드 B의 GPU group을 쓰고, model cache가 두 Pod에 모두 연결된 다중 노드 추론 구조

두 node의 GPU 개수만 맞아도 통신 topology·NIC·MTU·RDMA resource가 다르면 시작 실패나 낮은 처리량이 난다. 그래서 multi-node runtime은 GPU resource뿐 아니라 network label·resource와 model cache 준비 상태를 함께 계약해야 한다.

계층맡길 것
사내 Gateway APITLS·host·path·기본 network policy
KServe routemodel service·revision·inference pool 연결
LiteLLM최종 이용자 key·quota·alias·provider fallback

LLMInferenceService의 고급 router는 Envoy Gateway·Envoy AI Gateway·Gateway API Inference Extension을 사용한다. 사내 기본 Gateway가 다른 구현이라면 별도 GatewayClass로 격리하거나 Standard path를 유지한다. 기존 입구를 무심코 교체하지 않는다.

관측은 한 요청과 한 replica를 연결한다

섹션 제목: “관측은 한 요청과 한 replica를 연결한다”
request ID → LiteLLM model alias → InferenceService revision → Pod → node → GPU UUID · MIG instance

최소 신호는 다음과 같다.

  • LiteLLM request rate·error·latency·token usage
  • Gateway upstream error·stream duration·disconnect
  • KServe desired·ready replica와 revision
  • vLLM TTFT·ITL·queue·KV cache·input/output tokens
  • Pod restart·OOM·node pressure
  • DCGM memory·activity·temperature·XID·ECC·link
  • registry·model fetch throughput·cache hit·checksum failure

GPU utilization이 낮으면서 vLLM queue가 길다면 model server CPU·network·process health를 먼저 본다. replica가 Ready가 되기까지 오래 걸리면 image·model·load·compile 단계로 나눈다.