이용자 요청
LiteLLM → Gateway → KServe Service.
TLS·streaming·timeout·connection cancel이 핵심이다.
GPU가 빨라도 model이 늦게 도착하거나 요청 길이 막히면 비싼 장비가 기다린다
cold start콜드 스타트local cache로컬 캐시RDMARemote Direct Memory Accesseast-west동서 트래픽왼쪽 두 줄은 서버를 준비하는 길이고, 아래 한 줄은 준비된 서버를 호출하는 길이다. model download가 느리면 Ready가 늦고, Gateway가 stream을 buffer하면 Ready인 서버도 사용자에게 느리게 보인다.
Ready 시간 = image pull + model fetch + model load + kernel compile + warm-up · health stabilization| 줄일 대상 | 수단 | 함정 |
|---|---|---|
| image pull | 사내 registry·node pre-pull | ARM64·amd64 image를 같은 tag로 덮음 |
| model fetch | object mirror·local cache | mutable model name·checksum 누락 |
| model load | quantization·memory headroom | load 성공 뒤 실제 traffic OOM을 놓침 |
| compile | persistent compile cache | runtime·driver 변경 뒤 stale cache |
| readiness | 실제 warm-up request | 단순 TCP open을 Ready로 오인 |
Pod Ready 12분 하나만 기록하면 registry가 느린지 model load가 느린지 알 수 없다. 각 단계를 event·metric으로
분리하고 image digest·model revision을 함께 남긴다.
KServe의 LocalModelCache 지원 범위는 release와 serving API에 따라 달라질 수 있다. 선택한 version에서
Standard InferenceService와 LLMInferenceService의 지원을 따로 확인한다. 맞지 않으면 generic prefetch
DaemonSet이나 node image를 이용한다.
이용자 요청
LiteLLM → Gateway → KServe Service.
TLS·streaming·timeout·connection cancel이 핵심이다.
다중 노드 추론
한 model replica의 leader·worker와 GPU 사이.
NIC·MTU·RDMA·topology·collective 통신이 핵심이다.
artifact 입출력
registry·model store → node cache.
throughput·checksum·cache hit·장애 격리가 핵심이다.
RDMA와 Network Operator는 모든 serving의 필수 구성요소가 아니다. single-node Standard InferenceService는
먼저 일반 Pod network로 합격시킨다. model이 한 node에 들어가지 않아 LeaderWorkerSet 다중 노드 추론을
쓸 때 NIC driver·RDMA device·network attachment와 topology를 별도 기준선으로 올린다.
두 node의 GPU 개수만 맞아도 통신 topology·NIC·MTU·RDMA resource가 다르면 시작 실패나 낮은 처리량이 난다. 그래서 multi-node runtime은 GPU resource뿐 아니라 network label·resource와 model cache 준비 상태를 함께 계약해야 한다.
| 계층 | 맡길 것 |
|---|---|
| 사내 Gateway API | TLS·host·path·기본 network policy |
| KServe route | model service·revision·inference pool 연결 |
| LiteLLM | 최종 이용자 key·quota·alias·provider fallback |
LLMInferenceService의 고급 router는 Envoy Gateway·Envoy AI Gateway·Gateway API Inference Extension을
사용한다. 사내 기본 Gateway가 다른 구현이라면 별도 GatewayClass로 격리하거나 Standard path를 유지한다.
기존 입구를 무심코 교체하지 않는다.
request ID → LiteLLM model alias → InferenceService revision → Pod → node → GPU UUID · MIG instance최소 신호는 다음과 같다.
GPU utilization이 낮으면서 vLLM queue가 길다면 model server CPU·network·process health를 먼저 본다. replica가 Ready가 되기까지 오래 걸리면 image·model·load·compile 단계로 나눈다.