콘텐츠로 이동
Study Note온프렘 GPU 플랫폼

8. 마무리

전체를 기억할 때는 노드를 준비하는 흐름, 모델을 배포하는 흐름, 요청이 지나가는 흐름 세 줄이면 된다

이 장의 사용법4개
전체 지도
GPU node 준비부터 API 요청까지 세 흐름을 한 그림에서 본다.
선택표
Standard InferenceService와 LLMInferenceService 중 어디에서 시작할지 고른다.
도입 체크리스트
Spark·A100·B300 serving을 열기 전 빠뜨리지 않을 gate다.
장애 대응 카드
endpoint·model Pod·GPU 기반 장애의 첫 확인 순서를 모았다.
GPU Operator가 노드를 준비하는 A, Git과 ServingRuntime이 KServe controller를 거쳐 모델 Pod를 배포하는 B, 사내 앱이 LiteLLM과 Gateway를 지나 그 Pod와 GPU로 닿는 C를 한데 모은 덱 전체 요약도
  • A가 깨지면 Pod가 GPU를 요청하지 못하거나 CUDA가 실패한다.
  • B가 깨지면 model replica·Service·Ready 상태가 원하는 대로 만들어지지 않는다.
  • C가 깨지면 모델 Pod가 건강해도 이용자 요청이 도달하지 않거나 stream이 중간에서 끊긴다.

이 세 줄을 섞지 않으면 장애 범위를 빨리 좁힐 수 있다.

질문GPU OperatorKServe
관리 대상GPU node software stackmodel serving workload
대표 원하는 상태ClusterPolicy·NVIDIADriver·MIG configInferenceService·ServingRuntime·LLMInferenceService
주로 만드는 것driver·toolkit·device plugin·GFD·DCGM PodDeployment·Service·route·model Pod
성공 결과GPU resource·label·metric이 node에 보임model endpoint와 Ready 상태가 유지됨
하지 않는 일model replica·API route 관리driver 설치·GPU 자동 추천
첫 장애 확인validator·Allocatable·CDI·DCGMInferenceService condition·Deployment·Pod event

GPU Operator가 아래에서 실행 가능성을 만들고, KServe가 위에서 model server의 생명주기를 관리한다. 둘은 경쟁 제품이 아니라 서로 다른 층이다.

새 model serving이 한 Pod·한 node에 들어가면 Standard InferenceService로 가고 아니면 LeaderWorkerSet을 검토하며, Standard가 실제 부하에서 SLO를 못 맞추면 병목 종류에 따라 Inference Extension이나 prefill·decode 분리로 가는 판단 흐름
serving 요구기본 객체·방식GPU 단위
Spark single-node LLMStandard InferenceServiceSpark node 1대
A100·B300 single-node LLMStandard InferenceServicefull GPU 1·2·4·8 preset
embedding·rerankerStandard InferenceService고정 MIG 또는 full GPU profile
custom FastAPI modelcustom predictorCPU·MIG·full GPU
model이 한 node에 안 들어감LLMInferenceService + LWS고정 topology·RDMA 검증 pool
prefix·load-aware routingLLMInferenceService + Inference Extension여러 replica pool
prefill이 decode를 방해LLMInferenceService P/D 분리prefill·decode 전용 pool
장한 문장
0요청의 길과 운영의 길을 나누면 제품 역할이 보인다
1Spark·A100·B300은 같은 GPU 수가 아니라 서로 다른 자원 섬이다
2GPU Operator는 driver부터 resource·label·metric까지 node contract를 만든다
3KServe는 상위 model 선언을 하위 Kubernetes 객체로 바꾸고 계속 맞춘다
4LiteLLM은 이용자를, KServe는 model replica와 endpoint를 안다
5GPU 사용률보다 queue·TTFT·ITL·KV cache와 이용자 SLO를 먼저 본다
6image·model·Gateway·다중 노드 연결이 빠른 GPU를 굶길 수 있다
7endpoint와 node ownership을 나눠 Spark부터 검증하고 A100은 한 대씩 옮긴다
  1. scope — 상시 inference만 대상이며 batch·학습·Notebook은 별도 범위임을 합의했다.
  2. cluster 경계 — Spark와 datacenter GPU를 한 운영 모델·독립 가능한 실행 장애 영역으로 정했다.
  3. node contract — architecture·GPU family·partition·network label과 taint를 고정했다.
  4. GPU 기반 — driver·toolkit·CDI·device plugin·GFD·MIG·DCGM을 실제 CUDA Pod로 검증했다.
  5. runtime catalog — Spark·A100·B300용 image digest와 model·CUDA·driver matrix를 만들었다.
  6. Standard serving — vLLM·embedding·FastAPI를 InferenceService와 LiteLLM으로 연결했다.
  7. serving 안전선 — 실제 prompt·output 분포로 SLO·동시성·queue·timeout budget을 정했다.
  8. data·network — cold start budget·cache·streaming·Gateway·필요 시 RDMA 기준을 측정했다.
  9. 고급 KServe gate — Standard로 풀 수 없는 병목이 측정될 때만 LLMInferenceService 의존성을 설치한다.
  10. rollback — model endpoint와 A100 node 하나를 실제로 기존 환경으로 되돌렸다.
  11. B300 준비 — B300 같은 새 서버를 추가한다면 반입 전에 driver·image·full-node·fabric·burn-in 기준을 승인했다.

모든 inference가 안 된다

LiteLLM과 Gateway health를 먼저 나눈다.

여러 Service endpoint·DNS·TLS와 model Pod Ready를 비교한다.

모델 하나만 안 된다

InferenceService condition → Deployment → Pod event 순서로 내려간다.

image architecture·model fetch·readiness·OOM을 본다.

Pod가 Pending이다

GPU request와 node Allocatable을 비교한다.

label·taint·affinity·MIG resource 이름을 확인한다.

Pod는 떠도 GPU가 안 보인다

toolkit·CDI → device plugin Allocate → container CUDA 순서로 본다.

GPU Operator validator와 containerd event를 확인한다.

느려지고 timeout이 난다

waiting·queue time·TTFT·ITL을 먼저 맞춘다.

KV cache·preemption·Gateway timeout과 request cancel을 본다.

GPU 오류가 반복된다

DCGM XID·ECC·temperature·link와 Pod restart 시점을 맞춘다.

route에서 제외하고 event를 보존한 뒤 node를 drain한다.

첫 구현은 단순하다. GPU Operator로 node contract를 만들고, Standard InferenceService로 single-node serving을 안정화하고, LiteLLM에 그 endpoint를 연결한다. 여기까지가 기본 경로다.

그 다음은 기능 목록이 아니라 측정에서 출발한다. model이 한 node에 들어가지 않거나 prefix·부하를 고려한 routing, prefill/decode 분리가 실제 SLO를 개선할 때만 LLMInferenceService로 확장한다. 고급 기능을 쓰지 않는 것이 미완성이 아니라, 필요한 복잡도만 운영하는 것이다.