콘텐츠로 이동
Study Note온프렘 GPU 플랫폼

3. KServe 이해하기

KServe는 모델을 계산하는 엔진이 아니라 모델 서버의 생명주기를 관리하는 Kubernetes control plane이다

이 장에서 처음 나오는 말5개
CRDCustom Resource Definition
Kubernetes에 InferenceService 같은 새 객체 종류를 추가하는 schema다.
reconcile상태 맞추기
controller가 원하는 상태와 실제 상태의 차이를 반복해서 줄이는 동작이다.
InferenceService
model·runtime·resource·replica를 선언하는 KServe의 기본 serving 객체다.
ServingRuntime
어떤 container image와 command가 어떤 model format을 실행하는지 정한 재사용 template다.
LLMInferenceService
지능형 routing·다중 노드·prefill/decode 같은 고급 LLM topology를 위한 별도 CRD다.
운영자가 선언한 InferenceService와 ServingRuntime을 KServe controller가 받아 Deployment와 Service·HTTPRoute를 만들고, model artifact를 실은 vLLM·FastAPI Pod로 LiteLLM과 앱의 요청이 닿는 구조

KServe가 request를 받아 직접 추론하지 않는다. controller는 운영 경로에 있고, 실제 request는 이미 만들어진 Service를 지나 vLLM·FastAPI 같은 model server로 간다. 그래서 controller가 잠시 재시작돼도 기존 Pod와 Service가 살아 있다면 serving data plane은 계속 동작할 수 있다.

InferenceService · 무엇을

어떤 model을 어떤 resource·replica로 서비스할지 선언한다.

운영자가 주로 만들고 Git에서 version control할 객체다.

ServingRuntime · 어떻게

model format을 실행할 image·command·port·protocol template다.

platform team이 검증해 여러 서비스가 재사용한다.

KServe controller · 계속 맞춘다

두 선언을 읽어 Deployment·Service·route를 만들고 상태를 관찰한다.

drift나 장애가 생기면 원하는 상태로 되돌린다.

한 문장으로는 InferenceService = model 배포 인스턴스, ServingRuntime = 검증된 실행 방법이다.

선언 하나가 실제 서버가 되는 과정

섹션 제목: “선언 하나가 실제 서버가 되는 과정”
kind: InferenceService
metadata:
name: corp-chat
spec:
predictor:
model:
runtime: vllm-a100
storageUri: hf://org/model
resources:
limits:
nvidia.com/gpu: "1"
nodeSelector:
accelerator.pool: a100-serving
  1. API server가 저장한다 — CRD schema로 검증한 InferenceService를 etcd에 저장한다.
  2. KServe controller가 감지한다 — 새 객체나 spec 변경을 보고 reconcile을 시작한다.
  3. runtime을 합성한다 — vllm-a100의 image·command template에 model URI와 resource를 넣는다.
  4. Kubernetes 객체를 만든다 — Standard mode에서는 Deployment·Service와 선택한 network route를 만든다.
  5. Kubernetes가 Pod를 띄운다 — scheduler가 A100 pool과 빈 GPU를 보고 node를 정한다.
  6. model server가 준비된다 — image pull·model fetch·load·warm-up 뒤 readiness가 성공한다.
  7. 상태를 되돌려 준다 — KServe가 하위 상태를 모아 InferenceService.status와 Ready condition을 갱신한다.

spec을 바꾸면 이 흐름이 다시 돈다. Pod를 직접 고쳐도 controller의 원본은 InferenceService이므로 다음 reconcile에서 되돌아간다. 이것이 “선언형”의 실무적 의미다.

Standard mode는 일반 Kubernetes Deployment·Service를 사용한다. 상시 LLM은 model이 크고 cold start가 길며 stream이 오래 이어지므로 온프렘 환경에서는 가장 예측 가능한 출발점이다.

질문Standard InferenceService가 주는 답
replica를 몇 개 유지할까Deployment replica·autoscaling policy
죽은 model server는 누가 복구할까Kubernetes controller와 KServe 상태 관리
endpoint를 어떻게 유지할까Service와 Gateway API·Ingress
image·model을 어떻게 바꿀까spec 변경과 rollout
서로 다른 model runtime을 어떻게 표준화할까ServingRuntime catalog

Knative mode는 CPU 예측 모델의 scale-to-zero처럼 명확한 이유가 있을 때 검토한다. GPU LLM에서 scale-to-zero는 수분짜리 model load를 첫 사용자 요청에 떠넘길 수 있다.

runtime catalog가 플랫폼의 안전장치다

섹션 제목: “runtime catalog가 플랫폼의 안전장치다”
runtime대상과 고정할 계약
vllm-sparkGB10 · ARM64 image·Spark driver/CUDA·한 node
vllm-a100A100 full · x86_64 image·vLLM·GPU preset
vllm-a100-migA100 MIG · model 크기·MIG resource·memory 여유
vllm-b300B300 · driver/CUDA·kernel·full-node 기준선
fastapi-cudacustom model · port·readiness·SIGTERM·metric

latest image나 범용 runtime 하나로 모든 GPU를 받지 않는다. runtime name이 architecture·가속기 계약을 드러내고, policy가 잘못된 image·nodeSelector·resource 조합을 거부해야 한다.

LLMInferenceService는 언제 필요한가

섹션 제목: “LLMInferenceService는 언제 필요한가”

Standard path가 model Pod의 배포와 endpoint를 푼다면, LLMInferenceService는 큰 LLM을 여러 Pod와 지능형 router로 운영하는 topology를 푼다.

사용자 요청이 Gateway API와 inference scheduler를 지나 decode·prefill vLLM workload로 전달되고 LLMInferenceService controller가 이를 관리하는 구조
위에서 아래로 요청을 따라간다. Gateway API가 입구를 만들고 EPP scheduler가 KV cache·부하를 보고 endpoint를 고른다. 아래의 LLMInferenceService controller는 이 data plane을 생성·관리한다.출처: KServe 공식 문서 — Architecture
필요한 기능추가되는 핵심 조각
replica의 단순 round-robin보다 나은 routingGateway API Inference Extension의 InferencePool·EPP scheduler
여러 node가 한 replica를 구성LeaderWorkerSet가 leader·worker를 한 단위로 관리
긴 prompt와 token 생성을 다른 pool로 분리prefill·decode workload와 그 사이의 KV 전달
prefix cache·부하를 고려한 endpoint 선택scheduler plugin과 pod metric
Standard InferenceService로 SLO를 만족하면 그대로 두고, 아니라면 prefix·replica 부하는 Inference Extension, 모델이 한 노드에 안 들어가면 LeaderWorkerSet, prefill이 decode를 방해하면 분리로 가서 모두 LLMInferenceService로 모이는 판단 흐름

고급 경로에는 Gateway API CRD·Inference Extension·Gateway provider·LeaderWorkerSet 같은 의존성이 추가된다. 기능 이름이 매력적이라는 이유로 처음부터 설치하지 않는다. Standard에서 image·GPU·artifact· streaming을 먼저 합격시키고, 측정된 병목이 있을 때 이동한다.

  • 최종 이용자의 API key·team quota·공개 alias — LiteLLM
  • GPU driver·device discovery·MIG·DCGM — GPU Operator
  • model memory에 맞는 GPU 자동 추천 — platform preset과 benchmark
  • model 품질 평가·승인 — 별도 CI·registry workflow
  • batch queue·학습 lifecycle·Notebook — 이 덱의 범위 밖
InferenceService Ready condition
→ KServe controller event
→ Deployment desired / available
→ Pod event · readiness
→ model server log
→ GPU Operator와 node 상태

상위 상태부터 내려가면 “KServe가 문제인지, 일반 Kubernetes 배치가 문제인지, model process가 문제인지, GPU 기반이 문제인지”를 빠르게 나눌 수 있다.