Monolithic
하나의 litellm 서비스가 LLM traffic, 관리 API와 UI를 제공한다. 공식 litellm-helm chart의 단순 경로다.
작은 플랫폼 팀이 요청 흐름과 상태부터 익히기에 적합하다.
Kubernetes에서 LiteLLM Pod는 가볍게 복제할 수 있다 — 어렵고 중요한 것은 Pod 밖의 공유 상태다
monolithiccomponentizedPDBPodDisruptionBudgetmigration job첫 production은 monolithic chart, 2개 이상 replica, worker 1개/Pod로 시작하는 편이 운영 경계가 작다. 트래픽 계층과 관리 UI의 확장·보안·release 주기를 분리해야 할 이유가 생기면 componentized 모드로 간다.
Monolithic
하나의 litellm 서비스가 LLM traffic, 관리 API와 UI를 제공한다. 공식 litellm-helm chart의 단순 경로다.
작은 플랫폼 팀이 요청 흐름과 상태부터 익히기에 적합하다.
Componentized
gateway :4000, backend :4001, UI :3000을 나눈다. gateway만 크게 확장하거나 관리 경로를 별도 정책으로
노출할 수 있지만 Service·routing·probe·version 호환의 운영 대상이 늘어난다.
“microservices가 더 production답다”는 이유만으로 나누지 않는다. gateway와 관리 평면을 독립 확장하거나 장애·보안 경계를 분리할 명확한 요구가 있을 때 선택한다.
LiteLLM은 양쪽 네트워크가 모두 중요하다.
| 방향 | 연결 | 실패 시 보이는 현상 |
|---|---|---|
| northbound | client → Gateway → LiteLLM Service | DNS·TLS·401·연결 timeout |
| state | LiteLLM → Postgres·Redis | readiness 실패, 인증·limit 불일치 |
| southbound internal | LiteLLM → vLLM/KServe/GPUStack | 특정 model 5xx·timeout |
| southbound external | LiteLLM → egress proxy → provider | proxy auth·CA·방화벽·429 |
| telemetry | LiteLLM → Langfuse/OTel, Prometheus → LiteLLM | 요청은 성공하지만 trace·metric 유실 |
NetworkPolicy와 방화벽은 이 방향별 allowlist로 만든다. LiteLLM Pod에 목적지 제한 없는 인터넷 egress를 주는 것은 provider 추가를 편하게 하지만 데이터 반출 경계도 없앤다.
공식 production 지침은 Kubernetes에서 Uvicorn worker를 Pod마다 하나 두고 Pod를 수평 확장하는 구성을 권장한다. 여러 worker를 한 Pod에 넣으면 CPU·메모리·DB connection·background job 수가 worker 수만큼 늘고 HPA가 내부 process를 구분하지 못한다.
총 process 수 = replica 수 × Pod당 worker 수Redis 필요 여부와 DB connection 상한은 Pod 수가 아니라 이 process 수를 기준으로 본다.
/health/liveliness — process가 살아 있는지 본다./health/readiness — traffic을 받을 준비와 configured DB 연결을 본다.liveness에서 provider 호출을 하면 provider 장애 때 모든 Pod가 재시작되는 연쇄 장애가 난다. 반대로 readiness가 너무 느슨하면 DB를 못 읽는 Pod가 traffic을 받는다. 세 질문을 한 endpoint로 합치지 않는다.
종료 때는 readiness를 먼저 내리고 기존 streaming 요청을 drain할 시간을 준다. terminationGracePeriodSeconds,
preStop/lifecycle과 Gateway timeout을 실제 최대 stream 시간에 맞춘다.