07 · Serving-платформа7.1–7.7 · Runtime и оркестраторы7.6.1эксперт

vLLM для LLM serving

Зачем это нужно

Large Language Models (LLM) не помещаются в классический MLServer workflow: нужны GPU, PagedAttention, continuous batching, tokenizer pipeline, возможно multi-GPU. vLLM — высокопроизводительный open-source engine для serving LLM с OpenAI-compatible API.

MLOps для GenAI — отдельный контур; этот урок даёт базу, чтобы не путать «деплой BERT 2019» и «деплой Llama 3 8B».

Основные идеи

Почему не просто Flask + transformers.

Challenge vLLM approach
GPU memory fragmentation PagedAttention
Variable sequence length Continuous batching
Throughput at scale Optimized CUDA kernels
Standard client API OpenAI-compatible /v1/completions, /v1/chat/completions

Deployment modes.

  1. Standalonevllm serve in Deployment.

  2. KServe — custom predictor container or community ServingRuntime.

  3. Behind gateway — rate limit, auth, logging (module 3.6).

Minimal K8s sketch:

spec: containers: - name: vllm image: vllm/vllm-openai:v0.8.5 args: - "--model" - "meta-llama/Llama-3.2-3B-Instruct" - "--max-model-len" - "4096" resources: limits: nvidia.com/gpu: "1" memory: "24Gi"

Model weights. Download from Hugging Face (license!) or internal MinIO mirror; initContainer or pre-baked PVC. Air-gapped: model on PVC snapshot. Не используйте mutable latest: в production закрепляйте проверенный version tag и, по возможности, image digest.

Key parameters.

  • max-model-len — context window cap; affects VRAM.

  • tensor-parallel-size — multi-GPU single model.

  • gpu-memory-utilization — fraction of VRAM vLLM uses (default 0.9).

OpenAI-compatible client:

from openai import OpenAI client = OpenAI(base_url="https://llm.ml.example.com/v1", api_key="dummy") resp = client.chat.completions.create( model="meta-llama/Llama-3.2-3B-Instruct", messages=[{"role": "user", "content": "Explain Kafka in 3 sentences"}], )

Streaming. SSE token stream — UX for chat; observability counts time-to-first-token (TTFT) separately from total latency.

Observability.

  • RED on gateway + vLLM metrics (if exposed).

  • GPU USE (module 5.1.3).

  • Log prompt/response — PII policy; often hash or omit content in prod logs.

Safety & cost. Rate limits, max tokens, content filters upstream; GPU expensive — autoscale conservative, minReplicas tradeoff.

vs Triton for LLM. Triton + TensorRT-LLM backend — alternative; vLLM often faster iteration for HuggingFace models in university/lab settings.

Как это выглядит на практике

Lab MDP: one A100 node pool, single Llama 3.2 3B for course assistants. Ingress with JWT auth; students share endpoint with quota.

Prod pattern: model mirror in MinIO → initContainer → vLLM; HPA not automatic — manual scale for predictable load; KEDA on queue depth for async batch.

Incident: OOM at startup → reduce max-model-len or use smaller quant (AWQ/GPTQ) — advanced topic.

Что сделать после занятия

  • Объясните, зачем LLM нужен отдельный runtime от MLServer.

  • Выпишите 3 параметра vLLM, влияющих на GPU memory.

  • Опишите политику логирования prompts в prod (2–3 правила).

Официальные материалы