Зачем это нужно
Large Language Models (LLM) не помещаются в классический MLServer workflow: нужны GPU, PagedAttention, continuous batching, tokenizer pipeline, возможно multi-GPU. vLLM — высокопроизводительный open-source engine для serving LLM с OpenAI-compatible API.
MLOps для GenAI — отдельный контур; этот урок даёт базу, чтобы не путать «деплой BERT 2019» и «деплой Llama 3 8B».
Основные идеи
Почему не просто Flask + transformers.
| Challenge | vLLM approach |
|---|---|
| GPU memory fragmentation | PagedAttention |
| Variable sequence length | Continuous batching |
| Throughput at scale | Optimized CUDA kernels |
| Standard client API | OpenAI-compatible /v1/completions, /v1/chat/completions |
Deployment modes.
Standalone —
vllm servein Deployment.KServe — custom predictor container or community ServingRuntime.
Behind gateway — rate limit, auth, logging (module 3.6).
Minimal K8s sketch:
spec: containers: - name: vllm image: vllm/vllm-openai:v0.8.5 args: - "--model" - "meta-llama/Llama-3.2-3B-Instruct" - "--max-model-len" - "4096" resources: limits: nvidia.com/gpu: "1" memory: "24Gi"
Model weights. Download from Hugging Face (license!) or internal MinIO mirror; initContainer or pre-baked PVC. Air-gapped: model on PVC snapshot. Не используйте mutable latest: в production закрепляйте проверенный version tag и, по возможности, image digest.
Key parameters.
max-model-len— context window cap; affects VRAM.tensor-parallel-size— multi-GPU single model.gpu-memory-utilization— fraction of VRAM vLLM uses (default 0.9).
OpenAI-compatible client:
from openai import OpenAI client = OpenAI(base_url="https://llm.ml.example.com/v1", api_key="dummy") resp = client.chat.completions.create( model="meta-llama/Llama-3.2-3B-Instruct", messages=[{"role": "user", "content": "Explain Kafka in 3 sentences"}], )
Streaming. SSE token stream — UX for chat; observability counts time-to-first-token (TTFT) separately from total latency.
Observability.
RED on gateway + vLLM metrics (if exposed).
GPU USE (module 5.1.3).
Log prompt/response — PII policy; often hash or omit content in prod logs.
Safety & cost. Rate limits, max tokens, content filters upstream; GPU expensive — autoscale conservative, minReplicas tradeoff.
vs Triton for LLM. Triton + TensorRT-LLM backend — alternative; vLLM often faster iteration for HuggingFace models in university/lab settings.
Как это выглядит на практике
Lab MDP: one A100 node pool, single Llama 3.2 3B for course assistants. Ingress with JWT auth; students share endpoint with quota.
Prod pattern: model mirror in MinIO → initContainer → vLLM; HPA not automatic — manual scale for predictable load; KEDA on queue depth for async batch.
Incident: OOM at startup → reduce max-model-len or use smaller quant (AWQ/GPTQ) — advanced topic.
Что сделать после занятия
Объясните, зачем LLM нужен отдельный runtime от MLServer.
Выпишите 3 параметра vLLM, влияющих на GPU memory.
Опишите политику логирования prompts в prod (2–3 правила).