05 · Observability5.1 · Метрики5.1.2средний

Prometheus, VictoriaMetrics и Grafana

Зачем это нужно

Метрики USE/RED бесполезны, если их никто не собирает и не показывает. Стек курса: Prometheus (или совместимый VictoriaMetrics) для хранения time series, Grafana для dashboards и alert visualization. Этот урок — как pieces fit together в Kubernetes для ML-сервисов.

Основные идеи

Pull vs push.

  • Prometheus по умолчанию scrapes targets (HTTP GET /metrics) по interval.

  • Pushgateway — исключение для batch jobs (краткоживущие pods); не замена для сервисов.

Компоненты в K8s.

Component Role
exporter / app Exposes /metrics in Prometheus text format
ServiceMonitor (Prometheus Operator) Declares scrape config for Service
Prometheus / VM TSDB storage, PromQL queries
Grafana Dashboards, alerts UI
Alertmanager Routing notifications (урок 5.3.2)

Prometheus text format example.

`

HELP http_requests_total Total HTTP requests

TYPE http_requests_total counter

http_requests_total{method="POST",handler="/predict",code="200"} 1024 `

VictoriaMetrics (VM). Prometheus-compatible API и PromQL; часто дешевле по RAM/ disk at scale. Remote write из Prometheus → VM — типовой hybrid. Для курса: те же queries, другой backend.

ServiceMonitor pattern.

apiVersion: monitoring.coreos.com/v1 kind: ServiceMonitor metadata: name: churn-serving spec: selector: matchLabels: app: churn-serving endpoints: - port: metrics interval: 15s

Pod Service must expose port metrics; app or sidecar implements exporter.

PromQL essentials.

  • rate(counter[5m]) — per-second rate for counters.

  • histogram_quantile(0.95, ...) — percentiles from histograms.

  • sum by (pod) (...) — aggregation.

Percentile и SLO — не одно и то же. p95 < 200ms означает, что 95-й перцентиль latency ниже порога; request-based SLI requests < 200ms / all requests прямо считает долю good events. Они могут быть связаны, но SLO «99.5% запросов быстрее 200ms» нельзя называть «p95 < 200ms»: для него нужен порог 99.5% и явная формула good-event ratio.

Grafana.

  • Data source → Prometheus or VM URL.

  • Dashboard as code — JSON in Git (optional mature practice).

  • Variables: $namespace, $model_version для drill-down.

Istio metrics. Sidecar exposes request metrics; standard RED without app instrumentation. Labels: destination_service, response_code.

Retention and cardinality. Prometheus local storage ~15d common; long-term → VM or object storage. Drop high-cardinality labels at scrape or recording rules.

ML platform dashboards layers.

  1. Cluster/node (USE).

  2. Namespace workloads (pods, HPA).

  3. Per-model inference (RED + version label).

  4. Business KPI (often from other DB, not always Prometheus).

Как это выглядит на практике

Namespace ml-prod:

  1. Inference deployment with annotations for Prometheus or dedicated prometheus-fastapi-instrumentator on port 8080 /metrics.

  2. ServiceMonitor in monitoring namespace selects pods app=churn-serving.

  3. Prometheus Operator configures scrape.

  4. Grafana dashboard imported: «Kubernetes / Istio / Custom ML».

  5. Recording rule: job:churn:request_latency_p95:5m for faster alerts.

Query for error rate (Istio):

sum(rate(istio_requests_total{destination_service="churn-serving",response_code=~"5.."}[5m])) / sum(rate(istio_requests_total{destination_service="churn-serving"}[5m]))

Upgrade path: Prometheus scrapes → remote_write VictoriaMetrics → Grafana queries VM for heavy dashboards.

Что сделать после занятия

  • Найдите в кластере или docs, где живёт Prometheus/Grafana; откройте один dashboard.

  • Напишите PromQL для RPS и p95 latency вашего сервиса (conceptual labels).

  • Сравните: что даёт Istio metrics vs app-native /metrics.

Официальные материалы