didi-lot2-backend/backend/observability/README.md
2026-07-10 03:39:53 -07:00

8 KiB
Raw Permalink Blame History

DiDi Observability Stack

Stack complet de monitorizare, logging și tracing distribuit pentru platforma DiDi. Implementează cerința LOT 2 modul 8 (Observabilitate & Logging) din caietul de sarcini + diagrama A.9 din oferta EVOTECH.

Componente

Component Rol URL UI (LAN)
Prometheus Scraping metrici + alert rules http://10.11.10.12:9090
Grafana Dashboard-uri vizualizare metrici + loguri + traces http://10.11.10.12:3030
Loki Agregare loguri http://10.11.10.12:3100
Promtail Shipper Docker logs → Loki (no UI)
Jaeger UI tracing distribuit http://10.11.10.12:16686
OTel Collector Receiver OTLP (trace + metric) + processor + exporter http://10.11.10.12:4319 (gRPC), :4320 (HTTP)
Alertmanager Routing alerte (email, Slack/Teams, PagerDuty) http://10.11.10.12:9093

Quick start

cd /home/admin365/didi_mono/didi_mono/backend/observability

# Setup .env
cp .env.example .env
$EDITOR .env  # set GRAFANA_ADMIN_PASSWORD + SENDGRID_API_KEY

# Start stack
docker compose up -d

# Verify all healthy
docker compose ps

Apoi accesează Grafana la http://10.11.10.12:3030 (user admin, parola din GRAFANA_ADMIN_PASSWORD).

Datasources sunt deja provisioned (Prometheus, Loki, Jaeger). Adaugă dashboard-uri custom în grafana/dashboards/ (auto-provisioned la 30s).

Instrumentare servicii

Node.js (agent-v3, didi-framework, admin-dashboard)

Instalează prom-client + @opentelemetry/sdk-node:

npm install --save prom-client @opentelemetry/api @opentelemetry/sdk-node \
  @opentelemetry/auto-instrumentations-node @opentelemetry/exporter-trace-otlp-grpc \
  @opentelemetry/resources @opentelemetry/semantic-conventions

Adaugă în src/index.ts (înainte de orice alt import):

import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-grpc';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { Resource } from '@opentelemetry/resources';
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';

const sdk = new NodeSDK({
  resource: new Resource({
    [SemanticResourceAttributes.SERVICE_NAME]: 'agent-v3',
    [SemanticResourceAttributes.SERVICE_VERSION]: '3.0.0',
  }),
  traceExporter: new OTLPTraceExporter({
    url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || 'http://didi-otel-collector:4317',
  }),
  instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();

Adaugă endpoint /metrics în Express:

import { register, collectDefaultMetrics } from 'prom-client';
collectDefaultMetrics({ prefix: 'didi_agent_v3_' });

app.get('/metrics', async (req, res) => {
  res.set('Content-Type', register.contentType);
  res.end(await register.metrics());
});

Custom metrics relevante DiDi:

import { Counter, Histogram } from 'prom-client';

export const analysisCompleted = new Counter({
  name: 'didi_analyses_completed_total',
  help: 'Total number of analyses completed',
  labelNames: ['component', 'tier', 'media_type', 'verdict'],
});

export const pipelineDuration = new Histogram({
  name: 'didi_pipeline_duration_seconds',
  help: 'Pipeline duration in seconds',
  labelNames: ['component', 'tier'],
  buckets: [1, 5, 10, 30, 60, 120, 300],
});

// În executor:
const end = pipelineDuration.startTimer({ component: 'techniques', tier });
try {
  await runAnalysis();
  analysisCompleted.inc({ component: 'techniques', tier, media_type, verdict });
} finally {
  end();
}

Python (ai_platform modules)

Instalează prometheus_client + opentelemetry-instrumentation-fastapi:

pip install prometheus-client opentelemetry-api opentelemetry-sdk \
  opentelemetry-exporter-otlp opentelemetry-instrumentation-fastapi

Adaugă în app.py:

from prometheus_client import make_asgi_app
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.sdk.resources import Resource

resource = Resource(attributes={"service.name": "llm-inference"})
trace.set_tracer_provider(TracerProvider(resource=resource))
trace.get_tracer_provider().add_span_processor(
    BatchSpanProcessor(OTLPSpanExporter(endpoint="http://didi-otel-collector:4317", insecure=True))
)

app = FastAPI()
FastAPIInstrumentor.instrument_app(app)

# Mount /metrics
app.mount("/metrics", make_asgi_app())

Environment variables pe servicii

Adaugă în compose-urile fiecărui serviciu:

environment:
  OTEL_EXPORTER_OTLP_ENDPOINT: http://didi-otel-collector:4317
  OTEL_SERVICE_NAME: agent-v3
  OTEL_RESOURCE_ATTRIBUTES: cluster=didi-prod,environment=production

Dashboards predefinite

Adaugă fișiere JSON în grafana/dashboards/ (auto-importate). Recomandate:

  1. DiDi Platform Overview

    • KPI cards: analize/min, error rate, P50/P95/P99 pipeline latency, GPU util
    • Time-series: requests per service, queue backlog (RabbitMQ), Redis hit rate
  2. AI Platform

    • LLM inference latency per model, GPU memory (Qwen 3.5, BusterX)
    • Brain analysis_atom cache hit rate, scheduler health
  3. Cozi & Workers

    • RabbitMQ queue depth per component × tier
    • Worker throughput, retry count, DLQ messages
  4. Cost & business

    • Stripe webhook success rate, Sales Invoice rate, MRR proxy
  5. Infrastructure

    • CPU/RAM/Disk/Network per node, container restarts, healthcheck failures

Import dashboards exemple din comunitate:

  • ID 11074 (Node Exporter Full)
  • ID 13639 (Logs via Loki)
  • ID 17761 (Cadvisor)
  • ID 14570 (RabbitMQ Cluster)

În Grafana: + → Import → paste ID-ul → load.

Verificare end-to-end

# 1. Verifică Prometheus scrapes
curl -s http://10.11.10.12:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'

# 2. Verifică Loki primește loguri
curl -s 'http://10.11.10.12:3100/loki/api/v1/labels' | jq

# 3. Trimite un trace de test din shell (după ce ai un service instrumentat)
# Vizibil în Jaeger UI: http://10.11.10.12:16686

# 4. Trimite o alertă de test
curl -X POST http://10.11.10.12:9090/-/reload
# Așteaptă să se trigger ServiceDown alert (timer ~5min)

Retention

  • Prometheus: 30 zile (configurabil în compose --storage.tsdb.retention.time)
  • Loki: 7 zile (configurabil în loki-config.yaml retention_period)
  • Jaeger (Badger storage): persistent, ~10GB cap
  • Alertmanager: persistent state

SLA + escalation

Vezi alertmanager.yml pentru routing:

  • severity=critical → notify imediat la office@clossers.com
  • severity=warning → batch, repeat 12h
  • TODO: adaugă Slack/Teams webhook + PagerDuty key pentru on-call

Troubleshooting

Prometheus arată target-uri DOWN: verifică că serviciul are endpoint /metrics accesibil + că e pe didi-network.

Loki nu primește loguri: verifică promtail logs (docker logs didi-promtail) — probabil container labels nu match-uiesc.

Jaeger fără traces: verifică că serviciul are env OTEL_EXPORTER_OTLP_ENDPOINT setat corect + că face HTTP/gRPC către didi-otel-collector:4317.

Alertmanager nu trimite email: verifică SENDGRID_API_KEY în env + că from address alerts@didi365.eu e validat în SendGrid.

Roadmap

  • Adaugă postgres_exporter pentru metrici PostgreSQL (slow queries, replication lag)
  • Adaugă redis_exporter pentru metrici Redis
  • Adaugă nvidia_gpu_exporter pe GPU host pentru metrici VRAM/utilization
  • Instrumentare agent-v3 (PR follow-up)
  • Instrumentare ai_platform modules (PR follow-up)
  • Slack/Teams webhook integration
  • PagerDuty on-call rotation
  • Synthetic monitoring (uptime checks pe didi365.eu)

Referințe

  • Caiet sarcini LOT 2 modul 8 (Observabilitate & Logging)
  • Oferta EVOTECH §A.9 (diagrama observabilitate completă)
  • Cercetare industrială §C (validare experimentală + KPI-uri)