234 lines
8 KiB
Markdown
234 lines
8 KiB
Markdown
# DiDi Observability Stack
|
||
|
||
Stack complet de monitorizare, logging și tracing distribuit pentru platforma DiDi. Implementează cerința **LOT 2 modul 8 (Observabilitate & Logging)** din caietul de sarcini + diagrama A.9 din oferta EVOTECH.
|
||
|
||
## Componente
|
||
|
||
| Component | Rol | URL UI (LAN) |
|
||
|---|---|---|
|
||
| **Prometheus** | Scraping metrici + alert rules | http://10.11.10.12:9090 |
|
||
| **Grafana** | Dashboard-uri vizualizare metrici + loguri + traces | http://10.11.10.12:3030 |
|
||
| **Loki** | Agregare loguri | http://10.11.10.12:3100 |
|
||
| **Promtail** | Shipper Docker logs → Loki | (no UI) |
|
||
| **Jaeger** | UI tracing distribuit | http://10.11.10.12:16686 |
|
||
| **OTel Collector** | Receiver OTLP (trace + metric) + processor + exporter | http://10.11.10.12:4319 (gRPC), :4320 (HTTP) |
|
||
| **Alertmanager** | Routing alerte (email, Slack/Teams, PagerDuty) | http://10.11.10.12:9093 |
|
||
|
||
## Quick start
|
||
|
||
```bash
|
||
cd /home/admin365/didi_mono/didi_mono/backend/observability
|
||
|
||
# Setup .env
|
||
cp .env.example .env
|
||
$EDITOR .env # set GRAFANA_ADMIN_PASSWORD + SENDGRID_API_KEY
|
||
|
||
# Start stack
|
||
docker compose up -d
|
||
|
||
# Verify all healthy
|
||
docker compose ps
|
||
```
|
||
|
||
Apoi accesează **Grafana** la http://10.11.10.12:3030 (user `admin`, parola din `GRAFANA_ADMIN_PASSWORD`).
|
||
|
||
Datasources sunt deja provisioned (Prometheus, Loki, Jaeger). Adaugă dashboard-uri custom în `grafana/dashboards/` (auto-provisioned la 30s).
|
||
|
||
## Instrumentare servicii
|
||
|
||
### Node.js (agent-v3, didi-framework, admin-dashboard)
|
||
|
||
Instalează `prom-client` + `@opentelemetry/sdk-node`:
|
||
```bash
|
||
npm install --save prom-client @opentelemetry/api @opentelemetry/sdk-node \
|
||
@opentelemetry/auto-instrumentations-node @opentelemetry/exporter-trace-otlp-grpc \
|
||
@opentelemetry/resources @opentelemetry/semantic-conventions
|
||
```
|
||
|
||
Adaugă în `src/index.ts` (înainte de orice alt import):
|
||
```ts
|
||
import { NodeSDK } from '@opentelemetry/sdk-node';
|
||
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-grpc';
|
||
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
|
||
import { Resource } from '@opentelemetry/resources';
|
||
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';
|
||
|
||
const sdk = new NodeSDK({
|
||
resource: new Resource({
|
||
[SemanticResourceAttributes.SERVICE_NAME]: 'agent-v3',
|
||
[SemanticResourceAttributes.SERVICE_VERSION]: '3.0.0',
|
||
}),
|
||
traceExporter: new OTLPTraceExporter({
|
||
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || 'http://didi-otel-collector:4317',
|
||
}),
|
||
instrumentations: [getNodeAutoInstrumentations()],
|
||
});
|
||
sdk.start();
|
||
```
|
||
|
||
Adaugă endpoint `/metrics` în Express:
|
||
```ts
|
||
import { register, collectDefaultMetrics } from 'prom-client';
|
||
collectDefaultMetrics({ prefix: 'didi_agent_v3_' });
|
||
|
||
app.get('/metrics', async (req, res) => {
|
||
res.set('Content-Type', register.contentType);
|
||
res.end(await register.metrics());
|
||
});
|
||
```
|
||
|
||
Custom metrics relevante DiDi:
|
||
```ts
|
||
import { Counter, Histogram } from 'prom-client';
|
||
|
||
export const analysisCompleted = new Counter({
|
||
name: 'didi_analyses_completed_total',
|
||
help: 'Total number of analyses completed',
|
||
labelNames: ['component', 'tier', 'media_type', 'verdict'],
|
||
});
|
||
|
||
export const pipelineDuration = new Histogram({
|
||
name: 'didi_pipeline_duration_seconds',
|
||
help: 'Pipeline duration in seconds',
|
||
labelNames: ['component', 'tier'],
|
||
buckets: [1, 5, 10, 30, 60, 120, 300],
|
||
});
|
||
|
||
// În executor:
|
||
const end = pipelineDuration.startTimer({ component: 'techniques', tier });
|
||
try {
|
||
await runAnalysis();
|
||
analysisCompleted.inc({ component: 'techniques', tier, media_type, verdict });
|
||
} finally {
|
||
end();
|
||
}
|
||
```
|
||
|
||
### Python (ai_platform modules)
|
||
|
||
Instalează `prometheus_client` + `opentelemetry-instrumentation-fastapi`:
|
||
```bash
|
||
pip install prometheus-client opentelemetry-api opentelemetry-sdk \
|
||
opentelemetry-exporter-otlp opentelemetry-instrumentation-fastapi
|
||
```
|
||
|
||
Adaugă în `app.py`:
|
||
```python
|
||
from prometheus_client import make_asgi_app
|
||
from opentelemetry import trace
|
||
from opentelemetry.sdk.trace import TracerProvider
|
||
from opentelemetry.sdk.trace.export import BatchSpanProcessor
|
||
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
|
||
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
|
||
from opentelemetry.sdk.resources import Resource
|
||
|
||
resource = Resource(attributes={"service.name": "llm-inference"})
|
||
trace.set_tracer_provider(TracerProvider(resource=resource))
|
||
trace.get_tracer_provider().add_span_processor(
|
||
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://didi-otel-collector:4317", insecure=True))
|
||
)
|
||
|
||
app = FastAPI()
|
||
FastAPIInstrumentor.instrument_app(app)
|
||
|
||
# Mount /metrics
|
||
app.mount("/metrics", make_asgi_app())
|
||
```
|
||
|
||
### Environment variables pe servicii
|
||
|
||
Adaugă în compose-urile fiecărui serviciu:
|
||
```yaml
|
||
environment:
|
||
OTEL_EXPORTER_OTLP_ENDPOINT: http://didi-otel-collector:4317
|
||
OTEL_SERVICE_NAME: agent-v3
|
||
OTEL_RESOURCE_ATTRIBUTES: cluster=didi-prod,environment=production
|
||
```
|
||
|
||
## Dashboards predefinite
|
||
|
||
Adaugă fișiere JSON în `grafana/dashboards/` (auto-importate). Recomandate:
|
||
|
||
1. **DiDi Platform Overview**
|
||
- KPI cards: analize/min, error rate, P50/P95/P99 pipeline latency, GPU util
|
||
- Time-series: requests per service, queue backlog (RabbitMQ), Redis hit rate
|
||
|
||
2. **AI Platform**
|
||
- LLM inference latency per model, GPU memory (Qwen 3.5, BusterX)
|
||
- Brain analysis_atom cache hit rate, scheduler health
|
||
|
||
3. **Cozi & Workers**
|
||
- RabbitMQ queue depth per component × tier
|
||
- Worker throughput, retry count, DLQ messages
|
||
|
||
4. **Cost & business**
|
||
- Stripe webhook success rate, Sales Invoice rate, MRR proxy
|
||
|
||
5. **Infrastructure**
|
||
- CPU/RAM/Disk/Network per node, container restarts, healthcheck failures
|
||
|
||
Import dashboards exemple din comunitate:
|
||
- ID 11074 (Node Exporter Full)
|
||
- ID 13639 (Logs via Loki)
|
||
- ID 17761 (Cadvisor)
|
||
- ID 14570 (RabbitMQ Cluster)
|
||
|
||
În Grafana: **+ → Import → paste ID-ul → load**.
|
||
|
||
## Verificare end-to-end
|
||
|
||
```bash
|
||
# 1. Verifică Prometheus scrapes
|
||
curl -s http://10.11.10.12:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
|
||
|
||
# 2. Verifică Loki primește loguri
|
||
curl -s 'http://10.11.10.12:3100/loki/api/v1/labels' | jq
|
||
|
||
# 3. Trimite un trace de test din shell (după ce ai un service instrumentat)
|
||
# Vizibil în Jaeger UI: http://10.11.10.12:16686
|
||
|
||
# 4. Trimite o alertă de test
|
||
curl -X POST http://10.11.10.12:9090/-/reload
|
||
# Așteaptă să se trigger ServiceDown alert (timer ~5min)
|
||
```
|
||
|
||
## Retention
|
||
|
||
- **Prometheus**: 30 zile (configurabil în compose `--storage.tsdb.retention.time`)
|
||
- **Loki**: 7 zile (configurabil în `loki-config.yaml` `retention_period`)
|
||
- **Jaeger** (Badger storage): persistent, ~10GB cap
|
||
- **Alertmanager**: persistent state
|
||
|
||
## SLA + escalation
|
||
|
||
Vezi `alertmanager.yml` pentru routing:
|
||
- `severity=critical` → notify imediat la `office@clossers.com`
|
||
- `severity=warning` → batch, repeat 12h
|
||
- TODO: adaugă Slack/Teams webhook + PagerDuty key pentru on-call
|
||
|
||
## Troubleshooting
|
||
|
||
**Prometheus arată target-uri DOWN**: verifică că serviciul are endpoint `/metrics` accesibil + că e pe `didi-network`.
|
||
|
||
**Loki nu primește loguri**: verifică `promtail` logs (`docker logs didi-promtail`) — probabil container labels nu match-uiesc.
|
||
|
||
**Jaeger fără traces**: verifică că serviciul are env `OTEL_EXPORTER_OTLP_ENDPOINT` setat corect + că face HTTP/gRPC către `didi-otel-collector:4317`.
|
||
|
||
**Alertmanager nu trimite email**: verifică `SENDGRID_API_KEY` în env + că from address `alerts@didi365.eu` e validat în SendGrid.
|
||
|
||
## Roadmap
|
||
|
||
- [ ] Adaugă **postgres_exporter** pentru metrici PostgreSQL (slow queries, replication lag)
|
||
- [ ] Adaugă **redis_exporter** pentru metrici Redis
|
||
- [ ] Adaugă **nvidia_gpu_exporter** pe GPU host pentru metrici VRAM/utilization
|
||
- [ ] Instrumentare agent-v3 (PR follow-up)
|
||
- [ ] Instrumentare ai_platform modules (PR follow-up)
|
||
- [ ] Slack/Teams webhook integration
|
||
- [ ] PagerDuty on-call rotation
|
||
- [ ] Synthetic monitoring (uptime checks pe didi365.eu)
|
||
|
||
## Referințe
|
||
|
||
- Caiet sarcini LOT 2 modul 8 (Observabilitate & Logging)
|
||
- Oferta EVOTECH §A.9 (diagrama observabilitate completă)
|
||
- Cercetare industrială §C (validare experimentală + KPI-uri)
|