didi-lot2-backend/backend/observability/README.md
2026-07-10 03:39:53 -07:00

234 lines
8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# DiDi Observability Stack
Stack complet de monitorizare, logging și tracing distribuit pentru platforma DiDi. Implementează cerința **LOT 2 modul 8 (Observabilitate & Logging)** din caietul de sarcini + diagrama A.9 din oferta EVOTECH.
## Componente
| Component | Rol | URL UI (LAN) |
|---|---|---|
| **Prometheus** | Scraping metrici + alert rules | http://10.11.10.12:9090 |
| **Grafana** | Dashboard-uri vizualizare metrici + loguri + traces | http://10.11.10.12:3030 |
| **Loki** | Agregare loguri | http://10.11.10.12:3100 |
| **Promtail** | Shipper Docker logs → Loki | (no UI) |
| **Jaeger** | UI tracing distribuit | http://10.11.10.12:16686 |
| **OTel Collector** | Receiver OTLP (trace + metric) + processor + exporter | http://10.11.10.12:4319 (gRPC), :4320 (HTTP) |
| **Alertmanager** | Routing alerte (email, Slack/Teams, PagerDuty) | http://10.11.10.12:9093 |
## Quick start
```bash
cd /home/admin365/didi_mono/didi_mono/backend/observability
# Setup .env
cp .env.example .env
$EDITOR .env # set GRAFANA_ADMIN_PASSWORD + SENDGRID_API_KEY
# Start stack
docker compose up -d
# Verify all healthy
docker compose ps
```
Apoi accesează **Grafana** la http://10.11.10.12:3030 (user `admin`, parola din `GRAFANA_ADMIN_PASSWORD`).
Datasources sunt deja provisioned (Prometheus, Loki, Jaeger). Adaugă dashboard-uri custom în `grafana/dashboards/` (auto-provisioned la 30s).
## Instrumentare servicii
### Node.js (agent-v3, didi-framework, admin-dashboard)
Instalează `prom-client` + `@opentelemetry/sdk-node`:
```bash
npm install --save prom-client @opentelemetry/api @opentelemetry/sdk-node \
@opentelemetry/auto-instrumentations-node @opentelemetry/exporter-trace-otlp-grpc \
@opentelemetry/resources @opentelemetry/semantic-conventions
```
Adaugă în `src/index.ts` (înainte de orice alt import):
```ts
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-grpc';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { Resource } from '@opentelemetry/resources';
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: 'agent-v3',
[SemanticResourceAttributes.SERVICE_VERSION]: '3.0.0',
}),
traceExporter: new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || 'http://didi-otel-collector:4317',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
```
Adaugă endpoint `/metrics` în Express:
```ts
import { register, collectDefaultMetrics } from 'prom-client';
collectDefaultMetrics({ prefix: 'didi_agent_v3_' });
app.get('/metrics', async (req, res) => {
res.set('Content-Type', register.contentType);
res.end(await register.metrics());
});
```
Custom metrics relevante DiDi:
```ts
import { Counter, Histogram } from 'prom-client';
export const analysisCompleted = new Counter({
name: 'didi_analyses_completed_total',
help: 'Total number of analyses completed',
labelNames: ['component', 'tier', 'media_type', 'verdict'],
});
export const pipelineDuration = new Histogram({
name: 'didi_pipeline_duration_seconds',
help: 'Pipeline duration in seconds',
labelNames: ['component', 'tier'],
buckets: [1, 5, 10, 30, 60, 120, 300],
});
// În executor:
const end = pipelineDuration.startTimer({ component: 'techniques', tier });
try {
await runAnalysis();
analysisCompleted.inc({ component: 'techniques', tier, media_type, verdict });
} finally {
end();
}
```
### Python (ai_platform modules)
Instalează `prometheus_client` + `opentelemetry-instrumentation-fastapi`:
```bash
pip install prometheus-client opentelemetry-api opentelemetry-sdk \
opentelemetry-exporter-otlp opentelemetry-instrumentation-fastapi
```
Adaugă în `app.py`:
```python
from prometheus_client import make_asgi_app
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.sdk.resources import Resource
resource = Resource(attributes={"service.name": "llm-inference"})
trace.set_tracer_provider(TracerProvider(resource=resource))
trace.get_tracer_provider().add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://didi-otel-collector:4317", insecure=True))
)
app = FastAPI()
FastAPIInstrumentor.instrument_app(app)
# Mount /metrics
app.mount("/metrics", make_asgi_app())
```
### Environment variables pe servicii
Adaugă în compose-urile fiecărui serviciu:
```yaml
environment:
OTEL_EXPORTER_OTLP_ENDPOINT: http://didi-otel-collector:4317
OTEL_SERVICE_NAME: agent-v3
OTEL_RESOURCE_ATTRIBUTES: cluster=didi-prod,environment=production
```
## Dashboards predefinite
Adaugă fișiere JSON în `grafana/dashboards/` (auto-importate). Recomandate:
1. **DiDi Platform Overview**
- KPI cards: analize/min, error rate, P50/P95/P99 pipeline latency, GPU util
- Time-series: requests per service, queue backlog (RabbitMQ), Redis hit rate
2. **AI Platform**
- LLM inference latency per model, GPU memory (Qwen 3.5, BusterX)
- Brain analysis_atom cache hit rate, scheduler health
3. **Cozi & Workers**
- RabbitMQ queue depth per component × tier
- Worker throughput, retry count, DLQ messages
4. **Cost & business**
- Stripe webhook success rate, Sales Invoice rate, MRR proxy
5. **Infrastructure**
- CPU/RAM/Disk/Network per node, container restarts, healthcheck failures
Import dashboards exemple din comunitate:
- ID 11074 (Node Exporter Full)
- ID 13639 (Logs via Loki)
- ID 17761 (Cadvisor)
- ID 14570 (RabbitMQ Cluster)
În Grafana: **+ → Import → paste ID-ul → load**.
## Verificare end-to-end
```bash
# 1. Verifică Prometheus scrapes
curl -s http://10.11.10.12:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
# 2. Verifică Loki primește loguri
curl -s 'http://10.11.10.12:3100/loki/api/v1/labels' | jq
# 3. Trimite un trace de test din shell (după ce ai un service instrumentat)
# Vizibil în Jaeger UI: http://10.11.10.12:16686
# 4. Trimite o alertă de test
curl -X POST http://10.11.10.12:9090/-/reload
# Așteaptă să se trigger ServiceDown alert (timer ~5min)
```
## Retention
- **Prometheus**: 30 zile (configurabil în compose `--storage.tsdb.retention.time`)
- **Loki**: 7 zile (configurabil în `loki-config.yaml` `retention_period`)
- **Jaeger** (Badger storage): persistent, ~10GB cap
- **Alertmanager**: persistent state
## SLA + escalation
Vezi `alertmanager.yml` pentru routing:
- `severity=critical` → notify imediat la `office@clossers.com`
- `severity=warning` → batch, repeat 12h
- TODO: adaugă Slack/Teams webhook + PagerDuty key pentru on-call
## Troubleshooting
**Prometheus arată target-uri DOWN**: verifică că serviciul are endpoint `/metrics` accesibil + că e pe `didi-network`.
**Loki nu primește loguri**: verifică `promtail` logs (`docker logs didi-promtail`) — probabil container labels nu match-uiesc.
**Jaeger fără traces**: verifică că serviciul are env `OTEL_EXPORTER_OTLP_ENDPOINT` setat corect + că face HTTP/gRPC către `didi-otel-collector:4317`.
**Alertmanager nu trimite email**: verifică `SENDGRID_API_KEY` în env + că from address `alerts@didi365.eu` e validat în SendGrid.
## Roadmap
- [ ] Adaugă **postgres_exporter** pentru metrici PostgreSQL (slow queries, replication lag)
- [ ] Adaugă **redis_exporter** pentru metrici Redis
- [ ] Adaugă **nvidia_gpu_exporter** pe GPU host pentru metrici VRAM/utilization
- [ ] Instrumentare agent-v3 (PR follow-up)
- [ ] Instrumentare ai_platform modules (PR follow-up)
- [ ] Slack/Teams webhook integration
- [ ] PagerDuty on-call rotation
- [ ] Synthetic monitoring (uptime checks pe didi365.eu)
## Referințe
- Caiet sarcini LOT 2 modul 8 (Observabilitate & Logging)
- Oferta EVOTECH §A.9 (diagrama observabilitate completă)
- Cercetare industrială §C (validare experimentală + KPI-uri)