- CI/CD GitLab (.gitlab-ci.yml): verify -> test -> build -> publish -> deploy cu health-gate (/api/v3/health/all) + rollback manual - Publicare pe retele sociale (LinkedIn/Facebook) din Analysis History (social.ts, linkedin.ts, facebook.ts, SocialPostModal) - Specificatii OpenAPI: agent-v3 (87 op.) si didi-framework (287 op.) - Matrice testare acceptanta beneficiar + script dovezi API - Fix nginx SSL: redirect port 3001 (error_page 497, absolute_redirect off) - Adrese interne inlocuite cu hostname-uri generice (.local) |
||
|---|---|---|
| .. | ||
| alertmanager | ||
| grafana | ||
| loki | ||
| otel-collector | ||
| prometheus | ||
| promtail | ||
| .env.example | ||
| docker-compose.yml | ||
| README.md | ||
DiDi Observability Stack
Stack complet de monitorizare, logging și tracing distribuit pentru platforma DiDi. Implementează cerința LOT 2 modul 8 (Observabilitate & Logging) din caietul de sarcini + diagrama A.9 din oferta EVOTECH.
Componente
| Component | Rol | URL UI (LAN) |
|---|---|---|
| Prometheus | Scraping metrici + alert rules | http://10.11.10.12:9090 |
| Grafana | Dashboard-uri vizualizare metrici + loguri + traces | http://10.11.10.12:3030 |
| Loki | Agregare loguri | http://10.11.10.12:3100 |
| Promtail | Shipper Docker logs → Loki | (no UI) |
| Jaeger | UI tracing distribuit | http://10.11.10.12:16686 |
| OTel Collector | Receiver OTLP (trace + metric) + processor + exporter | http://10.11.10.12:4319 (gRPC), :4320 (HTTP) |
| Alertmanager | Routing alerte (email, Slack/Teams, PagerDuty) | http://10.11.10.12:9093 |
Quick start
cd /home/admin365/didi_mono/didi_mono/backend/observability
# Setup .env
cp .env.example .env
$EDITOR .env # set GRAFANA_ADMIN_PASSWORD + SENDGRID_API_KEY
# Start stack
docker compose up -d
# Verify all healthy
docker compose ps
Apoi accesează Grafana la http://10.11.10.12:3030 (user admin, parola din GRAFANA_ADMIN_PASSWORD).
Datasources sunt deja provisioned (Prometheus, Loki, Jaeger). Adaugă dashboard-uri custom în grafana/dashboards/ (auto-provisioned la 30s).
Instrumentare servicii
Node.js (agent-v3, didi-framework, admin-dashboard)
Instalează prom-client + @opentelemetry/sdk-node:
npm install --save prom-client @opentelemetry/api @opentelemetry/sdk-node \
@opentelemetry/auto-instrumentations-node @opentelemetry/exporter-trace-otlp-grpc \
@opentelemetry/resources @opentelemetry/semantic-conventions
Adaugă în src/index.ts (înainte de orice alt import):
import { NodeSDK } from '@opentelemetry/sdk-node';
import { OTLPTraceExporter } from '@opentelemetry/exporter-trace-otlp-grpc';
import { getNodeAutoInstrumentations } from '@opentelemetry/auto-instrumentations-node';
import { Resource } from '@opentelemetry/resources';
import { SemanticResourceAttributes } from '@opentelemetry/semantic-conventions';
const sdk = new NodeSDK({
resource: new Resource({
[SemanticResourceAttributes.SERVICE_NAME]: 'agent-v3',
[SemanticResourceAttributes.SERVICE_VERSION]: '3.0.0',
}),
traceExporter: new OTLPTraceExporter({
url: process.env.OTEL_EXPORTER_OTLP_ENDPOINT || 'http://didi-otel-collector:4317',
}),
instrumentations: [getNodeAutoInstrumentations()],
});
sdk.start();
Adaugă endpoint /metrics în Express:
import { register, collectDefaultMetrics } from 'prom-client';
collectDefaultMetrics({ prefix: 'didi_agent_v3_' });
app.get('/metrics', async (req, res) => {
res.set('Content-Type', register.contentType);
res.end(await register.metrics());
});
Custom metrics relevante DiDi:
import { Counter, Histogram } from 'prom-client';
export const analysisCompleted = new Counter({
name: 'didi_analyses_completed_total',
help: 'Total number of analyses completed',
labelNames: ['component', 'tier', 'media_type', 'verdict'],
});
export const pipelineDuration = new Histogram({
name: 'didi_pipeline_duration_seconds',
help: 'Pipeline duration in seconds',
labelNames: ['component', 'tier'],
buckets: [1, 5, 10, 30, 60, 120, 300],
});
// În executor:
const end = pipelineDuration.startTimer({ component: 'techniques', tier });
try {
await runAnalysis();
analysisCompleted.inc({ component: 'techniques', tier, media_type, verdict });
} finally {
end();
}
Python (ai_platform modules)
Instalează prometheus_client + opentelemetry-instrumentation-fastapi:
pip install prometheus-client opentelemetry-api opentelemetry-sdk \
opentelemetry-exporter-otlp opentelemetry-instrumentation-fastapi
Adaugă în app.py:
from prometheus_client import make_asgi_app
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
from opentelemetry.exporter.otlp.proto.grpc.trace_exporter import OTLPSpanExporter
from opentelemetry.instrumentation.fastapi import FastAPIInstrumentor
from opentelemetry.sdk.resources import Resource
resource = Resource(attributes={"service.name": "llm-inference"})
trace.set_tracer_provider(TracerProvider(resource=resource))
trace.get_tracer_provider().add_span_processor(
BatchSpanProcessor(OTLPSpanExporter(endpoint="http://didi-otel-collector:4317", insecure=True))
)
app = FastAPI()
FastAPIInstrumentor.instrument_app(app)
# Mount /metrics
app.mount("/metrics", make_asgi_app())
Environment variables pe servicii
Adaugă în compose-urile fiecărui serviciu:
environment:
OTEL_EXPORTER_OTLP_ENDPOINT: http://didi-otel-collector:4317
OTEL_SERVICE_NAME: agent-v3
OTEL_RESOURCE_ATTRIBUTES: cluster=didi-prod,environment=production
Dashboards predefinite
Adaugă fișiere JSON în grafana/dashboards/ (auto-importate). Recomandate:
-
DiDi Platform Overview
- KPI cards: analize/min, error rate, P50/P95/P99 pipeline latency, GPU util
- Time-series: requests per service, queue backlog (RabbitMQ), Redis hit rate
-
AI Platform
- LLM inference latency per model, GPU memory (Qwen 3.5, BusterX)
- Brain analysis_atom cache hit rate, scheduler health
-
Cozi & Workers
- RabbitMQ queue depth per component × tier
- Worker throughput, retry count, DLQ messages
-
Cost & business
- Stripe webhook success rate, Sales Invoice rate, MRR proxy
-
Infrastructure
- CPU/RAM/Disk/Network per node, container restarts, healthcheck failures
Import dashboards exemple din comunitate:
- ID 11074 (Node Exporter Full)
- ID 13639 (Logs via Loki)
- ID 17761 (Cadvisor)
- ID 14570 (RabbitMQ Cluster)
În Grafana: + → Import → paste ID-ul → load.
Verificare end-to-end
# 1. Verifică Prometheus scrapes
curl -s http://10.11.10.12:9090/api/v1/targets | jq '.data.activeTargets[] | {job: .labels.job, health: .health}'
# 2. Verifică Loki primește loguri
curl -s 'http://10.11.10.12:3100/loki/api/v1/labels' | jq
# 3. Trimite un trace de test din shell (după ce ai un service instrumentat)
# Vizibil în Jaeger UI: http://10.11.10.12:16686
# 4. Trimite o alertă de test
curl -X POST http://10.11.10.12:9090/-/reload
# Așteaptă să se trigger ServiceDown alert (timer ~5min)
Retention
- Prometheus: 30 zile (configurabil în compose
--storage.tsdb.retention.time) - Loki: 7 zile (configurabil în
loki-config.yamlretention_period) - Jaeger (Badger storage): persistent, ~10GB cap
- Alertmanager: persistent state
SLA + escalation
Vezi alertmanager.yml pentru routing:
severity=critical→ notify imediat laoffice@clossers.comseverity=warning→ batch, repeat 12h- TODO: adaugă Slack/Teams webhook + PagerDuty key pentru on-call
Troubleshooting
Prometheus arată target-uri DOWN: verifică că serviciul are endpoint /metrics accesibil + că e pe didi-network.
Loki nu primește loguri: verifică promtail logs (docker logs didi-promtail) — probabil container labels nu match-uiesc.
Jaeger fără traces: verifică că serviciul are env OTEL_EXPORTER_OTLP_ENDPOINT setat corect + că face HTTP/gRPC către didi-otel-collector:4317.
Alertmanager nu trimite email: verifică SENDGRID_API_KEY în env + că from address alerts@didi365.eu e validat în SendGrid.
Roadmap
- Adaugă postgres_exporter pentru metrici PostgreSQL (slow queries, replication lag)
- Adaugă redis_exporter pentru metrici Redis
- Adaugă nvidia_gpu_exporter pe GPU host pentru metrici VRAM/utilization
- Instrumentare agent-v3 (PR follow-up)
- Instrumentare ai_platform modules (PR follow-up)
- Slack/Teams webhook integration
- PagerDuty on-call rotation
- Synthetic monitoring (uptime checks pe didi365.eu)
Referințe
- Caiet sarcini LOT 2 modul 8 (Observabilitate & Logging)
- Oferta EVOTECH §A.9 (diagrama observabilitate completă)
- Cercetare industrială §C (validare experimentală + KPI-uri)