didi-lot1-ai/ai_platform/ENDPOINTS.md
EVOTECH IT SRL af0ec4a150 Arhitectură (diagrame) + imagini Docker module + contract furnizare + optimizări deploy
- docs/ARCHITECTURE.md (7 diagrame Mermaid) + architecture.html + diagrame originale (offer_diagram_1..7.png)
- artefacte_lot1/imagini_docker: imagini pre-construite pt toate cele 13 module (la zi)
- documentatie: Contract de furnizare nr.19
- optimizări deploy.sh/seed (MODELS_DIR, fix download python3); .gitignore (modele/.env/PV-uri excluse)
2026-06-29 16:23:03 +03:00

16 KiB
Raw Permalink Blame History

🚀 ML PROJECTS - Complete API Endpoints Reference

Last Updated: 2026-02-06 Status: All services operational

📊 Port Allocation Schema

Porturi alocate conform convenției datacenter:

  • 1xxxx = Production
  • 5xxxx = Development
  • x4xxx = AI/LLM Services
  • x1xxx = API Gateway

📊 Quick Status Overview

Production (1xxxx)

Service Port Description
Catalog API 11000 Main orchestrator gateway
Domain Check API 11000 (alias domain-check-api) T4 — source credibility (WHOIS/SSL/blacklist/DNS/IP/HTTP/mail-intel); isolated stack
vLLM Qwen3.5-35B-A3B 14001 Text + Vision LLM (MoE, GPU 0)
LLM Inference API 14011 Unified LLM router
Embeddings API 14100 OpenAI-compatible embeddings
- vLLM Embed Server 14101 GPU embedding backend
- llama.cpp Embed Server 14110 CPU/GGUF embedding backend
Rerank API 14200 Document reranking
- vLLM Rerank Server 14201 GPU reranking backend
- llama.cpp Rerank Server 14210 CPU/GGUF reranking backend

Development (5xxxx)

Service Port Description
Web API 51100 Fact-checking / web search
Embeddings API 54100 OpenAI-compatible embeddings
- vLLM Embed Server 54101 GPU embedding backend
- llama.cpp Embed Server 54110 CPU/GGUF embedding backend
Rerank API 54200 Document reranking
- vLLM Rerank Server 54201 GPU reranking backend
- llama.cpp Rerank Server 54210 CPU/GGUF reranking backend
Audio API 54300 Speech-to-text (Whisper)
BusterX vLLM 54500 Deepfake detection model
Video API 54600 Video analysis

🏭 PRODUCTION SERVICES (1xxxx)

1 Catalog API (Main Gateway)

Port: 11000
Base URL: http://localhost:11000
Purpose: Main orchestrator gateway - routes to all services
Decodare: 1+1+0+0+0 = Prod + API + Gateway

Endpoints

Method Path Description
GET /health Health check
GET /ready Readiness probe
GET /v1/status Status of all components

Quick Test

curl http://localhost:11000/health

2 vLLM Backend: Qwen3.5-35B-A3B

Port: 14001 Base URL: http://localhost:14001 Model: Qwen3.5-35B-A3B (MoE, native text + vision) GPU: GPU 0 (~57GB VRAM) Decodare: 1+4+0+0+1 = Prod + AI + TextInf + vLLM + instance1

Endpoints

Method Path Description
GET /v1/models List loaded model info
POST /v1/chat/completions OpenAI-compatible chat (text + vision)
POST /v1/completions Text completions
GET /health Health check
GET /version vLLM version

Quick Test

# Check model
curl http://localhost:14001/v1/models | jq '.data[0].id'
# Output: "qwen3.5"

# Text chat completion
curl -X POST http://localhost:14001/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5",
    "messages": [{"role": "user", "content": "Hello!"}],
    "max_tokens": 50
  }'

# Vision chat (with image URL)
curl -X POST http://localhost:14001/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}}
      ]
    }]
  }'

4 LLM Inference API (LLM Router)

Port: 14011
Base URL: http://localhost:14011
Purpose: Unified LLM inference gateway (routes to vLLM, LiteLLM, llamacpp)
Decodare: 1+4+0+1+1 = Prod + AI + TextInf + llama.cpp + instance1

Endpoints

Method Path Description
GET /health Health check
GET /ready Readiness probe
GET /v1/models List all models (all backends)
GET /v1/models?backend=vllm Filter by backend
GET /v1/backends List available backends
POST /v1/chat/completions Unified chat completions
POST /v1/models/load Load model (vLLM/llamacpp)
POST /v1/models/unload Unload model

Quick Test

curl http://localhost:14011/health
curl http://localhost:14011/v1/backends

5 Embeddings API

Port: 14100 (Prod) / 54100 (Dev) Base URL: http://localhost:14100 Purpose: OpenAI-compatible embeddings with multi-backend support (vLLM, llama.cpp) Model: BAAI/bge-m3 Decodare: 1+4+1+0+0 = Prod + AI + Embeddings

Backend Servers:

Port Server Description
14101 / 54101 vLLM Embed GPU-accelerated embedding server
14110 / 54110 llama.cpp Embed CPU/GGUF embedding server

Endpoints

Method Path Description
GET /health Health check (per-backend status)
GET /ready Readiness probe
GET /v1/models List embedding models
GET /v1/backends List available backends
POST /v1/embeddings Generate embeddings (OpenAI-compatible)

Quick Test

# Health check
curl http://localhost:14100/health

# Generate embeddings
curl -X POST http://localhost:14100/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "input": ["Hello world", "How are you?"],
    "model": "BAAI/bge-m3"
  }'

6 Rerank API

Port: 14200 (Prod) / 54200 (Dev) Base URL: http://localhost:14200 Purpose: Cohere/Jina-compatible document reranking with multi-backend support Model: BAAI/bge-reranker-v2-m3 Decodare: 1+4+2+0+0 = Prod + AI + Reranking

Backend Servers:

Port Server Description
14201 / 54201 vLLM Rerank GPU-accelerated reranking server
14210 / 54210 llama.cpp Rerank CPU/GGUF reranking server

Endpoints

Method Path Description
GET /health Health check (per-backend status)
GET /ready Readiness probe
GET /v1/models List reranking models
GET /v1/backends List available backends
POST /v1/rerank Rerank documents
POST /v2/rerank Rerank documents (v2 alias)

Quick Test

# Health check
curl http://localhost:14200/health

# Rerank documents
curl -X POST http://localhost:14200/v1/rerank \
  -H "Content-Type: application/json" \
  -d '{
    "model": "BAAI/bge-reranker-v2-m3",
    "query": "What is machine learning?",
    "documents": [
      "Machine learning is a subset of AI",
      "Cats are pets",
      "Deep learning uses neural networks"
    ]
  }'

🔧 DEVELOPMENT SERVICES (5xxxx)

7 Web API (Fact-checking)

Port: 51100
Base URL: http://localhost:51100
Purpose: Web search + evidence extraction for fact-checking
Decodare: 5+1+1+0+0 = Dev + API + Gateway + instance0

Endpoints

Method Path Description
GET /health Health check
GET /ready Readiness probe
POST /v1/gather Main: Complete fact-check pipeline
POST /v1/search Web search only (SearXNG free tier; provideri premium optionali: SerpAPI/Tavily/Brave/LinkUp via X-Search-Tier: premium)
POST /v1/fetch Fetch URLs with fallback

Quick Test

# Health check
curl http://localhost:51100/health

# Fact-check pipeline
curl -X POST http://localhost:51100/v1/gather \
  -H "Content-Type: application/json" \
  -d '{
    "claim": "Romania had highest GDP growth in EU 2024",
    "max_search_results": 5,
    "extract_snippets": true
  }' | jq '.evidence[0]'

8 Audio Transcription API

Port: 54300
Base URL: http://localhost:54300
Purpose: Speech-to-text using faster-whisper
Model: large-v3-turbo (int8 quantization)
GPU: GPU 0 (~2GB / 143GB VRAM)
Decodare: 5+4+3+0+0 = Dev + AI + Audio + vLLM + instance0

Endpoints

Method Path Description
GET /health Health check
GET /v1/models List Whisper models
POST /v1/audio/transcriptions Transcribe audio (OpenAI-compatible)

Quick Test

# Health check
curl http://localhost:54300/health

# Transcribe audio
curl -X POST http://localhost:54300/v1/audio/transcriptions \
  -F "file=@audio.mp3" \
  -F "response_format=json" \
  | jq '{text, language, duration}'

9 BusterX vLLM (Deepfake Vision)

Port: 54500
Base URL: http://localhost:54500
Model: BusterX (Qwen2.5-VL-7B fine-tuned for deepfake)
GPU: GPU 1 (22GB / 143GB VRAM)
Decodare: 5+4+5+0+0 = Dev + AI + Vision + vLLM + instance0

Endpoints

Method Path Description
GET /v1/models List loaded model info
POST /v1/chat/completions Deepfake detection
GET /health Health check

Quick Test

curl http://localhost:54500/v1/models | jq '.data[0].id'
# Used internally by video-analysis module

🔟 Video Analysis API

Port: 54600
Base URL: http://localhost:54600
Purpose: Deepfake detection + semantic video analysis
Decodare: 5+4+6+0+0 = Dev + AI + Video + instance0

Endpoints

Method Path Description
GET /health Health check
POST /analyze/video Deepfake detection (fast, 16 frames)
POST /analyze/video/semantic Semantic analysis (deep, 144+ frames)

Quick Test

# Health check
curl http://localhost:54600/health

# Deepfake detection
curl -X POST http://localhost:54600/analyze/video \
  -F "file=@video.mp4" \
  | jq '{verdict, explanation}'

# Semantic analysis
curl -X POST http://localhost:54600/analyze/video/semantic \
  -F "file=@video.mp4" \
  -F "chunk_duration_s=10.0" \
  -F "frames_per_chunk=24" \
  | jq '{num_chunks, final_summary}'

11 Extractors API (Feature Extraction)

Port: 54400
Base URL: http://localhost:54400
Purpose: Lightweight feature extraction: metadata, sentiment/OCR (delegate to LLM gateway), NER, object detection

Endpoints

Method Path Description
GET /health Health check
POST /v1/metadata EXIF / codec / spectrogram / ELA / integrity (C2PA)
POST /v1/sentiment Sentiment analysis (delegates to LLM gateway)
POST /v1/ocr OCR (delegates to Qwen3.5 vision via LLM gateway)
POST /v1/ner Named entity recognition (GLiNER, mdeberta backbone)
POST /v1/detect Object detection (YOLOv8n)

Quick Test

curl http://localhost:54400/health

12 Forensic Features API

Port: 8085
Base URL: http://localhost:8085
Purpose: Forensic detectors (rPPG, lip-sync, forgery heatmap, lighting) using MediaPipe

Endpoints

Method Path Description
GET /health Health check
POST /v1/analyze Run forensic detectors on a video/image

Quick Test

curl http://localhost:8085/health

13 Dashboard (Admin UI + API)

Port: 51300
Base URL: http://localhost:51300
Purpose: Admin monitoring UI + API: AI health, model catalog (DB-backed), runtime config, RBAC, cost/archive
Note: runs in STAGING MODE (DASHBOARD_STAGING_MODE=true) — Keycloak auth bypassed in delivered state.

Endpoints (selection)

Method Path Description
GET /health Health check
GET /api/stats/providers Provider status + live quota
GET /api/config/schema Runtime config schema (DB-overridable)
PUT /api/config/{key} Update runtime config (Bearer token required)
GET /audit Audit trail

Quick Test

curl http://localhost:51300/health

14 didi_brain API (Cache + RAG + Fact-checking)

Port: 8090
Base URL: http://localhost:8090
Purpose: Result cache (Postgres/pgvector) + RAG (Atomic) + fact-checking; consumed by backend

Endpoints (selection)

Method Path Description
GET /health Health check
POST /v1/gather Gather with cache (returns brain_meta.cache_status)
POST /v1/cache/invalidate Invalidate cache (rate-limited 10/h per actor)
GET /v1/fact_status/due_for_recheck Facts with next_check_at <= now, unlocked

Quick Test

curl http://localhost:8090/health

🧪 Complete Health Check Script

#!/bin/bash
echo "Testing all endpoints..."

# Production Services
echo "=== PRODUCTION ==="
curl -s http://localhost:11000/health && echo " ✓ Catalog API (11000)"
curl -s http://localhost:14001/health && echo " ✓ Qwen3.5-35B (14001)"
curl -s http://localhost:14011/health && echo " ✓ LLM API (14011)"
curl -s http://localhost:14100/health && echo " ✓ Embeddings API (14100)"
curl -s http://localhost:14200/health && echo " ✓ Rerank API (14200)"

# Development Services
echo "=== DEVELOPMENT ==="
curl -s http://localhost:51100/health && echo " ✓ Web API (51100)"
curl -s http://localhost:51300/health && echo " ✓ Dashboard (51300)"
curl -s http://localhost:54100/health && echo " ✓ Embeddings API (54100)"
curl -s http://localhost:54200/health && echo " ✓ Rerank API (54200)"
curl -s http://localhost:54300/health && echo " ✓ Audio API (54300)"
curl -s http://localhost:54400/health && echo " ✓ Extractors API (54400)"
curl -s http://localhost:54500/health && echo " ✓ BusterX (54500)"
curl -s http://localhost:54600/health && echo " ✓ Video API (54600)"
curl -s http://localhost:8085/health && echo " ✓ Forensic Features (8085)"
curl -s http://localhost:8090/health && echo " ✓ didi_brain (8090)"

echo ""
echo "All services operational ✅"

📈 GPU Allocation

GPU Model VRAM Used Total Utilization
GPU 0 Qwen3.5-35B-A3B (~57GB) + Whisper (~2GB) ~59GB 143GB 41%
GPU 1 BusterX / Qwen2.5-VL-7B (~22GB) + BAAI/bge-m3 + BAAI/bge-reranker-v2-m3 ~30GB 143GB 21%

🗺️ Port Map Summary

PRODUCTION (1xxxx):
├── 11000  Catalog API      (Main Gateway)
├── 14001  Qwen3.5-35B-A3B  (Text + Vision LLM)
├── 14011  LLM API          (LLM Router)
├── 14100  Embeddings API   (BGE-M3 Embeddings)
│   ├── 14101  vLLM Server
│   └── 14110  llama.cpp Server
├── 14200  Rerank API       (BGE Reranker)
│   ├── 14201  vLLM Server
│   └── 14210  llama.cpp Server

DEVELOPMENT (5xxxx):
├── 51100  Web API          (Fact-checking)
├── 51300  Dashboard        (Admin UI + API)
├── 54100  Embeddings API   (BGE-M3 Embeddings)
│   ├── 54101  vLLM Server
│   └── 54110  llama.cpp Server
├── 54200  Rerank API       (BGE Reranker)
│   ├── 54201  vLLM Server
│   └── 54210  llama.cpp Server
├── 54300  Audio API        (Whisper STT)
├── 54400  Extractors API   (metadata/sentiment/OCR/NER/detect)
├── 54500  BusterX          (Deepfake Vision)
├── 54600  Video API        (Video Analysis)
└── 8080   SearXNG (intern, nepublicat) (metasearch, free tier)

OTHER (outside 5-digit schema):
├── 8085   Forensic Features (rPPG/lip-sync/forgery)
├── 8090   didi_brain        (cache + RAG + fact-checking)
└── 11000  domain-check-api  (T4 — credibilitate sursă; stack izolat, doar pe didi-network)

  • LLM Inference: modules/llm-inference/API.md
  • Embeddings: modules/embeddings/API.md
  • Rerank: modules/rerank/API.md
  • Web (Fact-checking): modules/web/README.md
  • Video Analysis: modules/video-analysis/API.md
  • Audio Transcription: modules/audio/API.md
  • Domain Check (T4 — credibilitate sursă): modules/domain_check/use_api.mdPOST /api/v1/check/check pe domain-check-api:11000

📝 Notes

  • All Production services (1xxxx) are meant for external access
  • Development services (5xxxx) are for internal/testing use
  • All vLLM backends support OpenAI-compatible API
  • All services have health checks configured
  • Port schema follows datacenter convention for easy identification