# πŸš€ ML PROJECTS - Complete API Endpoints Reference > Last Updated: 2026-02-06 > Status: All services operational βœ… ## πŸ“Š Port Allocation Schema Porturi alocate conform convenΘ›iei datacenter: - **1xxxx** = Production - **5xxxx** = Development - **x4xxx** = AI/LLM Services - **x1xxx** = API Gateway --- ## πŸ“Š Quick Status Overview ### Production (1xxxx) | Service | Port | Description | |---------|------|-------------| | **Catalog API** | 11000 | Main orchestrator gateway | | **Domain Check API** | 11000 (alias `domain-check-api`) | T4 β€” source credibility (WHOIS/SSL/blacklist/DNS/IP/HTTP/mail-intel); isolated stack | | **vLLM Qwen3.5-35B-A3B** | 14001 | Text + Vision LLM (MoE, GPU 0) | | **LLM Inference API** | 14011 | Unified LLM router | | **Embeddings API** | 14100 | OpenAI-compatible embeddings | | - vLLM Embed Server | 14101 | GPU embedding backend | | - llama.cpp Embed Server | 14110 | CPU/GGUF embedding backend | | **Rerank API** | 14200 | Document reranking | | - vLLM Rerank Server | 14201 | GPU reranking backend | | - llama.cpp Rerank Server | 14210 | CPU/GGUF reranking backend | ### Development (5xxxx) | Service | Port | Description | |---------|------|-------------| | **Web API** | 51100 | Fact-checking / web search | | **Embeddings API** | 54100 | OpenAI-compatible embeddings | | - vLLM Embed Server | 54101 | GPU embedding backend | | - llama.cpp Embed Server | 54110 | CPU/GGUF embedding backend | | **Rerank API** | 54200 | Document reranking | | - vLLM Rerank Server | 54201 | GPU reranking backend | | - llama.cpp Rerank Server | 54210 | CPU/GGUF reranking backend | | **Audio API** | 54300 | Speech-to-text (Whisper) | | **BusterX vLLM** | 54500 | Deepfake detection model | | **Video API** | 54600 | Video analysis | --- ## 🏭 PRODUCTION SERVICES (1xxxx) ### 1️⃣ Catalog API (Main Gateway) **Port:** `11000` **Base URL:** `http://localhost:11000` **Purpose:** Main orchestrator gateway - routes to all services **Decodare:** 1+1+0+0+0 = Prod + API + Gateway #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `GET` | `/ready` | Readiness probe | | `GET` | `/v1/status` | Status of all components | #### Quick Test ```bash curl http://localhost:11000/health ``` --- ### 2️⃣ vLLM Backend: Qwen3.5-35B-A3B **Port:** `14001` **Base URL:** `http://localhost:14001` **Model:** Qwen3.5-35B-A3B (MoE, native text + vision) **GPU:** GPU 0 (~57GB VRAM) **Decodare:** 1+4+0+0+1 = Prod + AI + TextInf + vLLM + instance1 #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/v1/models` | List loaded model info | | `POST` | `/v1/chat/completions` | OpenAI-compatible chat (text + vision) | | `POST` | `/v1/completions` | Text completions | | `GET` | `/health` | Health check | | `GET` | `/version` | vLLM version | #### Quick Test ```bash # Check model curl http://localhost:14001/v1/models | jq '.data[0].id' # Output: "qwen3.5" # Text chat completion curl -X POST http://localhost:14001/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.5", "messages": [{"role": "user", "content": "Hello!"}], "max_tokens": 50 }' # Vision chat (with image URL) curl -X POST http://localhost:14001/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "qwen3.5", "messages": [{ "role": "user", "content": [ {"type": "text", "text": "What is in this image?"}, {"type": "image_url", "image_url": {"url": "data:image/jpeg;base64,..."}} ] }] }' ``` --- ### 4️⃣ LLM Inference API (LLM Router) **Port:** `14011` **Base URL:** `http://localhost:14011` **Purpose:** Unified LLM inference gateway (routes to vLLM, LiteLLM, llamacpp) **Decodare:** 1+4+0+1+1 = Prod + AI + TextInf + llama.cpp + instance1 #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `GET` | `/ready` | Readiness probe | | `GET` | `/v1/models` | List all models (all backends) | | `GET` | `/v1/models?backend=vllm` | Filter by backend | | `GET` | `/v1/backends` | List available backends | | `POST` | `/v1/chat/completions` | Unified chat completions | | `POST` | `/v1/models/load` | Load model (vLLM/llamacpp) | | `POST` | `/v1/models/unload` | Unload model | #### Quick Test ```bash curl http://localhost:14011/health curl http://localhost:14011/v1/backends ``` --- ### 5️⃣ Embeddings API **Port:** `14100` (Prod) / `54100` (Dev) **Base URL:** `http://localhost:14100` **Purpose:** OpenAI-compatible embeddings with multi-backend support (vLLM, llama.cpp) **Model:** BAAI/bge-m3 **Decodare:** 1+4+1+0+0 = Prod + AI + Embeddings **Backend Servers:** | Port | Server | Description | |------|--------|-------------| | 14101 / 54101 | vLLM Embed | GPU-accelerated embedding server | | 14110 / 54110 | llama.cpp Embed | CPU/GGUF embedding server | #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check (per-backend status) | | `GET` | `/ready` | Readiness probe | | `GET` | `/v1/models` | List embedding models | | `GET` | `/v1/backends` | List available backends | | `POST` | `/v1/embeddings` | **Generate embeddings** (OpenAI-compatible) | #### Quick Test ```bash # Health check curl http://localhost:14100/health # Generate embeddings curl -X POST http://localhost:14100/v1/embeddings \ -H "Content-Type: application/json" \ -d '{ "input": ["Hello world", "How are you?"], "model": "BAAI/bge-m3" }' ``` --- ### 6️⃣ Rerank API **Port:** `14200` (Prod) / `54200` (Dev) **Base URL:** `http://localhost:14200` **Purpose:** Cohere/Jina-compatible document reranking with multi-backend support **Model:** BAAI/bge-reranker-v2-m3 **Decodare:** 1+4+2+0+0 = Prod + AI + Reranking **Backend Servers:** | Port | Server | Description | |------|--------|-------------| | 14201 / 54201 | vLLM Rerank | GPU-accelerated reranking server | | 14210 / 54210 | llama.cpp Rerank | CPU/GGUF reranking server | #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check (per-backend status) | | `GET` | `/ready` | Readiness probe | | `GET` | `/v1/models` | List reranking models | | `GET` | `/v1/backends` | List available backends | | `POST` | `/v1/rerank` | **Rerank documents** | | `POST` | `/v2/rerank` | Rerank documents (v2 alias) | #### Quick Test ```bash # Health check curl http://localhost:14200/health # Rerank documents curl -X POST http://localhost:14200/v1/rerank \ -H "Content-Type: application/json" \ -d '{ "model": "BAAI/bge-reranker-v2-m3", "query": "What is machine learning?", "documents": [ "Machine learning is a subset of AI", "Cats are pets", "Deep learning uses neural networks" ] }' ``` --- ## πŸ”§ DEVELOPMENT SERVICES (5xxxx) ### 7️⃣ Web API (Fact-checking) **Port:** `51100` **Base URL:** `http://localhost:51100` **Purpose:** Web search + evidence extraction for fact-checking **Decodare:** 5+1+1+0+0 = Dev + API + Gateway + instance0 #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `GET` | `/ready` | Readiness probe | | `POST` | `/v1/gather` | **Main: Complete fact-check pipeline** | | `POST` | `/v1/search` | Web search only (SearXNG free tier; provideri premium optionali: SerpAPI/Tavily/Brave/LinkUp via `X-Search-Tier: premium`) | | `POST` | `/v1/fetch` | Fetch URLs with fallback | #### Quick Test ```bash # Health check curl http://localhost:51100/health # Fact-check pipeline curl -X POST http://localhost:51100/v1/gather \ -H "Content-Type: application/json" \ -d '{ "claim": "Romania had highest GDP growth in EU 2024", "max_search_results": 5, "extract_snippets": true }' | jq '.evidence[0]' ``` --- ### 8️⃣ Audio Transcription API **Port:** `54300` **Base URL:** `http://localhost:54300` **Purpose:** Speech-to-text using faster-whisper **Model:** large-v3-turbo (int8 quantization) **GPU:** GPU 0 (~2GB / 143GB VRAM) **Decodare:** 5+4+3+0+0 = Dev + AI + Audio + vLLM + instance0 #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `GET` | `/v1/models` | List Whisper models | | `POST` | `/v1/audio/transcriptions` | **Transcribe audio** (OpenAI-compatible) | #### Quick Test ```bash # Health check curl http://localhost:54300/health # Transcribe audio curl -X POST http://localhost:54300/v1/audio/transcriptions \ -F "file=@audio.mp3" \ -F "response_format=json" \ | jq '{text, language, duration}' ``` --- ### 9️⃣ BusterX vLLM (Deepfake Vision) **Port:** `54500` **Base URL:** `http://localhost:54500` **Model:** BusterX (Qwen2.5-VL-7B fine-tuned for deepfake) **GPU:** GPU 1 (22GB / 143GB VRAM) **Decodare:** 5+4+5+0+0 = Dev + AI + Vision + vLLM + instance0 #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/v1/models` | List loaded model info | | `POST` | `/v1/chat/completions` | Deepfake detection | | `GET` | `/health` | Health check | #### Quick Test ```bash curl http://localhost:54500/v1/models | jq '.data[0].id' # Used internally by video-analysis module ``` --- ### πŸ”Ÿ Video Analysis API **Port:** `54600` **Base URL:** `http://localhost:54600` **Purpose:** Deepfake detection + semantic video analysis **Decodare:** 5+4+6+0+0 = Dev + AI + Video + instance0 #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `POST` | `/analyze/video` | **Deepfake detection** (fast, 16 frames) | | `POST` | `/analyze/video/semantic` | **Semantic analysis** (deep, 144+ frames) | #### Quick Test ```bash # Health check curl http://localhost:54600/health # Deepfake detection curl -X POST http://localhost:54600/analyze/video \ -F "file=@video.mp4" \ | jq '{verdict, explanation}' # Semantic analysis curl -X POST http://localhost:54600/analyze/video/semantic \ -F "file=@video.mp4" \ -F "chunk_duration_s=10.0" \ -F "frames_per_chunk=24" \ | jq '{num_chunks, final_summary}' ``` --- ### 1️⃣1️⃣ Extractors API (Feature Extraction) **Port:** `54400` **Base URL:** `http://localhost:54400` **Purpose:** Lightweight feature extraction: metadata, sentiment/OCR (delegate to LLM gateway), NER, object detection #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `POST` | `/v1/metadata` | EXIF / codec / spectrogram / ELA / integrity (C2PA) | | `POST` | `/v1/sentiment` | Sentiment analysis (delegates to LLM gateway) | | `POST` | `/v1/ocr` | OCR (delegates to Qwen3.5 vision via LLM gateway) | | `POST` | `/v1/ner` | Named entity recognition (GLiNER, mdeberta backbone) | | `POST` | `/v1/detect` | Object detection (YOLOv8n) | #### Quick Test ```bash curl http://localhost:54400/health ``` --- ### 1️⃣2️⃣ Forensic Features API **Port:** `8085` **Base URL:** `http://localhost:8085` **Purpose:** Forensic detectors (rPPG, lip-sync, forgery heatmap, lighting) using MediaPipe #### Endpoints | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `POST` | `/v1/analyze` | Run forensic detectors on a video/image | #### Quick Test ```bash curl http://localhost:8085/health ``` --- ### 1️⃣3️⃣ Dashboard (Admin UI + API) **Port:** `51300` **Base URL:** `http://localhost:51300` **Purpose:** Admin monitoring UI + API: AI health, model catalog (DB-backed), runtime config, RBAC, cost/archive **Note:** runs in STAGING MODE (`DASHBOARD_STAGING_MODE=true`) β€” Keycloak auth bypassed in delivered state. #### Endpoints (selection) | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `GET` | `/api/stats/providers` | Provider status + live quota | | `GET` | `/api/config/schema` | Runtime config schema (DB-overridable) | | `PUT` | `/api/config/{key}` | Update runtime config (Bearer token required) | | `GET` | `/audit` | Audit trail | #### Quick Test ```bash curl http://localhost:51300/health ``` --- ### 1️⃣4️⃣ didi_brain API (Cache + RAG + Fact-checking) **Port:** `8090` **Base URL:** `http://localhost:8090` **Purpose:** Result cache (Postgres/pgvector) + RAG (Atomic) + fact-checking; consumed by backend #### Endpoints (selection) | Method | Path | Description | |--------|------|-------------| | `GET` | `/health` | Health check | | `POST` | `/v1/gather` | Gather with cache (returns `brain_meta.cache_status`) | | `POST` | `/v1/cache/invalidate` | Invalidate cache (rate-limited 10/h per actor) | | `GET` | `/v1/fact_status/due_for_recheck` | Facts with `next_check_at <= now`, unlocked | #### Quick Test ```bash curl http://localhost:8090/health ``` --- ## πŸ§ͺ Complete Health Check Script ```bash #!/bin/bash echo "Testing all endpoints..." # Production Services echo "=== PRODUCTION ===" curl -s http://localhost:11000/health && echo " βœ“ Catalog API (11000)" curl -s http://localhost:14001/health && echo " βœ“ Qwen3.5-35B (14001)" curl -s http://localhost:14011/health && echo " βœ“ LLM API (14011)" curl -s http://localhost:14100/health && echo " βœ“ Embeddings API (14100)" curl -s http://localhost:14200/health && echo " βœ“ Rerank API (14200)" # Development Services echo "=== DEVELOPMENT ===" curl -s http://localhost:51100/health && echo " βœ“ Web API (51100)" curl -s http://localhost:51300/health && echo " βœ“ Dashboard (51300)" curl -s http://localhost:54100/health && echo " βœ“ Embeddings API (54100)" curl -s http://localhost:54200/health && echo " βœ“ Rerank API (54200)" curl -s http://localhost:54300/health && echo " βœ“ Audio API (54300)" curl -s http://localhost:54400/health && echo " βœ“ Extractors API (54400)" curl -s http://localhost:54500/health && echo " βœ“ BusterX (54500)" curl -s http://localhost:54600/health && echo " βœ“ Video API (54600)" curl -s http://localhost:8085/health && echo " βœ“ Forensic Features (8085)" curl -s http://localhost:8090/health && echo " βœ“ didi_brain (8090)" echo "" echo "All services operational βœ…" ``` --- ## πŸ“ˆ GPU Allocation | GPU | Model | VRAM Used | Total | Utilization | |-----|-------|-----------|-------|-------------| | **GPU 0** | Qwen3.5-35B-A3B (~57GB) + Whisper (~2GB) | ~59GB | 143GB | 41% | | **GPU 1** | BusterX / Qwen2.5-VL-7B (~22GB) + BAAI/bge-m3 + BAAI/bge-reranker-v2-m3 | ~30GB | 143GB | 21% | --- ## πŸ—ΊοΈ Port Map Summary ``` PRODUCTION (1xxxx): β”œβ”€β”€ 11000 Catalog API (Main Gateway) β”œβ”€β”€ 14001 Qwen3.5-35B-A3B (Text + Vision LLM) β”œβ”€β”€ 14011 LLM API (LLM Router) β”œβ”€β”€ 14100 Embeddings API (BGE-M3 Embeddings) β”‚ β”œβ”€β”€ 14101 vLLM Server β”‚ └── 14110 llama.cpp Server β”œβ”€β”€ 14200 Rerank API (BGE Reranker) β”‚ β”œβ”€β”€ 14201 vLLM Server β”‚ └── 14210 llama.cpp Server DEVELOPMENT (5xxxx): β”œβ”€β”€ 51100 Web API (Fact-checking) β”œβ”€β”€ 51300 Dashboard (Admin UI + API) β”œβ”€β”€ 54100 Embeddings API (BGE-M3 Embeddings) β”‚ β”œβ”€β”€ 54101 vLLM Server β”‚ └── 54110 llama.cpp Server β”œβ”€β”€ 54200 Rerank API (BGE Reranker) β”‚ β”œβ”€β”€ 54201 vLLM Server β”‚ └── 54210 llama.cpp Server β”œβ”€β”€ 54300 Audio API (Whisper STT) β”œβ”€β”€ 54400 Extractors API (metadata/sentiment/OCR/NER/detect) β”œβ”€β”€ 54500 BusterX (Deepfake Vision) β”œβ”€β”€ 54600 Video API (Video Analysis) └── 8080 SearXNG (intern, nepublicat) (metasearch, free tier) OTHER (outside 5-digit schema): β”œβ”€β”€ 8085 Forensic Features (rPPG/lip-sync/forgery) β”œβ”€β”€ 8090 didi_brain (cache + RAG + fact-checking) └── 11000 domain-check-api (T4 β€” credibilitate sursΔƒ; stack izolat, doar pe didi-network) ``` --- ## πŸ”— API Documentation Links - **LLM Inference:** `modules/llm-inference/API.md` - **Embeddings:** `modules/embeddings/API.md` - **Rerank:** `modules/rerank/API.md` - **Web (Fact-checking):** `modules/web/README.md` - **Video Analysis:** `modules/video-analysis/API.md` - **Audio Transcription:** `modules/audio/API.md` - **Domain Check (T4 β€” credibilitate sursΔƒ):** `modules/domain_check/use_api.md` β€” `POST /api/v1/check/check` pe `domain-check-api:11000` --- ## πŸ“ Notes - All Production services (1xxxx) are meant for external access - Development services (5xxxx) are for internal/testing use - All vLLM backends support OpenAI-compatible API - All services have health checks configured - Port schema follows datacenter convention for easy identification