didi-lot1-ai/ai_platform/modules/video-analysis/INDEX.md

100 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# video-analysis — INDEX
Video analysis service for DIDI. Performs deepfake detection and semantic temporal analysis on video content via uniform / chunked frame extraction plus a vision-language model backend (vLLM). Exposes a small FastAPI surface used by `agent-v3` for video media sessions.
- **Stack:** Python 3.11+, FastAPI, uvicorn, OpenCV (headless), Pillow, NumPy, httpx/requests, ffmpeg toolchain (via OpenCV), pydantic-settings + YAML
- **URL (Dev):** `http://10.11.10.12:54600`
- **Container:** runs on GPU host (`network_mode: host` in `deploy/docker-compose.yml`); the service itself is CPU-only — GPU is consumed by the upstream vLLM server
- **Vision backend:** external vLLM server (default: BusterX 7B @ port `54500`); same endpoint also drives semantic analysis. Optional alternative: Qwen3-VL-30B @ port `14002` for richer semantic narratives. The wider DIDI vision cascade (Qwen Vision local → Gemini Flash → GPT-4o) lives in `agent-v3`; this service only talks to one vLLM at a time.
## Ce face
- Receives a video file via `multipart/form-data` upload (no URL/path indirection — file bytes are POSTed)
- Extracts frames via OpenCV (`video_sampling.py`) with two strategies:
- **Uniform sampling** for fast deepfake check (default 16 frames over the whole video)
- **Temporal chunking** for semantic analysis (default 10s chunks × 24 frames per chunk)
- Encodes frames to JPEG (configurable `max_side`, `jpeg_quality`) and ships them to the vLLM server as base64 image payloads
- Runs vision-language model inference and returns:
- `verdict` (REAL / FAKE / UNCERTAIN) plus `explanation` for the deepfake endpoint
- per-chunk `description` array + optional aggregated `final_summary` (narrative) for the semantic endpoint
- Persists request artifacts (frames, prompts, responses) to a `runs/<request_id>/` directory for reproducibility/debugging
- Used by `didi-backend` agent-v3 for video media sessions (techniques + ai_tampered components)
## API endpoints
Defined in `src/video_analysis/app.py`:
| Endpoint | Method | Description |
|---|---|---|
| `/health` | GET | Liveness probe — `{"status":"ok"}` |
| `/analyze/video` | POST | Deepfake detection. Form field `file` (video). Fast path, ~16 frames. Returns `verdict`, `explanation`, `usage`, `latency_s`, `meta`. |
| `/analyze/video/semantic` | POST | Semantic temporal analysis. Form fields: `file`, `chunk_duration_s` (default 10.0, range 160), `frames_per_chunk` (default 24, range 464), `enable_aggregation` (default true). Returns `chunk_results[]`, optional `final_summary`, `meta`. |
Auth: none (called over private network / through Kong upstream by agent-v3).
## How didi-backend uses it
- agent-v3 video pipeline calls `/analyze/video` for fast deepfake screening and (where enabled) `/analyze/video/semantic` for chunked scene narration
- Output feeds the `ai_tampered` component (verdict + explanation) and contributes visual cues to the `techniques` component
- Frame metadata (`fps`, `duration_s`, `sampled`, `indices`, `timestamps_s`) is surfaced upstream so the agent can correlate detections with timestamps
- Long inference times (~77s for 60s video, semantic mode) are the reason agent-v3 routes video through the **async** session path (not the sync pipeline)
## Frame extraction logic
- **Deepfake path:** uniform interval sampling — `interval = total_frames / frames` where `frames` defaults to 16; min effectively 1 frame, capped by video length. Tunable via `VIDEO_ANALYSIS_FRAMES`.
- **Semantic path:** temporal chunking — video is split into `chunk_duration_s` slices, each slice gets `frames_per_chunk` uniformly sampled frames; total frames analyzed scales with duration (e.g. 60s @ defaults → 6 chunks × 24 = 144 frames, ~8% of source).
- Frames are downscaled so the longer side ≤ `max_side` (default 960 px), encoded JPEG at `jpeg_quality` (default 85), then base64-embedded into the chat-completions request.
- All sampled frames + indices + timestamps are returned in `meta` and persisted under `runs/<request_id>/`.
## Vision cascade
This service does **not** implement a multi-provider cascade. It is a thin client over a single vLLM endpoint configured at startup:
- **Primary (deepfake):** BusterX (Qwen2.5-VL-7B fine-tune, `l8cv/BusterX_plusplus`) at `VIDEO_ANALYSIS_VLLM_BASE_URL` — typically `http://didiAI-video-vllm-buster:54500` on the GPU host
- **Optional (semantic):** Qwen3-VL-30B at `http://didiAI-llm-vllm-vision:14002` — swap by editing `deploy/.env` and rebuilding
- **DIDI-wide cascade** (Qwen Vision local → OpenRouter Gemini Flash → GPT-4o) is implemented in agent-v3, NOT here. This service is a leaf node in that chain — agent-v3 calls it as one of several vision options.
- Service refuses to start if `VIDEO_ANALYSIS_VLLM_BASE_URL` is not set or the vLLM endpoint is unreachable (see `buster_client.py`, `settings.py`).
## Configuration
Env vars (prefix `VIDEO_ANALYSIS_`), loaded from `deploy/.env`:
| Variable | Required | Description |
|---|---|---|
| `VIDEO_ANALYSIS_VLLM_BASE_URL` | yes | Upstream vLLM server URL |
| `VIDEO_ANALYSIS_VLLM_MODEL` | yes | Model name passed to vLLM (`busterx`, `qwen3-vl`, `l8cv/BusterX_plusplus`, …) |
| `VIDEO_ANALYSIS_RUNS_DIR` | yes | Where to drop per-request artifacts (default `/app/runs` in container) |
| `VIDEO_ANALYSIS_EXTERNAL_URL` | yes | External URL embedded in the OpenAPI spec |
| `HF_TOKEN`, `HF_CACHE_DIR` | yes (when running bundled vLLM) | HuggingFace creds + shared cache for the vLLM container |
| `VIDEO_ANALYSIS_FRAMES` | no (default 16) | Uniform-sampling frame count |
| `VIDEO_ANALYSIS_MAX_SIDE` | no (default 960) | Frame downscale cap |
| `VIDEO_ANALYSIS_JPEG_QUALITY` | no (default 85) | JPEG quality 1100 |
| `VIDEO_ANALYSIS_MAX_TOKENS` | no (default 750) | Model response cap |
| `VIDEO_ANALYSIS_TEMPERATURE` | no (default 0.000001) | Near-deterministic decoding |
| `VIDEO_ANALYSIS_REPETITION_PENALTY` | no (default 1.05) | Repetition penalty |
| `NGINX_CONNECT_TIMEOUT` / `_SEND_TIMEOUT` / `_READ_TIMEOUT` | no | nginx upstream timeouts (only `api-nginx` profile) |
Tuning defaults live in `deploy/config.yaml`; env vars override YAML.
## Deployment
- GPU host required for the vLLM upstream (NVIDIA driver 535+, NVIDIA Container Toolkit, ≥16 GB VRAM). The video-analysis container itself is CPU-only.
- Compose lives in `deploy/`:
- `deploy/docker-compose.yml` — services + profiles (`api`, `api-nginx`)
- `deploy/Dockerfile` — Python 3.11 + OpenCV-headless + uv
- `deploy/deploy.sh` — wrapper around `docker compose` (loads env strictly from `deploy/.env`)
- `deploy/nginx.conf` / `nginx.conf.template` — optional reverse proxy
- Typical bring-up:
- `cp .env.example deploy/.env && $EDITOR deploy/.env`
- `cd deploy && ./deploy.sh --profile api --detach`
- Port allocation: `54600` video-analysis API, `54500` BusterX vLLM (Dev).
- Restart: `cd deploy && docker compose restart video-analysis-api` (or `docker restart video_analysis`).
## Related
- **agent-v3 video pipeline** — `/home/admin365/didi_mono/backend/services/orchestration-layer/agent-v3` is the consumer; orchestrates async video sessions and merges this service's verdict into `ai_tampered` + `techniques` results
- **BusterX vLLM** (port `54500`) — sibling service in the AI platform; the actual GPU-backed model that this service queries (referenced in main `CLAUDE.md` ports section)
- **Qwen3-VL vision vLLM** (port `14002`) — alternative semantic backend (`didiAI-llm-vllm-vision`)
- **AI platform shared assets** — `../../README.md`, `../../ruff.toml`
- **Internal package layout:** `src/video_analysis/{app.py, buster_client.py, schemas.py, settings.py, video_sampling.py}`
- **Sibling docs:** `README.md`, `API.md`, `TESTING.md` in this folder