11 KiB
Video Analysis API Documentation
REST API for video deepfake detection using semantic analysis via vLLM backends.
Note: This API depends on an external vLLM server configured via VIDEO_ANALYSIS_VLLM_BASE_URL.
Base URL
{BASE_URL}
Common configurations:
- Local development:
http://localhost:54600 - Docker (internal):
http://didiAI-video-api:54600 - Production: Use your configured hostname
Authentication
No authentication required.
Vision & Aggregation Models
Both endpoints are backed by a single vision model — BusterX (l8cv/BusterX_plusplus, served as busterx) @ port 54500:
| Stage | Model | Endpoint | Role |
|---|---|---|---|
| Deepfake verdict | BusterX (Qwen2.5-VL-7B fine-tune) | http://didiAI-video-vllm-buster:54500 |
REAL / FAKE / UNCERTAIN + explanation |
| Semantic chunk descriptions | BusterX (same endpoint) | http://didiAI-video-vllm-buster:54500 |
per-chunk description |
| Semantic aggregation | DIDI text LLM (Qwen3.5) | http://didiAI-llm-api:14011 |
merges chunk descriptions into final_summary |
BusterX is self-contained — Qwen2.5-VL is bundled inside the fine-tune, so no separate vision base model is loaded. There is no separate Qwen3-VL vision backend in this deployment.
Endpoints
Health Check
Check if the service is running.
GET /health
Response
{
"status": "ok"
}
Example
curl http://localhost:54600/health
Analyze Video
Upload a video for deepfake analysis.
POST /analyze/video
Request
Content-Type: multipart/form-data
| Field | Type | Required | Description |
|---|---|---|---|
file |
file | Yes | Video file to analyze |
Example
curl -X POST http://localhost:54600/analyze/video \
-F "file=@/path/to/video.mp4"
Response
{
"request_id": "550e8400-e29b-41d4-a716-446655440000",
"run_dir": "runs/550e8400-e29b-41d4-a716-446655440000",
"verdict": "FAKE",
"explanation": "The video shows clear signs of manipulation...",
"frames_analyzed": 16,
"evidence": [
{"frame_index": 0, "timestamp_s": 0.0},
{"frame_index": 30, "timestamp_s": 1.0}
],
"usage": {
"prompt_tokens": 1250,
"completion_tokens": 150,
"total_tokens": 1400
},
"latency_s": {
"sampling_time_s": 0.234,
"encode_time_s": 0.567,
"model_inference_time_s": 12.345
},
"meta": {
"fps": 30.0,
"total_frames": 450,
"duration_s": 15.0,
"sampled": 16,
"indices": [0, 30, 60, 90, 120, 150, 180, 210, 240, 270, 300, 330, 360, 390, 420, 449],
"timestamps_s": [0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, 14.97]
}
}
Response Fields
| Field | Description |
|---|---|
request_id |
UUID identifying the analysis request |
run_dir |
Directory where analysis artifacts are stored |
verdict |
Classification result: REAL, FAKE, or UNCERTAIN |
explanation |
Model explanation for the verdict |
usage |
Token usage statistics from the vLLM backend |
latency_s |
Timing breakdown in seconds |
meta |
Video metadata and frame sampling information |
Semantic Video Analysis
Upload a video for deep semantic analysis with temporal chunking.
POST /analyze/video/semantic
This endpoint performs comprehensive content understanding by:
- Dividing the video into temporal chunks (default: 10s each)
- Densely sampling each chunk (default: 24 frames per chunk)
- Analyzing each chunk with the vision LLM
- Optionally aggregating chunk descriptions into a coherent narrative
Use cases:
- Content description and understanding
- Action recognition and tracking
- Scene analysis
- Narrative extraction from video
Request
Content-Type: multipart/form-data
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
file |
file | Yes | - | Video file to analyze |
chunk_duration_s |
float | No | 10.0 | Duration of each chunk in seconds (1-60) |
frames_per_chunk |
int | No | 24 | Number of frames to sample per chunk (4-64) |
enable_aggregation |
bool | No | true | Whether to aggregate chunks into final summary |
Examples
Basic semantic analysis (default settings):
curl -X POST http://localhost:54600/analyze/video/semantic \
-F "file=@meeting.mp4"
Custom chunk settings:
curl -X POST http://localhost:54600/analyze/video/semantic \
-F "file=@meeting.mp4" \
-F "chunk_duration_s=5.0" \
-F "frames_per_chunk=32"
Without aggregation (get only chunk descriptions):
curl -X POST http://localhost:54600/analyze/video/semantic \
-F "file=@meeting.mp4" \
-F "enable_aggregation=false"
Response
{
"request_id": "abc-123-def-456",
"run_dir": "/app/runs/abc-123-def-456",
"analysis_type": "semantic",
"video_duration_s": 60.0,
"num_chunks": 6,
"chunk_results": [
{
"chunk_idx": 0,
"time_range": "0.0s - 10.0s",
"description": "A person in business attire enters a conference room, walks to the head of the table, and places a laptop down. The room has white walls and a large window showing daylight outside.",
"frames_analyzed": 24,
"inference_time_s": 12.3,
"usage": {
"prompt_tokens": 5234,
"completion_tokens": 87,
"total_tokens": 5321
}
},
{
"chunk_idx": 1,
"time_range": "10.0s - 20.0s",
"description": "The person opens the laptop and begins gesturing while speaking. Two other people enter the room and take seats at the conference table. One person carries a notebook.",
"frames_analyzed": 24,
"inference_time_s": 12.1,
"usage": {
"prompt_tokens": 5198,
"completion_tokens": 92,
"total_tokens": 5290
}
},
{
"chunk_idx": 2,
"time_range": "20.0s - 30.0s",
"description": "The presenter is now showing content on the laptop screen to the group. All three people are focused on the screen. One person is taking notes.",
"frames_analyzed": 24,
"inference_time_s": 11.8,
"usage": {
"prompt_tokens": 5201,
"completion_tokens": 78,
"total_tokens": 5279
}
},
{
"chunk_idx": 3,
"time_range": "30.0s - 40.0s",
"description": "Discussion is ongoing. The presenter is gesturing toward the screen. One attendee raises their hand and appears to ask a question.",
"frames_analyzed": 24,
"inference_time_s": 12.0,
"usage": {
"prompt_tokens": 5187,
"completion_tokens": 71,
"total_tokens": 5258
}
},
{
"chunk_idx": 4,
"time_range": "40.0s - 50.0s",
"description": "The presenter responds to the question with gestures. All participants are engaged in the discussion. Papers are visible on the table.",
"frames_analyzed": 24,
"inference_time_s": 11.9,
"usage": {
"prompt_tokens": 5209,
"completion_tokens": 68,
"total_tokens": 5277
}
},
{
"chunk_idx": 5,
"time_range": "50.0s - 60.0s",
"description": "The meeting appears to be concluding. Participants are gathering their belongings. The presenter closes the laptop and people begin standing up.",
"frames_analyzed": 24,
"inference_time_s": 12.2,
"usage": {
"prompt_tokens": 5223,
"completion_tokens": 75,
"total_tokens": 5298
}
}
],
"final_summary": "The video captures a business meeting in a conference room. It begins with a presenter setting up and two colleagues joining. The presenter delivers a presentation using a laptop, with the group discussing the content. One attendee asks questions and takes notes throughout. The meeting concludes with participants gathering their items and preparing to leave. The entire sequence lasts approximately 60 seconds.",
"aggregation_time_s": 4.5,
"total_latency_s": 76.8,
"meta": {
"fps": 30.0,
"total_frames": 1800,
"duration_s": 60.0,
"chunk_duration_s": 10.0,
"frames_per_chunk": 24,
"total_frames_sampled": 144
}
}
Response Fields
| Field | Description |
|---|---|
request_id |
UUID identifying the analysis request |
run_dir |
Directory where analysis artifacts are stored |
analysis_type |
Always "semantic" for this endpoint |
video_duration_s |
Total video duration in seconds |
num_chunks |
Number of temporal chunks processed |
chunk_results |
Array of results for each chunk (see below) |
final_summary |
Aggregated narrative (null if aggregation disabled) |
aggregation_time_s |
Time spent on aggregation (null if disabled) |
total_latency_s |
Total processing time in seconds |
meta |
Video metadata and sampling configuration |
ChunkResult Fields:
| Field | Description |
|---|---|
chunk_idx |
Zero-based chunk index |
time_range |
Temporal range of this chunk (e.g., "0.0s - 10.0s") |
description |
Semantic description of what happens in this chunk |
frames_analyzed |
Number of frames analyzed for this chunk |
inference_time_s |
Time spent on VLM inference for this chunk |
usage |
Token usage statistics for this chunk |
Performance Characteristics
For a 60-second video with default settings:
- Chunks: 6 (10s each)
- Total frames analyzed: 144 (24 per chunk)
- Coverage: ~8% of all frames (vs 0.89% for deepfake detection)
- Latency: ~77s total
- Chunk processing: ~72s (6 × 12s per chunk)
- Aggregation: ~5s
- Overhead (sampling, encoding): <1s
Optimization options:
| Setting | Fast | Balanced | Detailed |
|---|---|---|---|
chunk_duration_s |
20.0 | 10.0 | 5.0 |
frames_per_chunk |
16 | 24 | 32 |
| Coverage (60s video) | ~3% | ~8% | ~21% |
| Latency estimate | ~40s | ~77s | ~150s |
Service Info
Return service catalog metadata (resources, models, functions). Consumed by the DIDI catalog-api.
GET /v1/info
Example
curl http://localhost:54600/v1/info
Error Responses
| Status | Description |
|---|---|
400 |
Bad Request (missing or invalid file) |
500 |
Internal Server Error |
Possible 500 Error Causes
- vLLM backend unreachable
VIDEO_ANALYSIS_VLLM_BASE_URLnot set- Model inference failed
Example
{
"detail": "VIDEO_ANALYSIS_VLLM_BASE_URL is not set"
}
Minimal Python Client Example
import httpx
def analyze_video(video_path: str) -> dict:
with open(video_path, "rb") as f:
response = httpx.post(
"http://localhost:54600/analyze/video",
files={"file": f},
timeout=180.0,
)
response.raise_for_status()
return response.json()
result = analyze_video("suspect_video.mp4")
print(f"Verdict: {result['verdict']}")
print(f"Explanation: {result['explanation']}")