6.1 KiB
Video Analysis - Testing Checklist
⚠️ NOTĂ: Acest fișier este un plan EXPLORATORIU / R&D, NU configurația livrată. Modelul
Qwen3-VL-30B(port 14002 /gpt-oss-120b/deploy-llm-api-1) nu există în deployment-ul livrat. Pipeline-ul REAL: deepfake + semantic folosesc BusterX (busterx, vLLMdidiAI-video-vllm-buster:54500), iar agregarea semantică pe text folosește Qwen3.5 (didiAI-llm-api:14011). Pentru testarea sistemului livrat veziai_platform/local_gpu_stack/TESTING.md.
R&D (opțional): Model Comparison for Semantic Analysis
Background
Sistemul livrat folosește BusterX (Qwen2.5-VL-7B) atât pentru deepfake cât și pentru cadrele din analiza semantică, cu agregare text pe Qwen3.5. Ca direcție de cercetare, s-ar putea evalua un model vision mai mare pentru partea semantică (dacă va fi disponibil în viitor).
Hypothesis
Analiza semantică (descriere scenă, narativ) ar putea beneficia de un model vision mai mare, în timp ce detecția deepfake rămâne pe modelul specializat BusterX.
Test Plan: Qwen3-VL-30B for Semantic Analysis
Current Configuration
# modules/video-analysis/deploy/.env
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-video-vllm-buster:54500 # Port 54500
VIDEO_ANALYSIS_VLLM_MODEL=busterx # 7B parameters
Test Configuration
# modules/video-analysis/deploy/.env
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-llm-vllm-vision:14002 # Port 14002
VIDEO_ANALYSIS_VLLM_MODEL=qwen3-vl # 30B parameters
Testing Procedure
Step 1: Baseline Test (BusterX 7B)
# Current configuration - no changes needed
cd /home/vasi/ml-projects/modules/video-analysis/deploy
# Test semantic analysis with BusterX
curl -X POST http://localhost:54600/analyze/video/semantic \
-F "file=@test_video_60s.mp4" \
-F "chunk_duration_s=10.0" \
-F "frames_per_chunk=24" \
-o baseline_busterx.json
# Review results
jq '{
model: "BusterX-7B",
num_chunks: .num_chunks,
chunk_descriptions: [.chunk_results[].description],
final_summary: .final_summary,
total_latency_s: .total_latency_s
}' baseline_busterx.json
Step 2: Test with Qwen3-VL-30B
cd /home/vasi/ml-projects/modules/video-analysis/deploy
# Backup current config
cp .env .env.backup
# Update to use Qwen3-VL
cat > .env << 'ENVFILE'
HF_TOKEN=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
HF_CACHE_DIR=/cai2_ds_storage/hf_cache
VIDEO_ANALYSIS_RUNS_DIR=/app/runs
# Switch to Qwen3-VL for testing
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-llm-vllm-vision:14002
VIDEO_ANALYSIS_VLLM_MODEL=qwen3-vl
# Semantic analysis settings
VIDEO_ANALYSIS_SEMANTIC_CHUNK_DURATION_S=10.0
VIDEO_ANALYSIS_SEMANTIC_FRAMES_PER_CHUNK=24
VIDEO_ANALYSIS_SEMANTIC_ENABLE_AGGREGATION=true
VIDEO_ANALYSIS_SEMANTIC_AGGREGATION_MODEL=gpt-oss-120b
VIDEO_ANALYSIS_SEMANTIC_LLM_BASE_URL=http://deploy-llm-api-1:8100
ENVFILE
# Rebuild and restart
docker compose build video-analysis-api
docker compose up -d video-analysis-api
# Wait for startup
sleep 10
# Test with same video
curl -X POST http://localhost:54600/analyze/video/semantic \
-F "file=@test_video_60s.mp4" \
-F "chunk_duration_s=10.0" \
-F "frames_per_chunk=24" \
-o test_qwen3vl.json
# Review results
jq '{
model: "Qwen3-VL-30B",
num_chunks: .num_chunks,
chunk_descriptions: [.chunk_results[].description],
final_summary: .final_summary,
total_latency_s: .total_latency_s
}' test_qwen3vl.json
Step 3: Compare Results
# Side-by-side comparison
echo "=== BusterX 7B ==="
jq -r '.chunk_results[0].description' baseline_busterx.json
echo ""
echo "=== Qwen3-VL 30B ==="
jq -r '.chunk_results[0].description' test_qwen3vl.json
echo ""
echo "=== Final Summaries ==="
echo "BusterX: $(jq -r '.final_summary' baseline_busterx.json)"
echo ""
echo "Qwen3-VL: $(jq -r '.final_summary' test_qwen3vl.json)"
Evaluation Criteria
Quality Metrics
-
Description Detail
- More specific object identification
- Better action recognition
- More context understanding
-
Narrative Coherence
- Logical flow between segments
- Temporal consistency
- Better story understanding
-
Accuracy
- Correct identification of people/objects
- Accurate scene descriptions
- Proper action sequences
Performance Metrics
| Metric | BusterX 7B | Qwen3-VL 30B | Difference |
|---|---|---|---|
| Latency per chunk | ~12s | ~??s | ?? |
| Total latency (60s video) | ~77s | ~??s | ?? |
| Token usage per chunk | ~5200 | ~????? | ?? |
Expected Outcomes
If Qwen3-VL is Better:
Action: Update default configuration to use Qwen3-VL for semantic analysis
# Keep two separate configs:
# 1. Deepfake endpoint → BusterX (specialized)
# 2. Semantic endpoint → Qwen3-VL (better understanding)
Implementation:
- Add
VIDEO_ANALYSIS_DEEPFAKE_MODELandVIDEO_ANALYSIS_SEMANTIC_MODELsettings - Configure different models per endpoint
If BusterX is Sufficient:
Action: Keep current configuration, document findings
Rationale:
- Latency advantage (30B model is slower)
- VRAM savings
- BusterX might be sufficient for semantic tasks
Test Videos
Suggested test scenarios:
-
Meeting/Conference (60s)
- Multiple people
- Complex interactions
- Scene changes
-
Action Sequence (60s)
- Fast movements
- Object manipulation
- Environmental changes
-
Indoor/Outdoor Transition (60s)
- Lighting changes
- Multiple scenes
- Context shifts
Rollback Procedure
If issues occur:
cd /home/vasi/ml-projects/modules/video-analysis/deploy
# Restore original config
cp .env.backup .env
# Rebuild
docker compose build video-analysis-api
docker compose up -d video-analysis-api
Status
- Baseline test completed (BusterX)
- Test with Qwen3-VL completed
- Results compared
- Decision made
- Configuration updated (if needed)
- Documentation updated
Notes
Add observations here during testing:
Date: ___________
Tester: _________
Observations:
Recommendation: