didi-lot1-ai/ai_platform/modules/video-analysis/TESTING.md

6.1 KiB

Video Analysis - Testing Checklist

⚠️ NOTĂ: Acest fișier este un plan EXPLORATORIU / R&D, NU configurația livrată. Modelul Qwen3-VL-30B (port 14002 / gpt-oss-120b / deploy-llm-api-1) nu există în deployment-ul livrat. Pipeline-ul REAL: deepfake + semantic folosesc BusterX (busterx, vLLM didiAI-video-vllm-buster:54500), iar agregarea semantică pe text folosește Qwen3.5 (didiAI-llm-api:14011). Pentru testarea sistemului livrat vezi ai_platform/local_gpu_stack/TESTING.md.

R&D (opțional): Model Comparison for Semantic Analysis

Background

Sistemul livrat folosește BusterX (Qwen2.5-VL-7B) atât pentru deepfake cât și pentru cadrele din analiza semantică, cu agregare text pe Qwen3.5. Ca direcție de cercetare, s-ar putea evalua un model vision mai mare pentru partea semantică (dacă va fi disponibil în viitor).

Hypothesis

Analiza semantică (descriere scenă, narativ) ar putea beneficia de un model vision mai mare, în timp ce detecția deepfake rămâne pe modelul specializat BusterX.


Test Plan: Qwen3-VL-30B for Semantic Analysis

Current Configuration

# modules/video-analysis/deploy/.env
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-video-vllm-buster:54500  # Port 54500
VIDEO_ANALYSIS_VLLM_MODEL=busterx                      # 7B parameters

Test Configuration

# modules/video-analysis/deploy/.env
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-llm-vllm-vision:14002  # Port 14002
VIDEO_ANALYSIS_VLLM_MODEL=qwen3-vl                     # 30B parameters

Testing Procedure

Step 1: Baseline Test (BusterX 7B)

# Current configuration - no changes needed
cd /home/vasi/ml-projects/modules/video-analysis/deploy

# Test semantic analysis with BusterX
curl -X POST http://localhost:54600/analyze/video/semantic \
  -F "file=@test_video_60s.mp4" \
  -F "chunk_duration_s=10.0" \
  -F "frames_per_chunk=24" \
  -o baseline_busterx.json

# Review results
jq '{
  model: "BusterX-7B",
  num_chunks: .num_chunks,
  chunk_descriptions: [.chunk_results[].description],
  final_summary: .final_summary,
  total_latency_s: .total_latency_s
}' baseline_busterx.json

Step 2: Test with Qwen3-VL-30B

cd /home/vasi/ml-projects/modules/video-analysis/deploy

# Backup current config
cp .env .env.backup

# Update to use Qwen3-VL
cat > .env << 'ENVFILE'
HF_TOKEN=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
HF_CACHE_DIR=/cai2_ds_storage/hf_cache
VIDEO_ANALYSIS_RUNS_DIR=/app/runs

# Switch to Qwen3-VL for testing
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-llm-vllm-vision:14002
VIDEO_ANALYSIS_VLLM_MODEL=qwen3-vl

# Semantic analysis settings
VIDEO_ANALYSIS_SEMANTIC_CHUNK_DURATION_S=10.0
VIDEO_ANALYSIS_SEMANTIC_FRAMES_PER_CHUNK=24
VIDEO_ANALYSIS_SEMANTIC_ENABLE_AGGREGATION=true
VIDEO_ANALYSIS_SEMANTIC_AGGREGATION_MODEL=gpt-oss-120b
VIDEO_ANALYSIS_SEMANTIC_LLM_BASE_URL=http://deploy-llm-api-1:8100
ENVFILE

# Rebuild and restart
docker compose build video-analysis-api
docker compose up -d video-analysis-api

# Wait for startup
sleep 10

# Test with same video
curl -X POST http://localhost:54600/analyze/video/semantic \
  -F "file=@test_video_60s.mp4" \
  -F "chunk_duration_s=10.0" \
  -F "frames_per_chunk=24" \
  -o test_qwen3vl.json

# Review results
jq '{
  model: "Qwen3-VL-30B",
  num_chunks: .num_chunks,
  chunk_descriptions: [.chunk_results[].description],
  final_summary: .final_summary,
  total_latency_s: .total_latency_s
}' test_qwen3vl.json

Step 3: Compare Results

# Side-by-side comparison
echo "=== BusterX 7B ==="
jq -r '.chunk_results[0].description' baseline_busterx.json
echo ""
echo "=== Qwen3-VL 30B ==="
jq -r '.chunk_results[0].description' test_qwen3vl.json
echo ""

echo "=== Final Summaries ==="
echo "BusterX: $(jq -r '.final_summary' baseline_busterx.json)"
echo ""
echo "Qwen3-VL: $(jq -r '.final_summary' test_qwen3vl.json)"

Evaluation Criteria

Quality Metrics

  1. Description Detail

    • More specific object identification
    • Better action recognition
    • More context understanding
  2. Narrative Coherence

    • Logical flow between segments
    • Temporal consistency
    • Better story understanding
  3. Accuracy

    • Correct identification of people/objects
    • Accurate scene descriptions
    • Proper action sequences

Performance Metrics

Metric BusterX 7B Qwen3-VL 30B Difference
Latency per chunk ~12s ~??s ??
Total latency (60s video) ~77s ~??s ??
Token usage per chunk ~5200 ~????? ??

Expected Outcomes

If Qwen3-VL is Better:

Action: Update default configuration to use Qwen3-VL for semantic analysis

# Keep two separate configs:
# 1. Deepfake endpoint → BusterX (specialized)
# 2. Semantic endpoint → Qwen3-VL (better understanding)

Implementation:

  • Add VIDEO_ANALYSIS_DEEPFAKE_MODEL and VIDEO_ANALYSIS_SEMANTIC_MODEL settings
  • Configure different models per endpoint

If BusterX is Sufficient:

Action: Keep current configuration, document findings

Rationale:

  • Latency advantage (30B model is slower)
  • VRAM savings
  • BusterX might be sufficient for semantic tasks

Test Videos

Suggested test scenarios:

  1. Meeting/Conference (60s)

    • Multiple people
    • Complex interactions
    • Scene changes
  2. Action Sequence (60s)

    • Fast movements
    • Object manipulation
    • Environmental changes
  3. Indoor/Outdoor Transition (60s)

    • Lighting changes
    • Multiple scenes
    • Context shifts

Rollback Procedure

If issues occur:

cd /home/vasi/ml-projects/modules/video-analysis/deploy

# Restore original config
cp .env.backup .env

# Rebuild
docker compose build video-analysis-api
docker compose up -d video-analysis-api

Status

  • Baseline test completed (BusterX)
  • Test with Qwen3-VL completed
  • Results compared
  • Decision made
  • Configuration updated (if needed)
  • Documentation updated

Notes

Add observations here during testing:

Date: ___________
Tester: _________

Observations:




Recommendation: