Livrare LOT 1 - Didi

This commit is contained in:
Dezvoltari Evotech 2026-06-25 14:13:25 -07:00
commit 5380c3fc63
990 changed files with 133308 additions and 0 deletions

View file

@ -0,0 +1,252 @@
# Video Analysis - Testing Checklist
## ⚠️ HIGH PRIORITY: Model Comparison for Semantic Analysis
### Background
Currently both deepfake detection and semantic analysis use **BusterX** (Qwen2.5-VL-7B, 7B parameters). However, we have access to a much larger model **Qwen3-VL-30B** (30B parameters) that could provide significantly better semantic understanding.
### Hypothesis
Semantic analysis (content understanding, scene description, narrative) would benefit from the larger Qwen3-VL-30B model, while deepfake detection should continue using the specialized BusterX model.
---
## Test Plan: Qwen3-VL-30B for Semantic Analysis
### Current Configuration
```bash
# modules/video-analysis/deploy/.env
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-video-vllm-buster:54500 # Port 54500
VIDEO_ANALYSIS_VLLM_MODEL=busterx # 7B parameters
```
### Test Configuration
```bash
# modules/video-analysis/deploy/.env
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-llm-vllm-vision:14002 # Port 14002
VIDEO_ANALYSIS_VLLM_MODEL=qwen3-vl # 30B parameters
```
---
## Testing Procedure
### Step 1: Baseline Test (BusterX 7B)
```bash
# Current configuration - no changes needed
cd /home/vasi/ml-projects/modules/video-analysis/deploy
# Test semantic analysis with BusterX
curl -X POST http://localhost:8007/analyze/video/semantic \
-F "file=@test_video_60s.mp4" \
-F "chunk_duration_s=10.0" \
-F "frames_per_chunk=24" \
-o baseline_busterx.json
# Review results
jq '{
model: "BusterX-7B",
num_chunks: .num_chunks,
chunk_descriptions: [.chunk_results[].description],
final_summary: .final_summary,
total_latency_s: .total_latency_s
}' baseline_busterx.json
```
### Step 2: Test with Qwen3-VL-30B
```bash
cd /home/vasi/ml-projects/modules/video-analysis/deploy
# Backup current config
cp .env .env.backup
# Update to use Qwen3-VL
cat > .env << 'ENVFILE'
HF_TOKEN=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
HF_CACHE_DIR=/cai2_ds_storage/hf_cache
VIDEO_ANALYSIS_RUNS_DIR=/app/runs
# Switch to Qwen3-VL for testing
VIDEO_ANALYSIS_VLLM_BASE_URL=http://didiAI-llm-vllm-vision:14002
VIDEO_ANALYSIS_VLLM_MODEL=qwen3-vl
# Semantic analysis settings
VIDEO_ANALYSIS_SEMANTIC_CHUNK_DURATION_S=10.0
VIDEO_ANALYSIS_SEMANTIC_FRAMES_PER_CHUNK=24
VIDEO_ANALYSIS_SEMANTIC_ENABLE_AGGREGATION=true
VIDEO_ANALYSIS_SEMANTIC_AGGREGATION_MODEL=gpt-oss-120b
VIDEO_ANALYSIS_SEMANTIC_LLM_BASE_URL=http://deploy-llm-api-1:8100
ENVFILE
# Rebuild and restart
docker compose build video-analysis-api
docker compose up -d video-analysis-api
# Wait for startup
sleep 10
# Test with same video
curl -X POST http://localhost:8007/analyze/video/semantic \
-F "file=@test_video_60s.mp4" \
-F "chunk_duration_s=10.0" \
-F "frames_per_chunk=24" \
-o test_qwen3vl.json
# Review results
jq '{
model: "Qwen3-VL-30B",
num_chunks: .num_chunks,
chunk_descriptions: [.chunk_results[].description],
final_summary: .final_summary,
total_latency_s: .total_latency_s
}' test_qwen3vl.json
```
### Step 3: Compare Results
```bash
# Side-by-side comparison
echo "=== BusterX 7B ==="
jq -r '.chunk_results[0].description' baseline_busterx.json
echo ""
echo "=== Qwen3-VL 30B ==="
jq -r '.chunk_results[0].description' test_qwen3vl.json
echo ""
echo "=== Final Summaries ==="
echo "BusterX: $(jq -r '.final_summary' baseline_busterx.json)"
echo ""
echo "Qwen3-VL: $(jq -r '.final_summary' test_qwen3vl.json)"
```
---
## Evaluation Criteria
### Quality Metrics
1. **Description Detail**
- [ ] More specific object identification
- [ ] Better action recognition
- [ ] More context understanding
2. **Narrative Coherence**
- [ ] Logical flow between segments
- [ ] Temporal consistency
- [ ] Better story understanding
3. **Accuracy**
- [ ] Correct identification of people/objects
- [ ] Accurate scene descriptions
- [ ] Proper action sequences
### Performance Metrics
| Metric | BusterX 7B | Qwen3-VL 30B | Difference |
|--------|-----------|--------------|------------|
| Latency per chunk | ~12s | ~??s | ?? |
| Total latency (60s video) | ~77s | ~??s | ?? |
| Token usage per chunk | ~5200 | ~????? | ?? |
---
## Expected Outcomes
### If Qwen3-VL is Better:
**Action:** Update default configuration to use Qwen3-VL for semantic analysis
```bash
# Keep two separate configs:
# 1. Deepfake endpoint → BusterX (specialized)
# 2. Semantic endpoint → Qwen3-VL (better understanding)
```
**Implementation:**
- Add `VIDEO_ANALYSIS_DEEPFAKE_MODEL` and `VIDEO_ANALYSIS_SEMANTIC_MODEL` settings
- Configure different models per endpoint
### If BusterX is Sufficient:
**Action:** Keep current configuration, document findings
**Rationale:**
- Latency advantage (30B model is slower)
- VRAM savings
- BusterX might be sufficient for semantic tasks
---
## Test Videos
Suggested test scenarios:
1. **Meeting/Conference** (60s)
- Multiple people
- Complex interactions
- Scene changes
2. **Action Sequence** (60s)
- Fast movements
- Object manipulation
- Environmental changes
3. **Indoor/Outdoor Transition** (60s)
- Lighting changes
- Multiple scenes
- Context shifts
---
## Rollback Procedure
If issues occur:
```bash
cd /home/vasi/ml-projects/modules/video-analysis/deploy
# Restore original config
cp .env.backup .env
# Rebuild
docker compose build video-analysis-api
docker compose up -d video-analysis-api
```
---
## Status
- [ ] Baseline test completed (BusterX)
- [ ] Test with Qwen3-VL completed
- [ ] Results compared
- [ ] Decision made
- [ ] Configuration updated (if needed)
- [ ] Documentation updated
---
## Notes
Add observations here during testing:
```
Date: ___________
Tester: _________
Observations:
Recommendation:
```

View file

@ -0,0 +1,3 @@
....... [100%]
------------------ generated xml file: /app/test-results.xml -------------------
7 passed in 0.23s

View file

@ -0,0 +1 @@
<?xml version="1.0" encoding="utf-8"?><testsuites name="pytest tests"><testsuite name="pytest" errors="0" failures="0" skipped="0" tests="7" time="0.235" timestamp="2026-06-25T15:16:23.358556+00:00" hostname="2b93193c8419"><testcase classname="tests.test_verdict" name="test_parse_real" time="0.000" /><testcase classname="tests.test_verdict" name="test_parse_fake" time="0.000" /><testcase classname="tests.test_verdict" name="test_parse_unclear_is_uncertain" time="0.000" /><testcase classname="tests.test_verdict" name="test_parse_empty_is_uncertain" time="0.000" /><testcase classname="tests.test_verdict" name="test_build_frame_evidence_aligns_indices_and_timestamps" time="0.000" /><testcase classname="tests.test_verdict" name="test_build_frame_evidence_handles_missing_timestamps" time="0.000" /><testcase classname="tests.test_verdict" name="test_build_frame_evidence_empty" time="0.000" /></testsuite></testsuites>