didi-lot1-ai/ai_platform/modules/audio/INDEX.md

137 lines
7.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Audio Transcription Service (M17-Whisper)
Audio transcription service for DIDI media analysis. Whisper-based (M17-Whisper, faster-whisper backend), used by agent-v3 media-preprocess worker for video/audio session pipelines (techniques, ai-tampered, claims). OpenAI-compatible API — drop-in for `client.audio.transcriptions.create()`.
## Stack
- Python 3.10+ / FastAPI / Uvicorn
- Backend: `faster-whisper` >= 1.0.0 (CTranslate2 optimized inference)
- Default model: `large-v3-turbo` (809M params, ~6GB VRAM int8)
- GPU: CUDA (shared GPU 0 with Qwen3.5-35B-A3B)
- URL (Dev): `http://10.11.10.17:54300/v1/audio/transcriptions`
- Container: `didiAI-audio` (GPU host)
- Auth: none on the service itself; agent-v3 uses bearer token via `M17_WHISPER_TOKEN` (enforced by gateway/nginx if configured)
## Ce face
Speech-to-text transcription pe fişiere audio sau URL-uri. agent-v3 trimite fie audio buffer (multipart upload), fie URL public (din MinIO), primeşte text + metadata (lang detect, duration, optional segments). Folosit în:
- **Video analysis pipeline** — agent-v3 extrage track-ul audio cu ffmpeg, trimite la M17-Whisper, foloseşte transcript pentru `claims` + `techniques` + `ai-tampered`
- **Audio-only sessions** — direct upload, transcripted, apoi pipeline normal de analiză text
- **Cascade fallback** — agent-v3 `transcription.ts` are cascadă: M17-Whisper (local GPU) → Groq Whisper → OpenAI Whisper. Dacă local pică sau întoarce string gol, trece la următorul provider.
## API endpoints
| Method | Path | Purpose |
|--------|------|---------|
| `GET` | `/health` | Health probe (returns `{"status":"ok"}`) |
| `GET` | `/v1/models` | List models, OpenAI-compatible |
| `POST` | `/v1/audio/transcriptions` | Transcribe audio (OpenAI-compatible) |
| `GET` | `/v1/info` | Service catalog metadata (used by didi catalog-api) |
### `POST /v1/audio/transcriptions`
Content-Type: `multipart/form-data`
| Param | Type | Default | Notes |
|-------|------|---------|-------|
| `file` | file | — | Audio file (MP3/WAV/M4A/FLAC/OGG/MP4 audio track…). XOR with `url`. |
| `url` | string | — | URL to download audio from (DIDI extension over OpenAI). XOR with `file`. |
| `model` | string | configured | Ignored — server uses `AUDIO_MODEL` env. |
| `language` | string | auto | ISO-639-1 code (`en`, `ro`, `es`, …). Auto-detect if omitted. |
| `prompt` | string | none | Optional initial prompt to bias style/vocab. |
| `response_format` | string | `json` | `json` \| `text` \| `verbose_json` (segments). |
| `temperature` | float | `0.0` | Sampling temperature 0.01.0. |
Response (json): `{ text, language, duration }` — verbose_json adds `segments[]` with `start/end/text/tokens/avg_logprob/no_speech_prob`.
Errors: `400` invalid params / download failed, `413` payload too large (default 500MB), `422` corrupted audio, `500` transcription failed, `503` model not loaded.
## Backends / Models
- Production: `large-v3-turbo` (8x faster than `large-v3`, similar quality, ~6GB VRAM int8)
- Available: `tiny`, `base`, `small`, `medium`, `large-v3`, `large-v3-turbo`
- VAD (Voice Activity Detection) enabled by default — skip silence
- Quantization: `int8` (recommended), `float16`, `int8_float16`
## How didi-backend uses it
- agent-v3 cascade: `backend/services/orchestration-layer/agent-v3/src/shared/media/transcription.ts`
- Order: **M17-Whisper (local GPU)** → Groq Whisper → OpenAI Whisper
- Empty-string from local triggers retry on next provider (Faza 1 din SAFETY_NETS_PLAN)
- Configured în agent-v3 docker-compose env:
- `M17_WHISPER_URL=http://10.11.10.17:54300`
- `M17_WHISPER_TOKEN=<gateway-token>` (optional, dacă există proxy auth)
- Apelat din **media-preprocess worker** la sesiuni cu `media_type` în `{audio, video}` — transcript devine input pentru `claims-routes.ts`, `routes.ts` (techniques), `ai-tampered-routes.ts`.
- Video flow: ffmpeg extract → upload temp / pass buffer → POST `/v1/audio/transcriptions` → text → pipeline analysis.
## Configuration
Env vars (prefix `AUDIO_`, set in `deploy/.env`):
**Required**
- `AUDIO_MODEL` — Whisper model (default `large-v3-turbo`)
- `AUDIO_DEVICE``cuda` | `cpu`
- `AUDIO_CACHE_DIR` — HuggingFace model cache path (`/cai2_ds_storage/hf_cache` typical)
- `AUDIO_EXTERNAL_URL` — used in OpenAPI servers spec
**Optional**
- `AUDIO_COMPUTE_TYPE``int8` (default), `float16`, `int8_float16`
- `AUDIO_BEAM_SIZE` — beam search width (default 5)
- `AUDIO_BEST_OF` — sampling candidates (default 5)
- `AUDIO_TEMPERATURE` — default 0.0
- `AUDIO_HOST` / `AUDIO_PORT` — default `0.0.0.0:54300`
- `AUDIO_LOG_LEVEL` — default `INFO`
- `AUDIO_MAX_FILE_SIZE_MB` — upload cap (default 500)
- `CUDA_VISIBLE_DEVICES=0` — pinned to GPU 0
## Deployment
Compose dir: `/home/admin365/didi_mono/ai_platform/modules/audio/deploy/`
```bash
cd modules/audio/deploy
cp ../.env.example .env # edit values
./deploy.sh --profile api --detach # API only
./deploy.sh --profile api-nginx --detach # with nginx reverse proxy
./deploy.sh --profile api --logs # tail logs
docker compose restart audio-api # quick restart
```
- Container name: `didiAI-audio`
- Image: `didiai-audio:audit`
- Network: `didi-network` (external, shared with other AI modules)
- GPU reservation: NVIDIA driver, device `0`
- Healthcheck: HTTP `GET /health` every 30s, 60s start period (model load)
- Model cache mount: `${AUDIO_CACHE_DIR}:/root/.cache/huggingface`
## Performance
- Real-Time Factor on H200 GPU + int8 + `large-v3-turbo`: **~0.050.08x** (10s audio = 0.51s transcription, 1h audio = 35min)
- VRAM: ~6GB for large-v3-turbo int8, ~8GB float16
- Throughput: sequential — process files one at a time per GPU
- File size cap: 500MB default (configurable)
- Recommended: pre-split audio > 1h cu ffmpeg segments
## Architecture (source layout)
```
modules/audio/
├── deploy/
│ ├── deploy.sh # CLI wrapper
│ ├── docker-compose.yml # didiAI-audio service
│ ├── Dockerfile # CUDA + faster-whisper image
│ └── .env # runtime config
├── src/audio/
│ ├── app.py # FastAPI routes + /v1/info catalog
│ ├── transcriber.py # faster-whisper wrapper (singleton)
│ ├── schemas.py # TranscriptionResponse / Segment
│ └── settings.py # pydantic-settings (AUDIO_* env)
├── tests/
├── pyproject.toml # fastapi, faster-whisper, httpx
├── README.md # full operator guide
├── API.md # full HTTP API reference
└── INDEX.md # this file
```
## Related
- agent-v3 transcription cascade — `backend/services/orchestration-layer/agent-v3/src/shared/media/transcription.ts`
- agent-v3 media routes — `routes.ts` / `ai-tampered-routes.ts` / `claims-routes.ts`
- media-preprocess worker (in agent-v3) — invokes this service
- Test files in MinIO (used during refactor verification):
- Audio: `https://didi365.eu/api/v3/media/file/uploads/test-user/1771883851173-audio_with_voice.mp3`
- Video: `https://didi365.eu/api/v3/media/file/uploads/test-user/1771883852186-voice_video.mp4`
- Faza 1 safety net (transcription retry on empty) — `agent-v3/SAFETY_NETS_PLAN.md`
- Upstream refs: [faster-whisper](https://github.com/SYSTRAN/faster-whisper), [OpenAI Whisper](https://github.com/openai/whisper)