# Audio Transcription Service (M17-Whisper) Audio transcription service for DIDI media analysis. Whisper-based (M17-Whisper, faster-whisper backend), used by agent-v3 media-preprocess worker for video/audio session pipelines (techniques, ai-tampered, claims). OpenAI-compatible API — drop-in for `client.audio.transcriptions.create()`. ## Stack - Python 3.10+ / FastAPI / Uvicorn - Backend: `faster-whisper` >= 1.0.0 (CTranslate2 optimized inference) - Default model: `large-v3-turbo` (809M params, ~6GB VRAM int8) - GPU: CUDA (shared GPU 0 with Qwen3.5-35B-A3B) - URL (Dev): `http://10.11.10.17:54300/v1/audio/transcriptions` - Container: `didiAI-audio` (GPU host) - Auth: none on the service itself; agent-v3 uses bearer token via `M17_WHISPER_TOKEN` (enforced by gateway/nginx if configured) ## Ce face Speech-to-text transcription pe fişiere audio sau URL-uri. agent-v3 trimite fie audio buffer (multipart upload), fie URL public (din MinIO), primeşte text + metadata (lang detect, duration, optional segments). Folosit în: - **Video analysis pipeline** — agent-v3 extrage track-ul audio cu ffmpeg, trimite la M17-Whisper, foloseşte transcript pentru `claims` + `techniques` + `ai-tampered` - **Audio-only sessions** — direct upload, transcripted, apoi pipeline normal de analiză text - **Cascade fallback** — agent-v3 `transcription.ts` are cascadă: M17-Whisper (local GPU) → Groq Whisper → OpenAI Whisper. Dacă local pică sau întoarce string gol, trece la următorul provider. ## API endpoints | Method | Path | Purpose | |--------|------|---------| | `GET` | `/health` | Health probe (returns `{"status":"ok"}`) | | `GET` | `/v1/models` | List models, OpenAI-compatible | | `POST` | `/v1/audio/transcriptions` | Transcribe audio (OpenAI-compatible) | | `GET` | `/v1/info` | Service catalog metadata (used by didi catalog-api) | ### `POST /v1/audio/transcriptions` Content-Type: `multipart/form-data` | Param | Type | Default | Notes | |-------|------|---------|-------| | `file` | file | — | Audio file (MP3/WAV/M4A/FLAC/OGG/MP4 audio track…). XOR with `url`. | | `url` | string | — | URL to download audio from (DIDI extension over OpenAI). XOR with `file`. | | `model` | string | configured | Ignored — server uses `AUDIO_MODEL` env. | | `language` | string | auto | ISO-639-1 code (`en`, `ro`, `es`, …). Auto-detect if omitted. | | `prompt` | string | none | Optional initial prompt to bias style/vocab. | | `response_format` | string | `json` | `json` \| `text` \| `verbose_json` (segments). | | `temperature` | float | `0.0` | Sampling temperature 0.0–1.0. | Response (json): `{ text, language, duration }` — verbose_json adds `segments[]` with `start/end/text/tokens/avg_logprob/no_speech_prob`. Errors: `400` invalid params / download failed, `413` payload too large (default 500MB), `422` corrupted audio, `500` transcription failed, `503` model not loaded. ## Backends / Models - Production: `large-v3-turbo` (8x faster than `large-v3`, similar quality, ~6GB VRAM int8) - Available: `tiny`, `base`, `small`, `medium`, `large-v3`, `large-v3-turbo` - VAD (Voice Activity Detection) enabled by default — skip silence - Quantization: `int8` (recommended), `float16`, `int8_float16` ## How didi-backend uses it - agent-v3 cascade: `backend/services/orchestration-layer/agent-v3/src/shared/media/transcription.ts` - Order: **M17-Whisper (local GPU)** → Groq Whisper → OpenAI Whisper - Empty-string from local triggers retry on next provider (Faza 1 din SAFETY_NETS_PLAN) - Configured în agent-v3 docker-compose env: - `M17_WHISPER_URL=http://10.11.10.17:54300` - `M17_WHISPER_TOKEN=` (optional, dacă există proxy auth) - Apelat din **media-preprocess worker** la sesiuni cu `media_type` în `{audio, video}` — transcript devine input pentru `claims-routes.ts`, `routes.ts` (techniques), `ai-tampered-routes.ts`. - Video flow: ffmpeg extract → upload temp / pass buffer → POST `/v1/audio/transcriptions` → text → pipeline analysis. ## Configuration Env vars (prefix `AUDIO_`, set in `deploy/.env`): **Required** - `AUDIO_MODEL` — Whisper model (default `large-v3-turbo`) - `AUDIO_DEVICE` — `cuda` | `cpu` - `AUDIO_CACHE_DIR` — HuggingFace model cache path (`/cai2_ds_storage/hf_cache` typical) - `AUDIO_EXTERNAL_URL` — used in OpenAPI servers spec **Optional** - `AUDIO_COMPUTE_TYPE` — `int8` (default), `float16`, `int8_float16` - `AUDIO_BEAM_SIZE` — beam search width (default 5) - `AUDIO_BEST_OF` — sampling candidates (default 5) - `AUDIO_TEMPERATURE` — default 0.0 - `AUDIO_HOST` / `AUDIO_PORT` — default `0.0.0.0:54300` - `AUDIO_LOG_LEVEL` — default `INFO` - `AUDIO_MAX_FILE_SIZE_MB` — upload cap (default 500) - `CUDA_VISIBLE_DEVICES=0` — pinned to GPU 0 ## Deployment Compose dir: `/home/admin365/didi_mono/ai_platform/modules/audio/deploy/` ```bash cd modules/audio/deploy cp ../.env.example .env # edit values ./deploy.sh --profile api --detach # API only ./deploy.sh --profile api-nginx --detach # with nginx reverse proxy ./deploy.sh --profile api --logs # tail logs docker compose restart audio-api # quick restart ``` - Container name: `didiAI-audio` - Image: `didiai-audio:audit` - Network: `didi-network` (external, shared with other AI modules) - GPU reservation: NVIDIA driver, device `0` - Healthcheck: HTTP `GET /health` every 30s, 60s start period (model load) - Model cache mount: `${AUDIO_CACHE_DIR}:/root/.cache/huggingface` ## Performance - Real-Time Factor on H200 GPU + int8 + `large-v3-turbo`: **~0.05–0.08x** (10s audio = 0.5–1s transcription, 1h audio = 3–5min) - VRAM: ~6GB for large-v3-turbo int8, ~8GB float16 - Throughput: sequential — process files one at a time per GPU - File size cap: 500MB default (configurable) - Recommended: pre-split audio > 1h cu ffmpeg segments ## Architecture (source layout) ``` modules/audio/ ├── deploy/ │ ├── deploy.sh # CLI wrapper │ ├── docker-compose.yml # didiAI-audio service │ ├── Dockerfile # CUDA + faster-whisper image │ └── .env # runtime config ├── src/audio/ │ ├── app.py # FastAPI routes + /v1/info catalog │ ├── transcriber.py # faster-whisper wrapper (singleton) │ ├── schemas.py # TranscriptionResponse / Segment │ └── settings.py # pydantic-settings (AUDIO_* env) ├── tests/ ├── pyproject.toml # fastapi, faster-whisper, httpx ├── README.md # full operator guide ├── API.md # full HTTP API reference └── INDEX.md # this file ``` ## Related - agent-v3 transcription cascade — `backend/services/orchestration-layer/agent-v3/src/shared/media/transcription.ts` - agent-v3 media routes — `routes.ts` / `ai-tampered-routes.ts` / `claims-routes.ts` - media-preprocess worker (in agent-v3) — invokes this service - Test files in MinIO (used during refactor verification): - Audio: `https://didi365.eu/api/v3/media/file/uploads/test-user/1771883851173-audio_with_voice.mp3` - Video: `https://didi365.eu/api/v3/media/file/uploads/test-user/1771883852186-voice_video.mp4` - Faza 1 safety net (transcription retry on empty) — `agent-v3/SAFETY_NETS_PLAN.md` - Upstream refs: [faster-whisper](https://github.com/SYSTRAN/faster-whisper), [OpenAI Whisper](https://github.com/openai/whisper)