289 lines
7.3 KiB
Markdown
289 lines
7.3 KiB
Markdown
# Audio Transcription
|
|
|
|
Speech-to-text transcription service using faster-whisper for optimized inference.
|
|
|
|
## Features
|
|
|
|
- **High-Performance Transcription**: Uses faster-whisper (4x faster than original Whisper)
|
|
- **Multiple Model Sizes**: Support for tiny, base, small, medium, large-v2, large-v3, and large-v3-turbo
|
|
- **GPU Acceleration**: Optimized CUDA inference with int8 quantization
|
|
- **OpenAI-Compatible API**: Drop-in replacement for OpenAI's transcription endpoint
|
|
- **Multiple Output Formats**: JSON, text, and verbose JSON with segments
|
|
- **Language Detection**: Automatic language detection for 99+ languages
|
|
- **VAD Filtering**: Voice Activity Detection for improved accuracy
|
|
|
|
## Prerequisites
|
|
|
|
**Required (all modules):**
|
|
- All global prerequisites (see main [README.md](../../README.md))
|
|
- NVIDIA GPU with CUDA support (recommended for production)
|
|
- ~6GB VRAM for large-v3-turbo with int8 quantization
|
|
|
|
**Optional:**
|
|
- CPU-only mode available (slower but no GPU required)
|
|
|
|
> **Note:** The audio module runs on GPU 0 (shared with Qwen3.5-35B-A3B).
|
|
|
|
## Installation
|
|
|
|
```bash
|
|
cd modules/audio
|
|
|
|
# Install dependencies
|
|
uv sync
|
|
|
|
# Install with dev dependencies
|
|
uv sync --extra dev
|
|
```
|
|
|
|
## Quick Start
|
|
|
|
### As an API Server
|
|
|
|
```bash
|
|
cd deploy/
|
|
|
|
# Create deployment env file
|
|
cp ../.env.example .env
|
|
# Edit deploy/.env with your configuration
|
|
|
|
# Start the server
|
|
./deploy.sh --profile api --detach
|
|
|
|
# Or with nginx reverse proxy
|
|
./deploy.sh --profile api-nginx --detach
|
|
```
|
|
|
|
### API Endpoints
|
|
|
|
| Endpoint | Method | Description |
|
|
|----------|--------|-------------|
|
|
| `/health` | GET | Health check |
|
|
| `/v1/models` | GET | List available models |
|
|
| `/v1/audio/transcriptions` | POST | Transcribe audio (OpenAI-compatible) |
|
|
|
|
### Example API Request
|
|
|
|
```bash
|
|
# Health check
|
|
curl http://localhost:8200/health
|
|
|
|
# Transcribe audio file
|
|
AUDIO="/path/to/audio.mp3"
|
|
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
|
-F "file=@${AUDIO}" \
|
|
-F "model=large-v3-turbo" \
|
|
-F "response_format=json"
|
|
|
|
# With language specification and verbose output
|
|
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
|
-F "file=@${AUDIO}" \
|
|
-F "language=en" \
|
|
-F "response_format=verbose_json"
|
|
```
|
|
|
|
### Example Response
|
|
|
|
**JSON format:**
|
|
```json
|
|
{
|
|
"text": "This is the full transcription of your audio file.",
|
|
"language": "en",
|
|
"duration": 45.5
|
|
}
|
|
```
|
|
|
|
**Verbose JSON format:**
|
|
```json
|
|
{
|
|
"text": "This is the full transcription.",
|
|
"language": "en",
|
|
"duration": 45.5,
|
|
"segments": [
|
|
{
|
|
"id": 0,
|
|
"start": 0.0,
|
|
"end": 3.5,
|
|
"text": "This is the full transcription.",
|
|
"tokens": [123, 456, 789],
|
|
"temperature": 0.0,
|
|
"avg_logprob": -0.25,
|
|
"compression_ratio": 1.5,
|
|
"no_speech_prob": 0.01
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
## Configuration
|
|
|
|
### Required Environment Variables
|
|
|
|
Configured via environment variables (prefix: `AUDIO_`). These are typically set in `deploy/.env`:
|
|
|
|
| Variable | Description |
|
|
|----------|-------------|
|
|
| `AUDIO_MODEL` | Whisper model name (e.g., `large-v3-turbo`) |
|
|
| `AUDIO_DEVICE` | Device: `cuda` or `cpu` |
|
|
| `AUDIO_CACHE_DIR` | Model cache directory |
|
|
|
|
### Optional Environment Variables
|
|
|
|
| Variable | Default | Description |
|
|
|----------|---------|-------------|
|
|
| `AUDIO_COMPUTE_TYPE` | `int8` | Compute type: `float16`, `int8`, `int8_float16` |
|
|
| `AUDIO_BEAM_SIZE` | `5` | Beam size for decoding (1-10) |
|
|
| `AUDIO_BEST_OF` | `5` | Number of candidates when sampling |
|
|
| `AUDIO_TEMPERATURE` | `0.0` | Sampling temperature (0.0-1.0) |
|
|
| `AUDIO_LOG_LEVEL` | `INFO` | Log level |
|
|
| `AUDIO_MAX_FILE_SIZE_MB` | `500` | Max upload file size in MB |
|
|
|
|
### Available Models
|
|
|
|
| Model | Parameters | VRAM (int8) | Speed | Quality |
|
|
|-------|-----------|-------------|-------|---------|
|
|
| `tiny` | 39M | ~1GB | 10x | Basic |
|
|
| `base` | 74M | ~1GB | 7x | Good |
|
|
| `small` | 244M | ~2GB | 4x | Better |
|
|
| `medium` | 769M | ~3-4GB | 2x | Very Good |
|
|
| `large-v3` | 1550M | ~6-8GB | 1x | Excellent |
|
|
| `large-v3-turbo` | 809M | ~6GB | 8x | Excellent |
|
|
|
|
**Recommendation:** Use `large-v3-turbo` for best balance of speed and accuracy.
|
|
|
|
## Deployment
|
|
|
|
```bash
|
|
cd deploy/
|
|
|
|
# Create deployment env file (REQUIRED)
|
|
cp ../.env.example .env
|
|
# Edit deploy/.env with your settings
|
|
|
|
# API only
|
|
./deploy.sh --profile api --detach
|
|
|
|
# API with nginx reverse proxy
|
|
./deploy.sh --profile api-nginx --detach
|
|
|
|
# View logs
|
|
./deploy.sh --profile api --logs
|
|
|
|
# Stop services
|
|
./deploy.sh --profile api --down
|
|
```
|
|
|
|
### Common Docker Commands
|
|
|
|
```bash
|
|
cd deploy/
|
|
|
|
# Restart the API service
|
|
docker compose restart audio-api
|
|
|
|
# Rebuild after code changes
|
|
docker compose build audio-api
|
|
docker compose up -d audio-api
|
|
|
|
# Check logs
|
|
docker compose logs -f audio-api
|
|
```
|
|
|
|
### Port Allocation
|
|
|
|
| Port | Service |
|
|
|------|---------|
|
|
| `8200` | Audio API |
|
|
| `54300` | Audio API (Dev + AI + Audio) |
|
|
|
|
## Development
|
|
|
|
```bash
|
|
# Install dev dependencies
|
|
uv sync --extra dev
|
|
|
|
# Run tests
|
|
uv run pytest
|
|
|
|
# Run tests with coverage
|
|
uv run pytest --cov=src/audio --cov-report=term-missing
|
|
|
|
# Lint and format
|
|
uv run ruff check .
|
|
uv run ruff format .
|
|
```
|
|
|
|
## Architecture
|
|
|
|
```
|
|
modules/audio/
|
|
├── deploy/
|
|
│ ├── deploy.sh # Deployment script
|
|
│ ├── docker-compose.yml # Docker services
|
|
│ ├── Dockerfile # Container image
|
|
│ └── nginx.conf # Nginx reverse proxy config (optional)
|
|
├── src/audio/
|
|
│ ├── __init__.py
|
|
│ ├── app.py # FastAPI application
|
|
│ ├── transcriber.py # faster-whisper wrapper
|
|
│ ├── schemas.py # Response schemas
|
|
│ └── settings.py # Configuration
|
|
├── tests/
|
|
├── .env.example # Environment template
|
|
├── API.md # API documentation
|
|
├── pyproject.toml # Dependencies
|
|
└── README.md # This file
|
|
```
|
|
|
|
## Supported Audio Formats
|
|
|
|
faster-whisper (via FFmpeg) supports:
|
|
- MP3, WAV, FLAC, OGG, M4A, AAC, WMA
|
|
- MP4, AVI, MKV (audio track extraction)
|
|
- And many more formats supported by FFmpeg
|
|
|
|
## Performance Tips
|
|
|
|
1. **Use int8 quantization** for GPU inference (40% memory savings, minimal accuracy loss)
|
|
2. **Use large-v3-turbo** for best speed/accuracy tradeoff (8x faster than large-v3)
|
|
3. **Enable VAD filtering** (enabled by default) to skip silence
|
|
4. **Batch processing**: Process multiple files sequentially for better GPU utilization
|
|
5. **Language specification**: Specify language code when known for faster processing
|
|
|
|
## Troubleshooting
|
|
|
|
### Issue: Out of Memory
|
|
|
|
**Solution:** Use a smaller model or reduce compute type:
|
|
```bash
|
|
# Use medium model instead
|
|
AUDIO_MODEL=medium
|
|
|
|
# Or use float16 instead of int8 (uses more VRAM but may work better)
|
|
AUDIO_COMPUTE_TYPE=float16
|
|
```
|
|
|
|
### Issue: Slow Transcription
|
|
|
|
**Solution:**
|
|
- Ensure GPU is being used (`AUDIO_DEVICE=cuda`)
|
|
- Use int8 compute type for faster inference
|
|
- Use large-v3-turbo instead of large-v3
|
|
- Check GPU utilization with `nvidia-smi`
|
|
|
|
### Issue: Model Download Fails
|
|
|
|
**Solution:** Check network connectivity and cache directory permissions:
|
|
```bash
|
|
# Verify cache directory exists and is writable
|
|
ls -la /cai2_ds_storage/hf_cache
|
|
|
|
# Or change to local directory
|
|
AUDIO_CACHE_DIR=/home/user/.cache/huggingface
|
|
```
|
|
|
|
## References
|
|
|
|
- [faster-whisper GitHub](https://github.com/SYSTRAN/faster-whisper)
|
|
- [OpenAI Whisper](https://github.com/openai/whisper)
|
|
- [Whisper Model Card](https://github.com/openai/whisper/blob/main/model-card.md)
|