Livrare LOT 1 - Didi
This commit is contained in:
commit
5380c3fc63
990 changed files with 133308 additions and 0 deletions
453
ai_platform/modules/audio/API.md
Normal file
453
ai_platform/modules/audio/API.md
Normal file
|
|
@ -0,0 +1,453 @@
|
|||
# Audio Transcription API Documentation
|
||||
|
||||
OpenAI-compatible speech-to-text transcription API using faster-whisper.
|
||||
|
||||
## Base URL
|
||||
|
||||
```
|
||||
{BASE_URL}
|
||||
```
|
||||
|
||||
- **Local development:** `http://localhost:8200`
|
||||
- **Docker (internal):** `http://audio-api:8200`
|
||||
- **Direct:** `http://localhost:54300`
|
||||
- **Production:** Use your configured hostname
|
||||
|
||||
## Authentication
|
||||
|
||||
No authentication required by default. Can be added via nginx or API gateway if needed.
|
||||
|
||||
## API Endpoints
|
||||
|
||||
### Health Check
|
||||
|
||||
Check API health status.
|
||||
|
||||
**Endpoint:** `GET /health`
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"status": "ok"
|
||||
}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
curl http://localhost:8200/health
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### List Models
|
||||
|
||||
List available Whisper models (OpenAI-compatible).
|
||||
|
||||
**Endpoint:** `GET /v1/models`
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "large-v3-turbo",
|
||||
"object": "model",
|
||||
"created": 1700000000,
|
||||
"owned_by": "openai"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
curl http://localhost:8200/v1/models
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Create Transcription
|
||||
|
||||
Transcribe audio file to text (OpenAI-compatible endpoint).
|
||||
|
||||
**Endpoint:** `POST /v1/audio/transcriptions`
|
||||
|
||||
**Content-Type:** `multipart/form-data`
|
||||
|
||||
**Request Parameters:**
|
||||
|
||||
| Parameter | Type | Required | Default | Description |
|
||||
|-----------|------|----------|---------|-------------|
|
||||
| `file` | file | Yes | - | Audio file to transcribe (MP3, WAV, M4A, etc.) |
|
||||
| `model` | string | No | `large-v3-turbo` | Model to use (currently ignored, uses configured model) |
|
||||
| `language` | string | No | `null` | Language code (ISO-639-1). Auto-detected if not specified. |
|
||||
| `prompt` | string | No | `null` | Optional text to guide the model's style |
|
||||
| `response_format` | string | No | `json` | Format: `json`, `text`, or `verbose_json` |
|
||||
| `temperature` | float | No | `0.0` | Sampling temperature (0.0-1.0). Use 0.0 for deterministic output. |
|
||||
|
||||
**Supported Languages (ISO-639-1 codes):**
|
||||
`en`, `es`, `fr`, `de`, `it`, `pt`, `nl`, `pl`, `tr`, `ru`, `ja`, `ko`, `zh`, `ar`, `hi`, and 90+ more languages.
|
||||
|
||||
**Response Formats:**
|
||||
|
||||
#### 1. JSON (default)
|
||||
```json
|
||||
{
|
||||
"text": "Full transcription text",
|
||||
"language": "en",
|
||||
"duration": 45.5
|
||||
}
|
||||
```
|
||||
|
||||
#### 2. Text
|
||||
```
|
||||
Full transcription text
|
||||
```
|
||||
|
||||
#### 3. Verbose JSON
|
||||
```json
|
||||
{
|
||||
"text": "Full transcription text",
|
||||
"language": "en",
|
||||
"duration": 45.5,
|
||||
"segments": [
|
||||
{
|
||||
"id": 0,
|
||||
"seek": 0,
|
||||
"start": 0.0,
|
||||
"end": 3.5,
|
||||
"text": "Hello, world!",
|
||||
"tokens": [15496, 11, 1002, 0],
|
||||
"temperature": 0.0,
|
||||
"avg_logprob": -0.25,
|
||||
"compression_ratio": 1.5,
|
||||
"no_speech_prob": 0.01
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Examples:**
|
||||
|
||||
**Basic transcription (JSON):**
|
||||
```bash
|
||||
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
||||
-F "file=@audio.mp3" \
|
||||
-F "response_format=json"
|
||||
```
|
||||
|
||||
**With language specification:**
|
||||
```bash
|
||||
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
||||
-F "file=@audio.mp3" \
|
||||
-F "language=en" \
|
||||
-F "response_format=json"
|
||||
```
|
||||
|
||||
**Text format:**
|
||||
```bash
|
||||
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
||||
-F "file=@audio.mp3" \
|
||||
-F "response_format=text"
|
||||
```
|
||||
|
||||
**Verbose JSON with segments:**
|
||||
```bash
|
||||
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
||||
-F "file=@audio.mp3" \
|
||||
-F "response_format=verbose_json"
|
||||
```
|
||||
|
||||
**With initial prompt (to guide style):**
|
||||
```bash
|
||||
curl -X POST "http://localhost:8200/v1/audio/transcriptions" \
|
||||
-F "file=@audio.mp3" \
|
||||
-F "prompt=This is a technical discussion about machine learning." \
|
||||
-F "response_format=json"
|
||||
```
|
||||
|
||||
**Python example:**
|
||||
```python
|
||||
import requests
|
||||
|
||||
url = "http://localhost:8200/v1/audio/transcriptions"
|
||||
|
||||
with open("audio.mp3", "rb") as f:
|
||||
files = {"file": f}
|
||||
data = {
|
||||
"model": "large-v3-turbo",
|
||||
"language": "en",
|
||||
"response_format": "verbose_json"
|
||||
}
|
||||
|
||||
response = requests.post(url, files=files, data=data)
|
||||
result = response.json()
|
||||
|
||||
print(f"Transcription: {result['text']}")
|
||||
print(f"Duration: {result['duration']}s")
|
||||
print(f"Segments: {len(result['segments'])}")
|
||||
```
|
||||
|
||||
**JavaScript example:**
|
||||
```javascript
|
||||
const formData = new FormData();
|
||||
formData.append('file', audioFile);
|
||||
formData.append('model', 'large-v3-turbo');
|
||||
formData.append('response_format', 'json');
|
||||
|
||||
const response = await fetch('http://localhost:8200/v1/audio/transcriptions', {
|
||||
method: 'POST',
|
||||
body: formData
|
||||
});
|
||||
|
||||
const result = await response.json();
|
||||
console.log('Transcription:', result.text);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Error Responses
|
||||
|
||||
### Standard Error Format
|
||||
|
||||
```json
|
||||
{
|
||||
"detail": "Error message"
|
||||
}
|
||||
```
|
||||
|
||||
### Common HTTP Status Codes
|
||||
|
||||
| Code | Meaning | Description |
|
||||
|------|---------|-------------|
|
||||
| 200 | OK | Request successful |
|
||||
| 400 | Bad Request | Invalid request parameters |
|
||||
| 413 | Payload Too Large | File exceeds max size limit (default 500MB) |
|
||||
| 422 | Unprocessable Entity | Invalid file format or corrupted audio |
|
||||
| 500 | Internal Server Error | Transcription failed |
|
||||
| 503 | Service Unavailable | Model not loaded or GPU unavailable |
|
||||
|
||||
**Example error response:**
|
||||
```json
|
||||
{
|
||||
"detail": "File too large: 550.0MB (max: 500MB)"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Rate Limiting
|
||||
|
||||
No built-in rate limiting. Can be added via nginx or API gateway.
|
||||
|
||||
**Recommended nginx configuration:**
|
||||
```nginx
|
||||
limit_req_zone $binary_remote_addr zone=audio_limit:10m rate=10r/m;
|
||||
|
||||
location / {
|
||||
limit_req zone=audio_limit burst=5;
|
||||
proxy_pass http://audio-api:8200;
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## File Size Limits
|
||||
|
||||
- **Default:** 500MB per file
|
||||
- **Configurable via:** `AUDIO_MAX_FILE_SIZE_MB` environment variable
|
||||
- **Recommended:** Keep files under 100MB for best performance
|
||||
|
||||
---
|
||||
|
||||
## Supported Audio Formats
|
||||
|
||||
All formats supported by FFmpeg, including:
|
||||
|
||||
**Audio files:**
|
||||
- MP3, WAV, FLAC, OGG, M4A, AAC, WMA, OPUS
|
||||
|
||||
**Video files (audio track extraction):**
|
||||
- MP4, AVI, MKV, MOV, WEBM, FLV
|
||||
|
||||
**Professional formats:**
|
||||
- PCM, AIFF, AU, AMR
|
||||
|
||||
---
|
||||
|
||||
## Performance Considerations
|
||||
|
||||
### Processing Time
|
||||
|
||||
Approximate transcription times for large-v3-turbo with int8 on H200 GPU:
|
||||
|
||||
| Audio Length | Processing Time | Real-Time Factor |
|
||||
|--------------|-----------------|------------------|
|
||||
| 1 minute | ~3-5 seconds | 0.05-0.08x |
|
||||
| 10 minutes | ~30-50 seconds | 0.05-0.08x |
|
||||
| 1 hour | ~3-5 minutes | 0.05-0.08x |
|
||||
|
||||
**Real-Time Factor (RTF):** Processing time / Audio duration. Lower is better.
|
||||
- RTF < 0.1x = Excellent (8x faster than real-time)
|
||||
- RTF < 0.5x = Good (2x faster than real-time)
|
||||
- RTF = 1.0x = Real-time
|
||||
|
||||
### Optimization Tips
|
||||
|
||||
1. **Use language parameter** when known (skip auto-detection)
|
||||
2. **Use appropriate model size** based on accuracy needs
|
||||
3. **Process shorter segments** for long recordings (split at silence)
|
||||
4. **Monitor GPU utilization** with `nvidia-smi`
|
||||
|
||||
---
|
||||
|
||||
## Comparison with OpenAI API
|
||||
|
||||
This API is **fully compatible** with OpenAI's transcription endpoint:
|
||||
|
||||
| Feature | OpenAI | This API |
|
||||
|---------|--------|----------|
|
||||
| Endpoint | `/v1/audio/transcriptions` | `/v1/audio/transcriptions` ✅ |
|
||||
| Request format | `multipart/form-data` | `multipart/form-data` ✅ |
|
||||
| Response formats | `json`, `text`, `verbose_json` | `json`, `text`, `verbose_json` ✅ |
|
||||
| Language detection | Auto | Auto ✅ |
|
||||
| Max file size | 25MB | 500MB (configurable) ✅ |
|
||||
| Cost | $0.006/minute | Free (self-hosted) ✅ |
|
||||
| Latency | Variable | ~0.05-0.08x RTF ✅ |
|
||||
|
||||
**Drop-in replacement:** Change only the base URL to switch from OpenAI to this API.
|
||||
|
||||
---
|
||||
|
||||
## SDK Examples
|
||||
|
||||
### OpenAI Python SDK
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
# Point to local API
|
||||
client = OpenAI(
|
||||
api_key="not-needed", # No auth required
|
||||
base_url="http://localhost:8200/v1"
|
||||
)
|
||||
|
||||
with open("audio.mp3", "rb") as f:
|
||||
transcript = client.audio.transcriptions.create(
|
||||
model="large-v3-turbo",
|
||||
file=f,
|
||||
response_format="verbose_json"
|
||||
)
|
||||
|
||||
print(transcript.text)
|
||||
```
|
||||
|
||||
### OpenAI Node.js SDK
|
||||
|
||||
```javascript
|
||||
import OpenAI from 'openai';
|
||||
import fs from 'fs';
|
||||
|
||||
const openai = new OpenAI({
|
||||
apiKey: 'not-needed',
|
||||
baseURL: 'http://localhost:8200/v1'
|
||||
});
|
||||
|
||||
const transcription = await openai.audio.transcriptions.create({
|
||||
file: fs.createReadStream('audio.mp3'),
|
||||
model: 'large-v3-turbo',
|
||||
response_format: 'verbose_json'
|
||||
});
|
||||
|
||||
console.log(transcription.text);
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Advanced Features
|
||||
|
||||
### Voice Activity Detection (VAD)
|
||||
|
||||
Enabled by default. Automatically skips silence for:
|
||||
- Faster processing
|
||||
- Better accuracy
|
||||
- Smaller output
|
||||
|
||||
### Beam Search
|
||||
|
||||
Configurable via `AUDIO_BEAM_SIZE` (default: 5). Higher values = better accuracy but slower.
|
||||
|
||||
### Temperature Sampling
|
||||
|
||||
- `0.0` (default): Deterministic output
|
||||
- `0.0-1.0`: More creative/diverse outputs (less reliable)
|
||||
|
||||
---
|
||||
|
||||
## Monitoring
|
||||
|
||||
### Health Check
|
||||
|
||||
```bash
|
||||
# Check if API is ready
|
||||
curl http://localhost:8200/health
|
||||
|
||||
# Expected response
|
||||
{"status": "ok"}
|
||||
```
|
||||
|
||||
### Logs
|
||||
|
||||
```bash
|
||||
# View real-time logs
|
||||
cd deploy/
|
||||
docker compose logs -f audio-api
|
||||
|
||||
# Check for errors
|
||||
docker compose logs audio-api | grep ERROR
|
||||
```
|
||||
|
||||
### GPU Usage
|
||||
|
||||
```bash
|
||||
# Monitor GPU while transcribing
|
||||
watch -n 1 nvidia-smi
|
||||
|
||||
# Check VRAM usage
|
||||
nvidia-smi --query-gpu=memory.used,memory.total --format=csv
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### Error: File too large
|
||||
|
||||
**Solution:** Increase max file size or split audio into smaller chunks:
|
||||
```bash
|
||||
# Increase limit
|
||||
AUDIO_MAX_FILE_SIZE_MB=1000
|
||||
|
||||
# Or split with ffmpeg
|
||||
ffmpeg -i long_audio.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
|
||||
```
|
||||
|
||||
### Error: Transcription failed
|
||||
|
||||
**Possible causes:**
|
||||
1. Corrupted audio file - verify with media player
|
||||
2. Unsupported format - convert to MP3/WAV
|
||||
3. GPU out of memory - use smaller model or CPU mode
|
||||
4. Audio is pure noise/music - Whisper is designed for speech
|
||||
|
||||
### Slow performance
|
||||
|
||||
**Check:**
|
||||
1. GPU is being used: `docker exec audio-api nvidia-smi`
|
||||
2. Model is correct: Check logs for model loading messages
|
||||
3. Compute type is int8: Faster than float16
|
||||
|
||||
---
|
||||
|
||||
For more information, see [README.md](README.md) and the [faster-whisper documentation](https://github.com/SYSTRAN/faster-whisper).
|
||||
Loading…
Add table
Add a link
Reference in a new issue