Livrare LOT 1 - Didi
This commit is contained in:
commit
5380c3fc63
990 changed files with 133308 additions and 0 deletions
336
ai_platform/modules/embeddings/API.md
Normal file
336
ai_platform/modules/embeddings/API.md
Normal file
|
|
@ -0,0 +1,336 @@
|
|||
# Embeddings API Reference
|
||||
|
||||
OpenAI-compatible embeddings API with multiple backend support.
|
||||
|
||||
## Base URL
|
||||
|
||||
```
|
||||
{BASE_URL}
|
||||
```
|
||||
|
||||
- **Local development:** `http://localhost:54100`
|
||||
- **Docker (internal):** `http://didiAI-embeddings-api:14100`
|
||||
- **Production:** Use your configured hostname
|
||||
|
||||
## Authentication
|
||||
|
||||
Authentication is **optional**. If `EMB_API_TOKENS` is set, requests require a Bearer token:
|
||||
|
||||
```
|
||||
Authorization: Bearer <token>
|
||||
```
|
||||
|
||||
Health endpoints (`/health`, `/ready`) are always public.
|
||||
|
||||
## Endpoints
|
||||
|
||||
### Create Embeddings
|
||||
|
||||
Generate embeddings for the given input texts.
|
||||
|
||||
**Endpoint:** `POST /v1/embeddings`
|
||||
|
||||
**Request Headers:**
|
||||
| Header | Required | Description |
|
||||
|--------|----------|-------------|
|
||||
| `Content-Type` | Yes | Must be `application/json` |
|
||||
| `Authorization` | If auth enabled | `Bearer <token>` |
|
||||
|
||||
**Request Body:**
|
||||
```json
|
||||
{
|
||||
"input": "text to embed",
|
||||
"model": "BAAI/bge-m3",
|
||||
"encoding_format": "float",
|
||||
"dimensions": null,
|
||||
"backend": null
|
||||
}
|
||||
```
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|-------|------|----------|-------------|
|
||||
| `input` | string or string[] | Yes | Text(s) to embed |
|
||||
| `model` | string | Yes | Model identifier |
|
||||
| `encoding_format` | string | No | `"float"` (default) or `"base64"` |
|
||||
| `dimensions` | integer | No | Desired embedding dimensions (if supported) |
|
||||
| `backend` | string | No | Override default backend: `"vllm"` or `"llamacpp"` |
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"object": "embedding",
|
||||
"index": 0,
|
||||
"embedding": [0.0023, -0.0142, 0.0083, ...]
|
||||
}
|
||||
],
|
||||
"model": "BAAI/bge-m3",
|
||||
"usage": {
|
||||
"prompt_tokens": 5,
|
||||
"total_tokens": 5
|
||||
},
|
||||
"backend": "vllm"
|
||||
}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
curl -X POST http://localhost:54100/v1/embeddings \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"input": ["Hello world", "How are you?"],
|
||||
"model": "BAAI/bge-m3"
|
||||
}'
|
||||
```
|
||||
|
||||
**Multiple texts:**
|
||||
```bash
|
||||
curl -X POST http://localhost:54100/v1/embeddings \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"input": [
|
||||
"First document to embed",
|
||||
"Second document to embed",
|
||||
"Third document to embed"
|
||||
],
|
||||
"model": "BAAI/bge-m3"
|
||||
}'
|
||||
```
|
||||
|
||||
**Base64 encoding:**
|
||||
```bash
|
||||
curl -X POST http://localhost:54100/v1/embeddings \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"input": "Hello world",
|
||||
"model": "BAAI/bge-m3",
|
||||
"encoding_format": "base64"
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### List Models
|
||||
|
||||
List available embedding models.
|
||||
|
||||
**Endpoint:** `GET /v1/models`
|
||||
|
||||
**Query Parameters:**
|
||||
| Parameter | Type | Required | Description |
|
||||
|-----------|------|----------|-------------|
|
||||
| `backend` | string | No | Filter by backend: `"vllm"` or `"llamacpp"` |
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "BAAI/bge-m3",
|
||||
"backend": "vllm",
|
||||
"loaded": true,
|
||||
"dimensions": null,
|
||||
"max_input_tokens": null
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
# List all models
|
||||
curl http://localhost:54100/v1/models
|
||||
|
||||
# List models from specific backend
|
||||
curl "http://localhost:54100/v1/models?backend=vllm"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### List Backends
|
||||
|
||||
List available backends.
|
||||
|
||||
**Endpoint:** `GET /v1/backends`
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"backends": ["vllm", "llamacpp"]
|
||||
}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
curl http://localhost:54100/v1/backends
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Health Check
|
||||
|
||||
Detailed health status including per-backend health.
|
||||
|
||||
**Endpoint:** `GET /health`
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"status": "healthy",
|
||||
"backends": [
|
||||
{
|
||||
"name": "vllm",
|
||||
"healthy": true,
|
||||
"message": null
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Status values:
|
||||
- `"healthy"` - All backends are healthy
|
||||
- `"degraded"` - Some backends are unhealthy
|
||||
- `"unhealthy"` - All backends are unhealthy
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
curl http://localhost:54100/health
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Readiness Probe
|
||||
|
||||
Simple readiness check for Kubernetes.
|
||||
|
||||
**Endpoint:** `GET /ready`
|
||||
|
||||
**Response:**
|
||||
```json
|
||||
{
|
||||
"ready": true
|
||||
}
|
||||
```
|
||||
|
||||
**Example:**
|
||||
```bash
|
||||
curl http://localhost:54100/ready
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Error Responses
|
||||
|
||||
All errors follow this format:
|
||||
|
||||
```json
|
||||
{
|
||||
"detail": "Error message describing what went wrong"
|
||||
}
|
||||
```
|
||||
|
||||
### HTTP Status Codes
|
||||
|
||||
| Code | Meaning |
|
||||
|------|---------|
|
||||
| 200 | Success |
|
||||
| 400 | Bad request (invalid parameters, backend not enabled) |
|
||||
| 401 | Authentication required or failed |
|
||||
| 429 | Rate limit exceeded |
|
||||
| 503 | Service unavailable (backend connection failed) |
|
||||
| 504 | Gateway timeout (backend request timed out) |
|
||||
|
||||
### Rate Limit Response
|
||||
|
||||
```json
|
||||
{
|
||||
"detail": "Too many requests",
|
||||
"retry_after": 1.5
|
||||
}
|
||||
```
|
||||
|
||||
Headers include: `Retry-After: 2`
|
||||
|
||||
### Authentication Error
|
||||
|
||||
```json
|
||||
{
|
||||
"detail": {
|
||||
"error": "Authentication required",
|
||||
"message": "Missing Authorization header"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Request Headers
|
||||
|
||||
| Header | Required | Description |
|
||||
|--------|----------|-------------|
|
||||
| `Content-Type` | Yes (POST) | Must be `application/json` |
|
||||
| `Authorization` | If auth enabled | `Bearer <token>` |
|
||||
| `X-Request-ID` | No | Request tracking ID (generated if not provided) |
|
||||
|
||||
Response always includes `X-Request-ID` header for tracking.
|
||||
|
||||
---
|
||||
|
||||
## SDK Examples
|
||||
|
||||
### Python (httpx)
|
||||
|
||||
```python
|
||||
import httpx
|
||||
|
||||
async def embed_texts(texts: list[str]) -> list[list[float]]:
|
||||
async with httpx.AsyncClient() as client:
|
||||
response = await client.post(
|
||||
"http://localhost:54100/v1/embeddings",
|
||||
json={
|
||||
"input": texts,
|
||||
"model": "BAAI/bge-m3",
|
||||
},
|
||||
headers={"Authorization": "Bearer your-token"},
|
||||
)
|
||||
response.raise_for_status()
|
||||
data = response.json()
|
||||
return [item["embedding"] for item in data["data"]]
|
||||
```
|
||||
|
||||
### Python (openai SDK)
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:54100/v1",
|
||||
api_key="your-token", # or "not-needed" if auth disabled
|
||||
)
|
||||
|
||||
response = client.embeddings.create(
|
||||
input=["Hello world"],
|
||||
model="BAAI/bge-m3",
|
||||
)
|
||||
|
||||
embedding = response.data[0].embedding
|
||||
print(f"Embedding dimensions: {len(embedding)}")
|
||||
```
|
||||
|
||||
### curl
|
||||
|
||||
```bash
|
||||
# Simple embedding
|
||||
curl -X POST http://localhost:54100/v1/embeddings \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Authorization: Bearer your-token" \
|
||||
-d '{"input": "Hello world", "model": "BAAI/bge-m3"}'
|
||||
|
||||
# With specific backend
|
||||
curl -X POST http://localhost:54100/v1/embeddings \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{"input": "Hello world", "model": "BAAI/bge-m3", "backend": "vllm"}'
|
||||
```
|
||||
Loading…
Add table
Add a link
Reference in a new issue