didi-lot1-ai/ai_platform/modules/embeddings/API.md

336 lines
6.3 KiB
Markdown

# Embeddings API Reference
OpenAI-compatible embeddings API with multiple backend support.
## Base URL
```
{BASE_URL}
```
- **Local development:** `http://localhost:54100`
- **Docker (internal):** `http://didiAI-embeddings-api:14100`
- **Production:** Use your configured hostname
## Authentication
Authentication is **optional**. If `EMB_API_TOKENS` is set, requests require a Bearer token:
```
Authorization: Bearer <token>
```
Health endpoints (`/health`, `/ready`) are always public.
## Endpoints
### Create Embeddings
Generate embeddings for the given input texts.
**Endpoint:** `POST /v1/embeddings`
**Request Headers:**
| Header | Required | Description |
|--------|----------|-------------|
| `Content-Type` | Yes | Must be `application/json` |
| `Authorization` | If auth enabled | `Bearer <token>` |
**Request Body:**
```json
{
"input": "text to embed",
"model": "BAAI/bge-m3",
"encoding_format": "float",
"dimensions": null,
"backend": null
}
```
| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `input` | string or string[] | Yes | Text(s) to embed |
| `model` | string | Yes | Model identifier |
| `encoding_format` | string | No | `"float"` (default) or `"base64"` |
| `dimensions` | integer | No | Desired embedding dimensions (if supported) |
| `backend` | string | No | Override default backend: `"vllm"` or `"llamacpp"` |
**Response:**
```json
{
"object": "list",
"data": [
{
"object": "embedding",
"index": 0,
"embedding": [0.0023, -0.0142, 0.0083, ...]
}
],
"model": "BAAI/bge-m3",
"usage": {
"prompt_tokens": 5,
"total_tokens": 5
},
"backend": "vllm"
}
```
**Example:**
```bash
curl -X POST http://localhost:54100/v1/embeddings \
-H "Content-Type: application/json" \
-d '{
"input": ["Hello world", "How are you?"],
"model": "BAAI/bge-m3"
}'
```
**Multiple texts:**
```bash
curl -X POST http://localhost:54100/v1/embeddings \
-H "Content-Type: application/json" \
-d '{
"input": [
"First document to embed",
"Second document to embed",
"Third document to embed"
],
"model": "BAAI/bge-m3"
}'
```
**Base64 encoding:**
```bash
curl -X POST http://localhost:54100/v1/embeddings \
-H "Content-Type: application/json" \
-d '{
"input": "Hello world",
"model": "BAAI/bge-m3",
"encoding_format": "base64"
}'
```
---
### List Models
List available embedding models.
**Endpoint:** `GET /v1/models`
**Query Parameters:**
| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `backend` | string | No | Filter by backend: `"vllm"` or `"llamacpp"` |
**Response:**
```json
{
"object": "list",
"data": [
{
"id": "BAAI/bge-m3",
"backend": "vllm",
"loaded": true,
"dimensions": null,
"max_input_tokens": null
}
]
}
```
**Example:**
```bash
# List all models
curl http://localhost:54100/v1/models
# List models from specific backend
curl "http://localhost:54100/v1/models?backend=vllm"
```
---
### List Backends
List available backends.
**Endpoint:** `GET /v1/backends`
**Response:**
```json
{
"backends": ["vllm", "llamacpp"]
}
```
**Example:**
```bash
curl http://localhost:54100/v1/backends
```
---
### Health Check
Detailed health status including per-backend health.
**Endpoint:** `GET /health`
**Response:**
```json
{
"status": "healthy",
"backends": [
{
"name": "vllm",
"healthy": true,
"message": null
}
]
}
```
Status values:
- `"healthy"` - All backends are healthy
- `"degraded"` - Some backends are unhealthy
- `"unhealthy"` - All backends are unhealthy
**Example:**
```bash
curl http://localhost:54100/health
```
---
### Readiness Probe
Simple readiness check for Kubernetes.
**Endpoint:** `GET /ready`
**Response:**
```json
{
"ready": true
}
```
**Example:**
```bash
curl http://localhost:54100/ready
```
---
## Error Responses
All errors follow this format:
```json
{
"detail": "Error message describing what went wrong"
}
```
### HTTP Status Codes
| Code | Meaning |
|------|---------|
| 200 | Success |
| 400 | Bad request (invalid parameters, backend not enabled) |
| 401 | Authentication required or failed |
| 429 | Rate limit exceeded |
| 503 | Service unavailable (backend connection failed) |
| 504 | Gateway timeout (backend request timed out) |
### Rate Limit Response
```json
{
"detail": "Too many requests",
"retry_after": 1.5
}
```
Headers include: `Retry-After: 2`
### Authentication Error
```json
{
"detail": {
"error": "Authentication required",
"message": "Missing Authorization header"
}
}
```
---
## Request Headers
| Header | Required | Description |
|--------|----------|-------------|
| `Content-Type` | Yes (POST) | Must be `application/json` |
| `Authorization` | If auth enabled | `Bearer <token>` |
| `X-Request-ID` | No | Request tracking ID (generated if not provided) |
Response always includes `X-Request-ID` header for tracking.
---
## SDK Examples
### Python (httpx)
```python
import httpx
async def embed_texts(texts: list[str]) -> list[list[float]]:
async with httpx.AsyncClient() as client:
response = await client.post(
"http://localhost:54100/v1/embeddings",
json={
"input": texts,
"model": "BAAI/bge-m3",
},
headers={"Authorization": "Bearer your-token"},
)
response.raise_for_status()
data = response.json()
return [item["embedding"] for item in data["data"]]
```
### Python (openai SDK)
```python
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:54100/v1",
api_key="your-token", # or "not-needed" if auth disabled
)
response = client.embeddings.create(
input=["Hello world"],
model="BAAI/bge-m3",
)
embedding = response.data[0].embedding
print(f"Embedding dimensions: {len(embedding)}")
```
### curl
```bash
# Simple embedding
curl -X POST http://localhost:54100/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer your-token" \
-d '{"input": "Hello world", "model": "BAAI/bge-m3"}'
# With specific backend
curl -X POST http://localhost:54100/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"input": "Hello world", "model": "BAAI/bge-m3", "backend": "vllm"}'
```