# Embeddings API Reference OpenAI-compatible embeddings API with multiple backend support. ## Base URL ``` {BASE_URL} ``` - **Local development:** `http://localhost:54100` - **Docker (internal):** `http://didiAI-embeddings-api:14100` - **Production:** Use your configured hostname ## Authentication Authentication is **optional**. If `EMB_API_TOKENS` is set, requests require a Bearer token: ``` Authorization: Bearer ``` Health endpoints (`/health`, `/ready`) are always public. ## Endpoints ### Create Embeddings Generate embeddings for the given input texts. **Endpoint:** `POST /v1/embeddings` **Request Headers:** | Header | Required | Description | |--------|----------|-------------| | `Content-Type` | Yes | Must be `application/json` | | `Authorization` | If auth enabled | `Bearer ` | **Request Body:** ```json { "input": "text to embed", "model": "BAAI/bge-m3", "encoding_format": "float", "dimensions": null, "backend": null } ``` | Field | Type | Required | Description | |-------|------|----------|-------------| | `input` | string or string[] | Yes | Text(s) to embed | | `model` | string | Yes | Model identifier | | `encoding_format` | string | No | `"float"` (default) or `"base64"` | | `dimensions` | integer | No | Desired embedding dimensions (if supported) | | `backend` | string | No | Override default backend: `"vllm"` or `"llamacpp"` | **Response:** ```json { "object": "list", "data": [ { "object": "embedding", "index": 0, "embedding": [0.0023, -0.0142, 0.0083, ...] } ], "model": "BAAI/bge-m3", "usage": { "prompt_tokens": 5, "total_tokens": 5 }, "backend": "vllm" } ``` **Example:** ```bash curl -X POST http://localhost:54100/v1/embeddings \ -H "Content-Type: application/json" \ -d '{ "input": ["Hello world", "How are you?"], "model": "BAAI/bge-m3" }' ``` **Multiple texts:** ```bash curl -X POST http://localhost:54100/v1/embeddings \ -H "Content-Type: application/json" \ -d '{ "input": [ "First document to embed", "Second document to embed", "Third document to embed" ], "model": "BAAI/bge-m3" }' ``` **Base64 encoding:** ```bash curl -X POST http://localhost:54100/v1/embeddings \ -H "Content-Type: application/json" \ -d '{ "input": "Hello world", "model": "BAAI/bge-m3", "encoding_format": "base64" }' ``` --- ### List Models List available embedding models. **Endpoint:** `GET /v1/models` **Query Parameters:** | Parameter | Type | Required | Description | |-----------|------|----------|-------------| | `backend` | string | No | Filter by backend: `"vllm"` or `"llamacpp"` | **Response:** ```json { "object": "list", "data": [ { "id": "BAAI/bge-m3", "backend": "vllm", "loaded": true, "dimensions": null, "max_input_tokens": null } ] } ``` **Example:** ```bash # List all models curl http://localhost:54100/v1/models # List models from specific backend curl "http://localhost:54100/v1/models?backend=vllm" ``` --- ### List Backends List available backends. **Endpoint:** `GET /v1/backends` **Response:** ```json { "backends": ["vllm", "llamacpp"] } ``` **Example:** ```bash curl http://localhost:54100/v1/backends ``` --- ### Health Check Detailed health status including per-backend health. **Endpoint:** `GET /health` **Response:** ```json { "status": "healthy", "backends": [ { "name": "vllm", "healthy": true, "message": null } ] } ``` Status values: - `"healthy"` - All backends are healthy - `"degraded"` - Some backends are unhealthy - `"unhealthy"` - All backends are unhealthy **Example:** ```bash curl http://localhost:54100/health ``` --- ### Readiness Probe Simple readiness check for Kubernetes. **Endpoint:** `GET /ready` **Response:** ```json { "ready": true } ``` **Example:** ```bash curl http://localhost:54100/ready ``` --- ## Error Responses All errors follow this format: ```json { "detail": "Error message describing what went wrong" } ``` ### HTTP Status Codes | Code | Meaning | |------|---------| | 200 | Success | | 400 | Bad request (invalid parameters, backend not enabled) | | 401 | Authentication required or failed | | 429 | Rate limit exceeded | | 503 | Service unavailable (backend connection failed) | | 504 | Gateway timeout (backend request timed out) | ### Rate Limit Response ```json { "detail": "Too many requests", "retry_after": 1.5 } ``` Headers include: `Retry-After: 2` ### Authentication Error ```json { "detail": { "error": "Authentication required", "message": "Missing Authorization header" } } ``` --- ## Request Headers | Header | Required | Description | |--------|----------|-------------| | `Content-Type` | Yes (POST) | Must be `application/json` | | `Authorization` | If auth enabled | `Bearer ` | | `X-Request-ID` | No | Request tracking ID (generated if not provided) | Response always includes `X-Request-ID` header for tracking. --- ## SDK Examples ### Python (httpx) ```python import httpx async def embed_texts(texts: list[str]) -> list[list[float]]: async with httpx.AsyncClient() as client: response = await client.post( "http://localhost:54100/v1/embeddings", json={ "input": texts, "model": "BAAI/bge-m3", }, headers={"Authorization": "Bearer your-token"}, ) response.raise_for_status() data = response.json() return [item["embedding"] for item in data["data"]] ``` ### Python (openai SDK) ```python from openai import OpenAI client = OpenAI( base_url="http://localhost:54100/v1", api_key="your-token", # or "not-needed" if auth disabled ) response = client.embeddings.create( input=["Hello world"], model="BAAI/bge-m3", ) embedding = response.data[0].embedding print(f"Embedding dimensions: {len(embedding)}") ``` ### curl ```bash # Simple embedding curl -X POST http://localhost:54100/v1/embeddings \ -H "Content-Type: application/json" \ -H "Authorization: Bearer your-token" \ -d '{"input": "Hello world", "model": "BAAI/bge-m3"}' # With specific backend curl -X POST http://localhost:54100/v1/embeddings \ -H "Content-Type: application/json" \ -d '{"input": "Hello world", "model": "BAAI/bge-m3", "backend": "vllm"}' ```