Livrare LOT 1 - Didi
This commit is contained in:
commit
5380c3fc63
990 changed files with 133308 additions and 0 deletions
611
ai_platform/modules/llm-inference/API.md
Normal file
611
ai_platform/modules/llm-inference/API.md
Normal file
|
|
@ -0,0 +1,611 @@
|
|||
# LLM Inference API Documentation
|
||||
|
||||
OpenAI-compatible REST API for unified LLM inference across multiple backends.
|
||||
|
||||
## Base URL
|
||||
|
||||
```
|
||||
{BASE_URL}
|
||||
```
|
||||
|
||||
Common configurations:
|
||||
- **Local development:** `http://localhost:14011`
|
||||
- **Docker (internal):** `http://didiAI-llm-api:14011`
|
||||
- **Docker (external):** `http://<host>:14011`
|
||||
- **VPN/Production:** Use your configured hostname or IP
|
||||
|
||||
## Authentication
|
||||
|
||||
### Bearer Token Authentication (Optional)
|
||||
|
||||
The API supports optional Bearer token authentication. When enabled, protected endpoints require a valid token.
|
||||
|
||||
**Enable by setting:**
|
||||
```bash
|
||||
LLM_API_TOKENS=token1,token2,token3 # comma-separated for multiple tokens
|
||||
```
|
||||
|
||||
**Request header:**
|
||||
```
|
||||
Authorization: Bearer <your-token>
|
||||
```
|
||||
|
||||
**Public endpoints (no auth required):**
|
||||
- `GET /health`
|
||||
- `GET /ready`
|
||||
|
||||
**Protected endpoints (require auth when enabled):**
|
||||
- `POST /v1/chat/completions`
|
||||
- `GET /v1/models`
|
||||
- `POST /v1/models/load`
|
||||
- `POST /v1/models/unload`
|
||||
- `GET /v1/backends`
|
||||
|
||||
**Error response (401):**
|
||||
```json
|
||||
{
|
||||
"detail": {
|
||||
"error": "Authentication required",
|
||||
"message": "Missing Authorization header"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
Response includes `WWW-Authenticate: Bearer` header.
|
||||
|
||||
### Backend API Keys
|
||||
|
||||
Backend services require their own API keys:
|
||||
- **LiteLLM**: `OPENROUTER_API_KEY`, `OPENAI_API_KEY`, or `ANTHROPIC_API_KEY`
|
||||
- **vLLM**: Optional `LLM_VLLM_API_KEY` for vLLM server
|
||||
|
||||
## Rate Limiting
|
||||
|
||||
The API uses token bucket rate limiting:
|
||||
|
||||
| Header | Description |
|
||||
|--------|-------------|
|
||||
| `Retry-After` | Seconds to wait when rate limited (429 response) |
|
||||
| `X-Request-ID` | Unique request identifier (auto-generated or pass via header) |
|
||||
|
||||
Default limits:
|
||||
- **10 requests/second** with burst of 20
|
||||
- **10 concurrent completion requests**
|
||||
|
||||
Configure via environment variables:
|
||||
- `LLM_RATE_LIMIT_RPS`: Requests per second
|
||||
- `LLM_RATE_LIMIT_BURST`: Maximum burst size
|
||||
- `LLM_MAX_CONCURRENT_COMPLETIONS`: Concurrent request limit
|
||||
|
||||
---
|
||||
|
||||
## Endpoints
|
||||
|
||||
### Chat Completions
|
||||
|
||||
Create a chat completion with optional streaming.
|
||||
|
||||
```
|
||||
POST /v1/chat/completions
|
||||
```
|
||||
|
||||
#### Request Body
|
||||
|
||||
| Field | Type | Required | Default | Description |
|
||||
|-------|------|----------|---------|-------------|
|
||||
| `messages` | array | Yes | - | List of messages (1-1000) |
|
||||
| `model` | string | Yes | - | Model identifier |
|
||||
| `temperature` | float | No | 0.7 | Sampling temperature (0.0-2.0) |
|
||||
| `max_tokens` | integer | No | null | Maximum tokens to generate (1-1,000,000) |
|
||||
| `stream` | boolean | No | false | Enable streaming response |
|
||||
| `backend` | string | No | null | Backend override: `litellm`, `vllm`, `llamacpp` |
|
||||
| `top_p` | float | No | null | Top-p sampling (0.0-1.0) |
|
||||
| `frequency_penalty` | float | No | null | Frequency penalty (-2.0 to 2.0) |
|
||||
| `presence_penalty` | float | No | null | Presence penalty (-2.0 to 2.0) |
|
||||
| `stop` | string/array | No | null | Stop sequences |
|
||||
|
||||
#### Message Object
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|-------|------|----------|-------------|
|
||||
| `role` | string | Yes | One of: `system`, `user`, `assistant`, `function`, `tool` |
|
||||
| `content` | string | Yes* | Message content (*can be null for assistant role) |
|
||||
| `name` | string | No | Optional author name |
|
||||
|
||||
#### Response (Non-Streaming)
|
||||
|
||||
```json
|
||||
{
|
||||
"id": "chatcmpl-abc123",
|
||||
"object": "chat.completion",
|
||||
"created": 1704067200,
|
||||
"model": "gpt-3.5-turbo",
|
||||
"choices": [
|
||||
{
|
||||
"index": 0,
|
||||
"message": {
|
||||
"role": "assistant",
|
||||
"content": "Hello! How can I help you today?"
|
||||
},
|
||||
"finish_reason": "stop"
|
||||
}
|
||||
],
|
||||
"usage": {
|
||||
"prompt_tokens": 10,
|
||||
"completion_tokens": 15,
|
||||
"total_tokens": 25
|
||||
},
|
||||
"backend": "litellm"
|
||||
}
|
||||
```
|
||||
|
||||
#### Response (Streaming)
|
||||
|
||||
Server-Sent Events (SSE) format. Each event contains a JSON chunk:
|
||||
|
||||
```
|
||||
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
|
||||
|
||||
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
|
||||
|
||||
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}
|
||||
|
||||
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
|
||||
|
||||
data: [DONE]
|
||||
```
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
# Non-streaming
|
||||
curl -X POST http://localhost:14011/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"messages": [
|
||||
{"role": "system", "content": "You are a helpful assistant."},
|
||||
{"role": "user", "content": "What is the capital of France?"}
|
||||
],
|
||||
"model": "gpt-3.5-turbo",
|
||||
"temperature": 0.7,
|
||||
"max_tokens": 100
|
||||
}'
|
||||
|
||||
# Streaming
|
||||
curl -X POST http://localhost:14011/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-H "Accept: text/event-stream" \
|
||||
-d '{
|
||||
"messages": [{"role": "user", "content": "Tell me a short story"}],
|
||||
"model": "gpt-4",
|
||||
"stream": true
|
||||
}'
|
||||
|
||||
# With specific backend
|
||||
curl -X POST http://localhost:14011/v1/chat/completions \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"messages": [{"role": "user", "content": "Hello!"}],
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"backend": "vllm"
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### List Models
|
||||
|
||||
List available models across all or specific backends.
|
||||
|
||||
```
|
||||
GET /v1/models
|
||||
GET /v1/models?backend=litellm
|
||||
```
|
||||
|
||||
#### Query Parameters
|
||||
|
||||
| Parameter | Type | Required | Description |
|
||||
|-----------|------|----------|-------------|
|
||||
| `backend` | string | No | Filter by backend: `litellm`, `vllm`, `llamacpp` |
|
||||
|
||||
#### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"object": "list",
|
||||
"data": [
|
||||
{
|
||||
"id": "gpt-3.5-turbo",
|
||||
"backend": "litellm",
|
||||
"loaded": true,
|
||||
"context_length": 16384,
|
||||
"capabilities": ["chat"]
|
||||
},
|
||||
{
|
||||
"id": "gpt-4",
|
||||
"backend": "litellm",
|
||||
"loaded": true,
|
||||
"context_length": 128000,
|
||||
"capabilities": ["chat"]
|
||||
},
|
||||
{
|
||||
"id": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"backend": "vllm",
|
||||
"loaded": true,
|
||||
"context_length": 4096,
|
||||
"capabilities": ["chat"]
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
# All models
|
||||
curl http://localhost:14011/v1/models
|
||||
|
||||
# Models from specific backend
|
||||
curl "http://localhost:14011/v1/models?backend=vllm"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Load Model
|
||||
|
||||
Load a model on a local backend (vLLM or llama.cpp only).
|
||||
|
||||
```
|
||||
POST /v1/models/load
|
||||
```
|
||||
|
||||
#### Request Body
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|-------|------|----------|-------------|
|
||||
| `model` | string | Yes | Model identifier to load |
|
||||
| `backend` | string | Yes | Target backend: `vllm` or `llamacpp` |
|
||||
|
||||
#### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"backend": "vllm",
|
||||
"message": "Model loaded successfully"
|
||||
}
|
||||
```
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:14011/v1/models/load \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"backend": "vllm"
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Unload Model
|
||||
|
||||
Unload a model from a local backend (vLLM or llama.cpp only).
|
||||
|
||||
```
|
||||
POST /v1/models/unload
|
||||
```
|
||||
|
||||
#### Request Body
|
||||
|
||||
| Field | Type | Required | Description |
|
||||
|-------|------|----------|-------------|
|
||||
| `model` | string | Yes | Model identifier to unload |
|
||||
| `backend` | string | Yes | Backend to unload from: `vllm` or `llamacpp` |
|
||||
|
||||
#### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"success": true,
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"backend": "vllm",
|
||||
"message": "Model unloaded successfully"
|
||||
}
|
||||
```
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
curl -X POST http://localhost:14011/v1/models/unload \
|
||||
-H "Content-Type: application/json" \
|
||||
-d '{
|
||||
"model": "meta-llama/Llama-2-7b-chat-hf",
|
||||
"backend": "vllm"
|
||||
}'
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### List Backends
|
||||
|
||||
List available backends.
|
||||
|
||||
```
|
||||
GET /v1/backends
|
||||
```
|
||||
|
||||
#### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"backends": ["litellm", "vllm", "llamacpp"]
|
||||
}
|
||||
```
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
curl http://localhost:14011/v1/backends
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Health Check
|
||||
|
||||
Check overall system health and per-backend status.
|
||||
|
||||
```
|
||||
GET /health
|
||||
```
|
||||
|
||||
#### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"status": "healthy",
|
||||
"backends": [
|
||||
{
|
||||
"name": "litellm",
|
||||
"healthy": true,
|
||||
"message": null
|
||||
},
|
||||
{
|
||||
"name": "vllm",
|
||||
"healthy": true,
|
||||
"message": null
|
||||
},
|
||||
{
|
||||
"name": "llamacpp",
|
||||
"healthy": false,
|
||||
"message": "Connection refused"
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
#### Status Values
|
||||
|
||||
| Status | Description |
|
||||
|--------|-------------|
|
||||
| `healthy` | All backends are healthy |
|
||||
| `degraded` | Some backends are healthy |
|
||||
| `unhealthy` | No backends are healthy |
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
curl http://localhost:14011/health
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
### Readiness Probe
|
||||
|
||||
Kubernetes-style readiness probe. Returns ready if the default backend is healthy.
|
||||
|
||||
```
|
||||
GET /ready
|
||||
```
|
||||
|
||||
#### Response
|
||||
|
||||
```json
|
||||
{
|
||||
"ready": true
|
||||
}
|
||||
```
|
||||
|
||||
#### HTTP Status Codes
|
||||
|
||||
| Code | Description |
|
||||
|------|-------------|
|
||||
| 200 | Service is ready |
|
||||
| 503 | Service is not ready |
|
||||
|
||||
#### Example Request
|
||||
|
||||
```bash
|
||||
curl http://localhost:14011/ready
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Error Responses
|
||||
|
||||
All errors follow a consistent format:
|
||||
|
||||
```json
|
||||
{
|
||||
"detail": "Error message description"
|
||||
}
|
||||
```
|
||||
|
||||
### HTTP Status Codes
|
||||
|
||||
| Code | Description |
|
||||
|------|-------------|
|
||||
| 400 | Bad Request - Invalid parameters or backend |
|
||||
| 401 | Unauthorized - Missing or invalid Bearer token (when auth enabled) |
|
||||
| 422 | Validation Error - Request body validation failed |
|
||||
| 429 | Too Many Requests - Rate limit exceeded |
|
||||
| 500 | Internal Server Error - Completion or backend failure |
|
||||
| 503 | Service Unavailable - Concurrency limit exceeded or service not ready |
|
||||
|
||||
### Error Examples
|
||||
|
||||
**Validation Error (422)**
|
||||
```json
|
||||
{
|
||||
"detail": [
|
||||
{
|
||||
"type": "value_error",
|
||||
"loc": ["body", "messages", 0, "content"],
|
||||
"msg": "Message at index 0 with role 'user' cannot have empty content",
|
||||
"input": ""
|
||||
}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
**Rate Limit (429)**
|
||||
```json
|
||||
{
|
||||
"detail": "Rate limit exceeded"
|
||||
}
|
||||
```
|
||||
Response includes `Retry-After` header with seconds to wait.
|
||||
|
||||
**Backend Error (400)**
|
||||
```json
|
||||
{
|
||||
"detail": "Backend 'vllm' is not available"
|
||||
}
|
||||
```
|
||||
|
||||
**Concurrency Limit (503)**
|
||||
```json
|
||||
{
|
||||
"detail": "Concurrency limit exceeded"
|
||||
}
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Request Headers
|
||||
|
||||
| Header | Required | Description |
|
||||
|--------|----------|-------------|
|
||||
| `Content-Type` | Yes (POST) | Must be `application/json` |
|
||||
| `Authorization` | When auth enabled | Bearer token: `Bearer <your-token>` |
|
||||
| `Accept` | No | Use `text/event-stream` for streaming |
|
||||
| `X-Request-ID` | No | Custom request ID (auto-generated if not provided) |
|
||||
|
||||
---
|
||||
|
||||
## Supported Backends
|
||||
|
||||
### LiteLLM (Default)
|
||||
|
||||
Supports 100+ LLM providers through a unified interface.
|
||||
|
||||
**Popular models:**
|
||||
- `gpt-3.5-turbo`, `gpt-4`, `gpt-4-turbo` (OpenAI)
|
||||
- `claude-3-opus`, `claude-3-sonnet`, `claude-3-haiku` (Anthropic)
|
||||
- `openrouter/meta-llama/llama-3-70b` (OpenRouter)
|
||||
|
||||
### vLLM
|
||||
|
||||
High-throughput GPU inference for open-source models.
|
||||
|
||||
**Requirements:** NVIDIA GPU with 16GB+ VRAM
|
||||
|
||||
**Example models:**
|
||||
- `meta-llama/Llama-2-7b-chat-hf`
|
||||
- `mistralai/Mistral-7B-Instruct-v0.2`
|
||||
- `microsoft/phi-2`
|
||||
|
||||
### llama.cpp
|
||||
|
||||
CPU/Metal inference for GGUF format models.
|
||||
|
||||
**Requirements:** 16GB+ RAM, CPU with AVX2 support
|
||||
|
||||
**Example models:** Local GGUF files
|
||||
|
||||
---
|
||||
|
||||
## Python Client Example
|
||||
|
||||
```python
|
||||
import httpx
|
||||
|
||||
async def chat_completion():
|
||||
async with httpx.AsyncClient() as client:
|
||||
response = await client.post(
|
||||
"http://localhost:14011/v1/chat/completions",
|
||||
json={
|
||||
"messages": [{"role": "user", "content": "Hello!"}],
|
||||
"model": "gpt-3.5-turbo",
|
||||
},
|
||||
)
|
||||
return response.json()
|
||||
|
||||
async def streaming_completion():
|
||||
async with httpx.AsyncClient() as client:
|
||||
async with client.stream(
|
||||
"POST",
|
||||
"http://localhost:14011/v1/chat/completions",
|
||||
json={
|
||||
"messages": [{"role": "user", "content": "Tell me a story"}],
|
||||
"model": "gpt-4",
|
||||
"stream": True,
|
||||
},
|
||||
) as response:
|
||||
async for line in response.aiter_lines():
|
||||
if line.startswith("data: "):
|
||||
data = line[6:]
|
||||
if data != "[DONE]":
|
||||
chunk = json.loads(data)
|
||||
content = chunk["choices"][0]["delta"].get("content", "")
|
||||
print(content, end="", flush=True)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## OpenAI SDK Compatibility
|
||||
|
||||
The API is compatible with the OpenAI Python SDK:
|
||||
|
||||
```python
|
||||
from openai import OpenAI
|
||||
|
||||
# Without auth (when LLM_API_TOKENS not set)
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:14011/v1",
|
||||
api_key="not-needed", # Required by SDK but not used
|
||||
)
|
||||
|
||||
# With auth (when LLM_API_TOKENS is set)
|
||||
client = OpenAI(
|
||||
base_url="http://localhost:14011/v1",
|
||||
api_key="your-api-token", # Your LLM_API_TOKENS value
|
||||
)
|
||||
|
||||
response = client.chat.completions.create(
|
||||
model="gpt-3.5-turbo",
|
||||
messages=[{"role": "user", "content": "Hello!"}],
|
||||
)
|
||||
print(response.choices[0].message.content)
|
||||
|
||||
# Streaming
|
||||
for chunk in client.chat.completions.create(
|
||||
model="gpt-4",
|
||||
messages=[{"role": "user", "content": "Tell me a story"}],
|
||||
stream=True,
|
||||
):
|
||||
if chunk.choices[0].delta.content:
|
||||
print(chunk.choices[0].delta.content, end="")
|
||||
```
|
||||
Loading…
Add table
Add a link
Reference in a new issue