16 KiB
LLM Inference API Documentation
OpenAI-compatible REST API for unified LLM inference across multiple backends.
Base URL
{BASE_URL}
Common configurations:
- Local development:
http://localhost:14011 - Docker (internal):
http://didiAI-llm-api:14011 - Docker (external):
http://<host>:14011 - VPN/Production: Use your configured hostname or IP
Authentication
Bearer Token Authentication (Optional)
The API supports optional Bearer token authentication. When enabled, protected endpoints require a valid token.
Enable by setting:
LLM_API_TOKENS=token1,token2,token3 # comma-separated for multiple tokens
Request header:
Authorization: Bearer <your-token>
Public endpoints (no auth required):
GET /healthGET /ready
Protected endpoints (require auth when enabled):
POST /v1/chat/completionsPOST /v1/completionsGET /v1/modelsPOST /v1/models/loadPOST /v1/models/unloadGET /v1/backends
Error response (401):
{
"detail": {
"error": "Authentication required",
"message": "Missing Authorization header"
}
}
Response includes WWW-Authenticate: Bearer header.
Backend API Keys
Backend services require their own API keys:
- LiteLLM:
OPENROUTER_API_KEY,OPENAI_API_KEY, orANTHROPIC_API_KEY - vLLM: Optional
LLM_VLLM_API_KEYfor vLLM server
Reasoning model note (
LLM_VLLM_DISABLE_THINKING, defaulttrue): the local modelQwen/Qwen3.5-35B-A3B(served asqwen3.5) is a reasoning model. By default the gateway injectschat_template_kwargs={"enable_thinking": false}on vLLM chat requests so the model returns the final answer directly instead of athinkingpreamble — important for callers that parse JSON. Callers may override by passing their ownchat_template_kwargsin the request body.
Rate Limiting
The API uses token bucket rate limiting:
| Header | Description |
|---|---|
Retry-After |
Seconds to wait when rate limited (429 response) |
X-Request-ID |
Unique request identifier (auto-generated or pass via header) |
Default limits:
- 10 requests/second with burst of 20
- 10 concurrent completion requests
Configure via environment variables:
LLM_RATE_LIMIT_RPS: Requests per secondLLM_RATE_LIMIT_BURST: Maximum burst sizeLLM_MAX_CONCURRENT_COMPLETIONS: Concurrent request limit
Endpoints
Chat Completions
Create a chat completion with optional streaming.
POST /v1/chat/completions
Request Body
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
messages |
array | Yes | - | List of messages (1-1000) |
model |
string | Yes | - | Model identifier |
temperature |
float | No | 0.7 | Sampling temperature (0.0-2.0) |
max_tokens |
integer | No | null | Maximum tokens to generate (1-1,000,000) |
stream |
boolean | No | false | Enable streaming response |
backend |
string | No | null | Backend override: litellm, vllm, llamacpp |
top_p |
float | No | null | Top-p sampling (0.0-1.0) |
frequency_penalty |
float | No | null | Frequency penalty (-2.0 to 2.0) |
presence_penalty |
float | No | null | Presence penalty (-2.0 to 2.0) |
stop |
string/array | No | null | Stop sequences |
Message Object
| Field | Type | Required | Description |
|---|---|---|---|
role |
string | Yes | One of: system, user, assistant, function, tool |
content |
string | Yes* | Message content (*can be null for assistant role) |
name |
string | No | Optional author name |
Response (Non-Streaming)
{
"id": "chatcmpl-abc123",
"object": "chat.completion",
"created": 1704067200,
"model": "gpt-3.5-turbo",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 15,
"total_tokens": 25
},
"backend": "litellm"
}
Response (Streaming)
Server-Sent Events (SSE) format. Each event contains a JSON chunk:
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{"role":"assistant"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{"content":"Hello"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{"content":"!"},"finish_reason":null}]}
data: {"id":"chatcmpl-abc123","object":"chat.completion.chunk","created":1704067200,"model":"gpt-3.5-turbo","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: [DONE]
Example Request
# Non-streaming
curl -X POST http://localhost:14011/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"}
],
"model": "gpt-3.5-turbo",
"temperature": 0.7,
"max_tokens": 100
}'
# Streaming
curl -X POST http://localhost:14011/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-d '{
"messages": [{"role": "user", "content": "Tell me a short story"}],
"model": "gpt-4",
"stream": true
}'
# With specific backend
curl -X POST http://localhost:14011/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [{"role": "user", "content": "Hello!"}],
"model": "meta-llama/Llama-2-7b-chat-hf",
"backend": "vllm"
}'
Text Completions (legacy)
Create a non-streaming text completion from a raw prompt. OpenAI-compatible legacy /v1/completions. Routes to the model's backend (vLLM primary, cloud via LiteLLM). Backends that do not support text completion return 501.
POST /v1/completions
Request Body
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
prompt |
string/array | Yes | - | Prompt(s) to complete |
model |
string | No | server default | Model (alias) to use |
temperature |
float | No | 0.7 | Sampling temperature (0.0-2.0) |
max_tokens |
integer | No | null | Maximum tokens to generate (1-1,000,000) |
backend |
string | No | null | Backend override: litellm, vllm, llamacpp |
top_p |
float | No | null | Top-p sampling (0.0-1.0) |
frequency_penalty |
float | No | null | Frequency penalty (-2.0 to 2.0) |
presence_penalty |
float | No | null | Presence penalty (-2.0 to 2.0) |
stop |
string/array | No | null | Stop sequences |
Response
{
"id": "cmpl-abc123",
"object": "text_completion",
"created": 1704067200,
"model": "qwen3.5",
"choices": [
{
"index": 0,
"text": "Paris is the capital of France.",
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 8,
"completion_tokens": 7,
"total_tokens": 15
},
"backend": "vllm"
}
Example Request
curl -X POST http://localhost:14011/v1/completions \
-H "Content-Type: application/json" \
-d '{
"prompt": "The capital of France is",
"model": "qwen3.5",
"max_tokens": 16
}'
List Models
List available models across all or specific backends.
GET /v1/models
GET /v1/models?backend=litellm
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
backend |
string | No | Filter by backend: litellm, vllm, llamacpp |
Response
{
"object": "list",
"data": [
{
"id": "gpt-3.5-turbo",
"backend": "litellm",
"loaded": true,
"context_length": 16384,
"capabilities": ["chat"]
},
{
"id": "gpt-4",
"backend": "litellm",
"loaded": true,
"context_length": 128000,
"capabilities": ["chat"]
},
{
"id": "meta-llama/Llama-2-7b-chat-hf",
"backend": "vllm",
"loaded": true,
"context_length": 4096,
"capabilities": ["chat"]
}
]
}
Example Request
# All models
curl http://localhost:14011/v1/models
# Models from specific backend
curl "http://localhost:14011/v1/models?backend=vllm"
Load Model
Load a model on a local backend (vLLM or llama.cpp only).
POST /v1/models/load
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Model identifier to load |
backend |
string | Yes | Target backend: vllm or llamacpp |
Response
{
"success": true,
"model": "meta-llama/Llama-2-7b-chat-hf",
"backend": "vllm",
"message": "Model loaded successfully"
}
Example Request
curl -X POST http://localhost:14011/v1/models/load \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-2-7b-chat-hf",
"backend": "vllm"
}'
Unload Model
Unload a model from a local backend (vLLM or llama.cpp only).
POST /v1/models/unload
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Model identifier to unload |
backend |
string | Yes | Backend to unload from: vllm or llamacpp |
Response
{
"success": true,
"model": "meta-llama/Llama-2-7b-chat-hf",
"backend": "vllm",
"message": "Model unloaded successfully"
}
Example Request
curl -X POST http://localhost:14011/v1/models/unload \
-H "Content-Type: application/json" \
-d '{
"model": "meta-llama/Llama-2-7b-chat-hf",
"backend": "vllm"
}'
List Backends
List available backends.
GET /v1/backends
Response
{
"backends": ["litellm", "vllm", "llamacpp"]
}
Example Request
curl http://localhost:14011/v1/backends
Health Check
Check overall system health and per-backend status.
GET /health
Response
{
"status": "healthy",
"backends": [
{
"name": "litellm",
"healthy": true,
"message": null
},
{
"name": "vllm",
"healthy": true,
"message": null
},
{
"name": "llamacpp",
"healthy": false,
"message": "Connection refused"
}
]
}
Status Values
| Status | Description |
|---|---|
healthy |
All backends are healthy |
degraded |
Some backends are healthy |
unhealthy |
No backends are healthy |
Example Request
curl http://localhost:14011/health
Readiness Probe
Kubernetes-style readiness probe. Returns ready if the default backend is healthy.
GET /ready
Response
{
"ready": true
}
HTTP Status Codes
| Code | Description |
|---|---|
| 200 | Service is ready |
| 503 | Service is not ready |
Example Request
curl http://localhost:14011/ready
Service Info (catalog)
Service metadata for cross-module catalog integration. Returns resource, models (from all enabled backends), and functions (API endpoint descriptors). Consumed by the catalog-api.
GET /v1/info
Example Request
curl http://localhost:14011/v1/info
Error Responses
All errors follow a consistent format:
{
"detail": "Error message description"
}
HTTP Status Codes
| Code | Description |
|---|---|
| 400 | Bad Request - Invalid parameters or backend |
| 401 | Unauthorized - Missing or invalid Bearer token (when auth enabled) |
| 422 | Validation Error - Request body validation failed |
| 429 | Too Many Requests - Rate limit exceeded |
| 500 | Internal Server Error - Completion or backend failure |
| 503 | Service Unavailable - Concurrency limit exceeded or service not ready |
Error Examples
Validation Error (422)
{
"detail": [
{
"type": "value_error",
"loc": ["body", "messages", 0, "content"],
"msg": "Message at index 0 with role 'user' cannot have empty content",
"input": ""
}
]
}
Rate Limit (429)
{
"detail": "Rate limit exceeded"
}
Response includes Retry-After header with seconds to wait.
Backend Error (400)
{
"detail": "Backend 'vllm' is not available"
}
Concurrency Limit (503)
{
"detail": "Concurrency limit exceeded"
}
Request Headers
| Header | Required | Description |
|---|---|---|
Content-Type |
Yes (POST) | Must be application/json |
Authorization |
When auth enabled | Bearer token: Bearer <your-token> |
Accept |
No | Use text/event-stream for streaming |
X-Request-ID |
No | Custom request ID (auto-generated if not provided) |
Supported Backends
LiteLLM (cloud)
Supports 100+ LLM providers through a unified interface. (LLM_DEFAULT_BACKEND is required and has no default; the DIDI deployment runs vllm with the local qwen3.5 model and uses LiteLLM for cloud/premium models.)
Popular models:
gpt-3.5-turbo,gpt-4,gpt-4-turbo(OpenAI)claude-3-opus,claude-3-sonnet,claude-3-haiku(Anthropic)openrouter/meta-llama/llama-3-70b(OpenRouter)
vLLM
High-throughput GPU inference for open-source models.
Requirements: NVIDIA GPU with 16GB+ VRAM
Example models:
meta-llama/Llama-2-7b-chat-hfmistralai/Mistral-7B-Instruct-v0.2microsoft/phi-2
llama.cpp
CPU/Metal inference for GGUF format models.
Requirements: 16GB+ RAM, CPU with AVX2 support
Example models: Local GGUF files
Python Client Example
import httpx
async def chat_completion():
async with httpx.AsyncClient() as client:
response = await client.post(
"http://localhost:14011/v1/chat/completions",
json={
"messages": [{"role": "user", "content": "Hello!"}],
"model": "gpt-3.5-turbo",
},
)
return response.json()
async def streaming_completion():
async with httpx.AsyncClient() as client:
async with client.stream(
"POST",
"http://localhost:14011/v1/chat/completions",
json={
"messages": [{"role": "user", "content": "Tell me a story"}],
"model": "gpt-4",
"stream": True,
},
) as response:
async for line in response.aiter_lines():
if line.startswith("data: "):
data = line[6:]
if data != "[DONE]":
chunk = json.loads(data)
content = chunk["choices"][0]["delta"].get("content", "")
print(content, end="", flush=True)
OpenAI SDK Compatibility
The API is compatible with the OpenAI Python SDK:
from openai import OpenAI
# Without auth (when LLM_API_TOKENS not set)
client = OpenAI(
base_url="http://localhost:14011/v1",
api_key="not-needed", # Required by SDK but not used
)
# With auth (when LLM_API_TOKENS is set)
client = OpenAI(
base_url="http://localhost:14011/v1",
api_key="your-api-token", # Your LLM_API_TOKENS value
)
response = client.chat.completions.create(
model="gpt-3.5-turbo",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
# Streaming
for chunk in client.chat.completions.create(
model="gpt-4",
messages=[{"role": "user", "content": "Tell me a story"}],
stream=True,
):
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")