didi-lot1-ai/ai_platform/modules/llm-inference
2026-06-27 06:42:02 -07:00
..
deploy LOT 1 - Optimizare script build -Instalare mono comanda 2026-06-27 06:42:02 -07:00
src/llm_inference LOT 1 - Optimizare script build -Instalare mono comanda 2026-06-27 06:42:02 -07:00
tests Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
.env.example Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
API.md LOT 1 - Optimizare script build -Instalare mono comanda 2026-06-27 06:42:02 -07:00
INDEX.md LOT 1 - Optimizare script build -Instalare mono comanda 2026-06-27 06:42:02 -07:00
pyproject.toml Livrare LOT 1 - Didi 2026-06-25 14:13:25 -07:00
README.md LOT 1 - Optimizare script build -Instalare mono comanda 2026-06-27 06:42:02 -07:00

LLM Inference

Unified LLM inference module with multiple backends: LiteLLM (OpenRouter, OpenAI, Anthropic), vLLM, and llama.cpp.

Prerequisites

Required:

  • All global prerequisites (see main README.md)
  • At least one LLM API key (OpenRouter, OpenAI, or Anthropic)

For vLLM backend (GPU inference):

  • NVIDIA GPU with 16GB+ VRAM (24GB+ recommended for larger models)
  • NVIDIA Driver 535+
  • NVIDIA Container Toolkit

For llama.cpp backend (CPU inference):

  • 16GB+ RAM (depends on model size)
  • CPU with AVX2 support (most modern CPUs)

Features

  • Multiple Backends: LiteLLM (100+ providers), vLLM (GPU inference), llama.cpp (CPU/Metal)
  • Unified Interface: Single API for all backends, switch seamlessly
  • Streaming Support: Server-Sent Events (SSE) for real-time responses
  • Model Management: List, load, and unload models (local backends)
  • OpenAI Compatible: Drop-in replacement for OpenAI API clients

Installation

cd modules/llm-inference

# Install core dependencies
uv sync

# Install with optional backends
uv sync --extra vllm       # For vLLM support
uv sync --extra llamacpp   # For llama.cpp support
uv sync --extra local      # For all local backends
uv sync --extra dev        # For development

Quick Start

As a Python Library

from llm_inference import LLMClient, BackendType

# Initialize client
client = LLMClient()

# Simple completion
response = await client.complete(
    messages=[{"role": "user", "content": "Hello!"}],
    model="gpt-3.5-turbo",
)
print(response.choices[0].message.content)

# Streaming
async for chunk in client.stream(
    messages=[{"role": "user", "content": "Tell me a story"}],
    model="gpt-4",
):
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

# Use specific backend
response = await client.complete(
    messages=[{"role": "user", "content": "Hello!"}],
    model="meta-llama/Llama-2-7b-chat-hf",
    backend=BackendType.VLLM,
)

As an API Server

cd deploy/

# Copy and configure environment
cp ../.env.example .env
# Edit .env with your API keys and settings

# Start the server
./deploy.sh --profile api -d

API Endpoints

Endpoint Method Description
/v1/chat/completions POST Chat completion (supports streaming)
/v1/completions POST Legacy text completion (non-streaming; 501 if backend unsupported)
/v1/models GET List available models
/v1/models/load POST Load a model (local backends)
/v1/models/unload POST Unload a model (local backends)
/v1/backends GET List available backends
/v1/info GET Service/catalog metadata
/health GET Health check
/ready GET Readiness probe

Example API Request

# Non-streaming
curl -X POST http://localhost:14011/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Hello!"}],
    "model": "gpt-3.5-turbo"
  }'

# Streaming
curl -X POST http://localhost:14011/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [{"role": "user", "content": "Tell me a story"}],
    "model": "gpt-4",
    "stream": true
  }'

Configuration

Configure via environment variables (prefix: LLM_):

Variable Default Description
LLM_DEFAULT_BACKEND required Default backend (litellm, vllm, llamacpp) — no default, must be set; the DIDI deployment runs vllm
LLM_DEFAULT_MODEL qwen3.5 Default model alias used when a request omits model
LLM_MODEL_ALIASES {} JSON map of alias → served id, e.g. {"qwen3.5":"Qwen/Qwen3.5-35B-A3B"}
LLM_ENABLE_FALLBACK true Cross-backend fallback cascade on backend failure
LLM_FALLBACK_ORDER [vllm,llamacpp,litellm] Backend order tried in the fallback cascade
LLM_PORT 14011 API server port
LLM_HOST 0.0.0.0 API server host
LLM_ENABLE_VLLM required Enable vLLM backend (no default)
LLM_ENABLE_LLAMACPP required Enable llama.cpp backend (no default)
LLM_VLLM_BASE_URL http://localhost:14001 vLLM server URL
LLM_VLLM_DISABLE_THINKING true Inject enable_thinking=false for the Qwen3.5 reasoning model so it returns the final answer directly (no thinking preamble)
LLM_LLAMACPP_BASE_URL http://localhost:8080 llama.cpp server URL
OPENROUTER_API_KEY - OpenRouter API key
OPENAI_API_KEY - OpenAI API key
ANTHROPIC_API_KEY - Anthropic API key

Rate Limiting

The API server includes built-in rate limiting using a token bucket algorithm. Configure via environment variables:

Variable Default Description
LLM_RATE_LIMIT_RPS 10.0 Requests per second
LLM_RATE_LIMIT_BURST 20 Maximum burst size
LLM_MAX_CONCURRENT_COMPLETIONS 10 Maximum concurrent completion requests

Important: Rate limiting is per-process. In multi-replica deployments (e.g., Kubernetes), each replica has its own independent limit. For distributed rate limiting, use an external solution like Redis, an API gateway (Kong, nginx), or cloud provider rate limiting.

Deployment

cd deploy/

# Copy and configure environment (REQUIRED)
cp ../.env.example .env
# Edit .env with your settings

# API only (uses external LLM providers)
./deploy.sh --profile api -d

# With vLLM (requires NVIDIA GPU)
./deploy.sh --profile vllm -d

# With llama.cpp (CPU inference)
./deploy.sh --profile llamacpp -d

# Full stack
./deploy.sh --profile full -d

# View logs
./deploy.sh --profile api --logs

# Stop services
./deploy.sh --profile api --down

Port Allocation

Port Service
14011 LLM Inference API
14001 vLLM Qwen3.5-35B-A3B

Development

# Install dev dependencies
uv sync --extra dev

# Run tests
uv run pytest

# Run tests with coverage
uv run pytest --cov=src/llm_inference --cov-report=term-missing

# Lint and format
uv run ruff check .
uv run ruff format .

# Type check
uv run mypy src/

Architecture

src/llm_inference/
├── __init__.py           # Package exports
├── config.py             # Pydantic Settings
├── types.py              # Core types and enums
├── schemas.py            # API request/response schemas
├── exceptions.py         # Custom exceptions
├── client.py             # High-level LLMClient
├── cli.py                # CLI entry point
├── backends/
│   ├── base.py           # Abstract LLMBackend
│   ├── registry.py       # BackendRegistry
│   ├── litellm_backend.py
│   ├── vllm_backend.py
│   └── llamacpp_backend.py
└── api/
    ├── app.py            # FastAPI app factory
    ├── dependencies.py
    └── routes/
        ├── completions.py
        ├── models.py
        └── health.py

License

MIT