didi-lot1-ai/ai_platform/modules/embeddings/README.md

4.1 KiB

Embeddings Module

OpenAI-compatible embeddings API with support for vLLM and llama.cpp backends.

Features

  • OpenAI-compatible API: Drop-in replacement for OpenAI's /v1/embeddings endpoint
  • Multiple backends: Support for vLLM and llama.cpp
  • High performance: Built on FastAPI with async support
  • Production-ready: Rate limiting, authentication, health checks

Prerequisites

Required

  • Linux - Ubuntu 22.04+ or similar
  • Python 3.10+ - Managed via uv
  • uv - Fast Python package manager

Optional (for backends)

  • vLLM - Requires NVIDIA GPU with CUDA 12.x
  • llama.cpp - Can run on CPU or GPU

Quick Start

1. Install dependencies

cd modules/embeddings
uv sync --all-extras

2. Configure environment

cp .env.example .env
# Edit .env with your settings

3. Start backend server

vLLM (GPU):

python -m vllm.entrypoints.openai.api_server \
    --model BAAI/bge-m3 \
    --host 0.0.0.0 --port 54101 \
    --task embed

llama.cpp (CPU):

llama-server \
    --model /models/bge-m3-q4_k_m.gguf \
    --host 0.0.0.0 --port 54110 \
    --embedding

4. Start API server

# Set required environment variables
export EMB_DEFAULT_BACKEND=vllm
export EMB_ENABLE_VLLM=true
export EMB_ENABLE_LLAMACPP=false
export EMB_EXTERNAL_URL=http://localhost:54100

# Run the server
uv run python -m embeddings.cli --port 54100

5. Test the API

curl -X POST http://localhost:54100/v1/embeddings \
  -H "Content-Type: application/json" \
  -d '{
    "input": "Hello, world!",
    "model": "BAAI/bge-m3"
  }'

Docker Deployment

cd modules/embeddings/deploy

# Copy and configure .env
cp ../.env.example .env
# Edit .env with your settings

# Start with vLLM backend
./deploy.sh --profile vllm -d

# Or start with llama.cpp backend
./deploy.sh --profile llamacpp -d

# View logs
./deploy.sh --profile vllm --logs

# Stop
./deploy.sh --profile vllm --down

Python Library Usage

from embeddings import EmbeddingClient

# Initialize client (reads config from environment)
client = EmbeddingClient()

# Generate embeddings
response = await client.embed(
    texts=["Hello, world!", "How are you?"],
    model="BAAI/bge-m3",
)

# Access embeddings
for item in response.data:
    print(f"Index {item.index}: {len(item.embedding)} dimensions")

# List available models
models = await client.list_models()
for model in models:
    print(f"{model.id} on {model.backend}")

Configuration

All configuration is via environment variables with the EMB_ prefix:

Variable Required Default Description
EMB_DEFAULT_BACKEND Yes - Default backend: vllm or llamacpp
EMB_ENABLE_VLLM Yes - Enable vLLM backend
EMB_ENABLE_LLAMACPP Yes - Enable llama.cpp backend
EMB_EXTERNAL_URL Yes - External URL for OpenAPI spec
EMB_PORT No 54100 API server port
EMB_VLLM_BASE_URL No http://localhost:54101 vLLM server URL
EMB_LLAMACPP_BASE_URL No http://localhost:54110 llama.cpp server URL
EMB_API_TOKENS No - Comma-separated API tokens
EMB_RATE_LIMIT_RPS No 20.0 Requests per second limit
EMB_MAX_CONCURRENT_REQUESTS No 20 Max concurrent requests

See .env.example for the complete list.

Port Allocation

Following the datacenter port schema (x41xx = Embeddings):

Port Service Environment
14100 Embeddings API Production
54100 Embeddings API Development
14101 vLLM Embed Server Production
54101 vLLM Embed Server Development
14110 llama.cpp Embed Server Production
54110 llama.cpp Embed Server Development

Development

# Install dev dependencies
uv sync --all-extras

# Run tests
uv run pytest

# Run tests with coverage
uv run pytest --cov=src/embeddings --cov-report=term-missing

# Lint and format
uv run ruff check .
uv run ruff format .

# Type check
uv run mypy src/

API Reference

See API.md for the complete API documentation.