didi-lot1-ai/ai_platform/modules/rerank/README.md

4.7 KiB

Rerank Module

Cohere/Jina-compatible reranking API with support for vLLM and llama.cpp backends.

Features

  • Cohere/Jina-compatible API: Drop-in replacement for Cohere's /v1/rerank endpoint
  • Multiple backends: Support for vLLM and llama.cpp
  • High performance: Built on FastAPI with async support
  • Production-ready: Rate limiting, authentication, health checks

Prerequisites

Required

  • Linux - Ubuntu 22.04+ or similar
  • Python 3.10+ - Managed via uv
  • uv - Fast Python package manager

Optional (for backends)

  • vLLM - Requires NVIDIA GPU with CUDA 12.x
  • llama.cpp - Can run on CPU or GPU

Quick Start

1. Install dependencies

cd modules/rerank
uv sync --all-extras

2. Configure environment

cp .env.example .env
# Edit .env with your settings

3. Start backend server

vLLM (GPU):

python -m vllm.entrypoints.openai.api_server \
    --model BAAI/bge-reranker-v2-m3 \
    --host 0.0.0.0 --port 54201 \
    --task score

llama.cpp (CPU):

llama-server \
    --model /models/bge-reranker-v2-m3-q4_k_m.gguf \
    --host 0.0.0.0 --port 54210 \
    --reranking

4. Start API server

# Set required environment variables
export RERANK_DEFAULT_BACKEND=vllm
export RERANK_ENABLE_VLLM=true
export RERANK_ENABLE_LLAMACPP=false
export RERANK_EXTERNAL_URL=http://localhost:54200

# Run the server
uv run python -m rerank.cli --port 54200

5. Test the API

curl -X POST http://localhost:54200/v1/rerank \
  -H "Content-Type: application/json" \
  -d '{
    "model": "BAAI/bge-reranker-v2-m3",
    "query": "What is machine learning?",
    "documents": [
      "Machine learning is a subset of AI",
      "Cats are pets",
      "Deep learning uses neural networks"
    ]
  }'

Docker Deployment

cd modules/rerank/deploy

# Copy and configure .env
cp ../.env.example .env
# Edit .env with your settings

# Start with vLLM backend
./deploy.sh --profile vllm -d

# Or start with llama.cpp backend
./deploy.sh --profile llamacpp -d

# View logs
./deploy.sh --profile vllm --logs

# Stop
./deploy.sh --profile vllm --down

Python Library Usage

from rerank import RerankClient

# Initialize client (reads config from environment)
client = RerankClient()

# Rerank documents
response = await client.rerank(
    query="What is machine learning?",
    documents=[
        "ML is a subset of AI",
        "Cats are pets",
        "Deep learning uses neural networks",
    ],
    model="BAAI/bge-reranker-v2-m3",
)

# Access results
for result in response.results:
    print(f"Index {result.index}: {result.relevance_score:.4f}")

# Get top N results
response = await client.rerank(
    query="programming",
    documents=["Python", "Java", "Cooking", "C++"],
    model="BAAI/bge-reranker-v2-m3",
    top_n=2,
)

# Include documents in response
response = await client.rerank(
    query="AI",
    documents=["Machine learning", "Deep learning"],
    model="BAAI/bge-reranker-v2-m3",
    return_documents=True,
)
for result in response.results:
    print(f"{result.document}: {result.relevance_score:.4f}")

Configuration

All configuration is via environment variables with the RERANK_ prefix:

Variable Required Default Description
RERANK_DEFAULT_BACKEND Yes - Default backend: vllm or llamacpp
RERANK_ENABLE_VLLM Yes - Enable vLLM backend
RERANK_ENABLE_LLAMACPP Yes - Enable llama.cpp backend
RERANK_EXTERNAL_URL Yes - External URL for OpenAPI spec
RERANK_PORT No 54200 API server port
RERANK_VLLM_BASE_URL No http://localhost:54201 vLLM server URL
RERANK_LLAMACPP_BASE_URL No http://localhost:54210 llama.cpp server URL
RERANK_API_TOKENS No - Comma-separated API tokens
RERANK_RATE_LIMIT_RPS No 20.0 Requests per second limit
RERANK_MAX_CONCURRENT_RERANKS No 20 Max concurrent requests

See .env.example for the complete list.

Port Allocation

Following the datacenter port schema (x42xx = Reranking):

Port Service Environment
14200 Rerank API Production
54200 Rerank API Development
14201 vLLM Rerank Server Production
54201 vLLM Rerank Server Development
14210 llama.cpp Rerank Server Production
54210 llama.cpp Rerank Server Development

Development

# Install dev dependencies
uv sync --all-extras

# Run tests
uv run pytest

# Run tests with coverage
uv run pytest --cov=src/rerank --cov-report=term-missing

# Lint and format
uv run ruff check .
uv run ruff format .

# Type check
uv run mypy src/

API Reference

See API.md for the complete API documentation.