didi-lot1-ai/ai_platform/modules/rerank/INDEX.md

12 KiB
Raw Blame History

Rerank Service — INDEX

Cross-encoder reranking API for DIDI semantic search precision. Serves the BAAI/bge-reranker-v2-m3 model behind a Cohere/Jina-compatible HTTP API. Used after kNN retrieval to score (query, candidate) pairs and reorder evidence by true relevance, instead of relying solely on pgvector cosine similarity from embeddings.

  • Stack: Python 3.10+, FastAPI, Uvicorn, Pydantic v2, httpx, vLLM (or llama.cpp) backend
  • URLs:
    • Dev API: http://10.11.10.12:54200
    • Prod API: http://10.11.10.12:14200
    • vLLM internal: :54201 (Dev) / :14201 (Prod)
    • llama.cpp internal: :54210 (Dev) / :14210 (Prod)
  • Containers: didiAI-rerank-api (FastAPI wrapper) + didiAI-rerank-vllm (or didiAI-rerank-llamacpp)
  • Model: BAAI/bge-reranker-v2-m3 (multilingual cross-encoder, ~8K context window)
  • Port schema: x42xx family (Reranking)

Ce face

Re-scores candidate documents against a query using a cross-encoder model — directly attending over both texts together — for higher precision than embedding cosine similarity alone. The cross-encoder gives a true relevance score in [0..1].

Pipeline role inside DIDI:

  1. didi-brain /v1/gather runs an initial kNN retrieval over pgvector embeddings (dim 1024, separate embeddings module).
  2. The top-K candidates (typically 50200) are passed to this rerank service.
  3. The reranker scores each (query, candidate) pair and returns a relevance-sorted list.
  4. didi-brain takes the top-N (e.g. top 10) reranked items as final evidence.

This two-stage retrieval (kNN → cross-encoder rerank) is the standard high-quality semantic search pattern: embeddings handle scale, the cross-encoder handles precision.


API endpoints

Cohere/Jina-compatible. Auth is optional (Bearer token) — enabled if RERANK_API_TOKENS is set. Health endpoints stay public.

POST /v1/rerank (alias POST /v2/rerank)

Rerank a list of documents against a query.

Request body:

{
  "model": "BAAI/bge-reranker-v2-m3",
  "query": "What is machine learning?",
  "documents": ["Machine learning is...", "Cats are pets", "Deep learning..."],
  "top_n": 3,
  "return_documents": false,
  "backend": null
}
Field Type Required Description
model string yes Model identifier (e.g. BAAI/bge-reranker-v2-m3)
query string yes Search query
documents string[] yes Documents to rerank (11000)
top_n int no Return only top N results (default: all)
return_documents bool no Echo document text in response
backend string no Override default: "vllm" or "llamacpp"

Response:

{
  "id": "rerank-abc123def456",
  "model": "BAAI/bge-reranker-v2-m3",
  "results": [
    {"index": 2, "relevance_score": 0.9523, "document": null},
    {"index": 0, "relevance_score": 0.8876, "document": null}
  ],
  "usage": {"total_tokens": 150},
  "backend": "vllm"
}

GET /v1/models

List loaded models. Optional ?backend=vllm|llamacpp filter.

GET /v1/backends

List enabled backends, e.g. {"backends": ["vllm", "llamacpp"]}.

GET /health

Detailed per-backend health. Status is healthy / degraded / unhealthy.

GET /ready

Simple K8s-style readiness probe — returns {"ready": true}.

Error codes

400 bad request · 401 auth failed · 429 rate-limited (with Retry-After header) · 503 backend unavailable · 504 backend timeout. All errors return {"detail": "..."}. Every response carries an X-Request-ID header for tracing.

Full reference in API.md.


Backends

The API is a thin FastAPI router in front of one of two cross-encoder servers. Both backends speak OpenAI-compatible HTTP, so the wrapper unifies them.

Backend When Pros Cons
vLLM Production, GPU host High throughput, batched scoring, lowest latency under load Requires NVIDIA GPU + CUDA 12.x
llama.cpp Dev / CPU fallback Runs on CPU or modest GPU, GGUF quantized models, low memory Lower throughput

The wrapper routes requests via a backend registry. Set RERANK_DEFAULT_BACKEND to choose, or override per-request with the backend field in the body. Both can be enabled simultaneously (RERANK_ENABLE_VLLM=true, RERANK_ENABLE_LLAMACPP=true).

vLLM is launched with --task score (vLLM 0.8.x cross-encoder mode). llama.cpp is launched with --reranking.


How didi-brain uses it

didi-brain's gather/retrieval flow:

  1. services/gather.py builds an initial candidate set via pgvector kNN over the embeddings module (1024-dim vectors).
  2. Calls the reranker via shared/reranker_client.py — typically wrapping the /v1/rerank endpoint with the brain's HTTPS/auth config.
  3. Picks the top-N reranked candidates as final evidence chunks.
  4. Passes them to the LLM as grounded context.

The reranker is therefore in the critical path of every brain /v1/gather call. Its latency budget is small (a few hundred ms), so vLLM batching matters in production.


Structura fisiere

rerank/
├── README.md                # Quick start, install, env vars table
├── API.md                   # Full HTTP API reference + curl/Python examples
├── INDEX.md                 # This file
├── pyproject.toml           # Package metadata, deps (fastapi, httpx, openai SDK)
├── .env.example             # All RERANK_* env vars documented
├── deploy/
│   ├── Dockerfile           # API wrapper image
│   ├── docker-compose.yml   # api / vllm / llamacpp profiles
│   └── deploy.sh            # Helper: ./deploy.sh --profile vllm -d
├── src/rerank/
│   ├── __init__.py          # Public exports (RerankClient, ...)
│   ├── cli.py               # `rerank` entrypoint — `python -m rerank.cli`
│   ├── client.py            # RerankClient (Python SDK to call this API)
│   ├── config.py            # Pydantic Settings, RERANK_ env prefix
│   ├── schemas.py           # Pydantic request/response models
│   ├── exceptions.py        # Custom error types
│   ├── logging.py           # Structured logging setup
│   ├── types.py             # Backend literal types
│   ├── api/
│   │   ├── app.py           # FastAPI app factory
│   │   ├── dependencies.py  # Auth + rate-limit DI
│   │   ├── middleware.py    # X-Request-ID, rate limit, error handlers
│   │   └── routes/
│   │       ├── rerank.py    # POST /v1/rerank, /v2/rerank
│   │       ├── models.py    # GET /v1/models, /v1/backends
│   │       └── health.py    # GET /health, /ready
│   └── backends/
│       ├── base.py          # Abstract BackendBase (rerank, health, list_models)
│       ├── vllm_backend.py  # vLLM via OpenAI SDK (--task score)
│       ├── llamacpp_backend.py # llama.cpp --reranking endpoint
│       └── registry.py      # Backend registry + selector
└── tests/
    ├── conftest.py
    ├── test_config.py
    └── test_schemas.py

Configuration

All env vars use the RERANK_ prefix and are read via Pydantic Settings (config.py).

Required (no defaults):

Var Example
RERANK_DEFAULT_BACKEND vllm or llamacpp
RERANK_ENABLE_VLLM true / false
RERANK_ENABLE_LLAMACPP true / false
RERANK_EXTERNAL_URL http://localhost:14200 (used in OpenAPI spec)

API server:

Var Default Notes
RERANK_HOST 0.0.0.0
RERANK_PORT 14200 Prod 14200 / Dev 54200
RERANK_API_TOKENS (empty) Comma-separated; if empty, auth is disabled
RERANK_RATE_LIMIT_RPS 20.0
RERANK_RATE_LIMIT_BURST 50
RERANK_MAX_CONCURRENT_RERANKS 20
RERANK_REQUEST_TIMEOUT 120.0 s
RERANK_CONNECT_TIMEOUT 10.0 s
RERANK_LOG_LEVEL INFO
RERANK_LOG_JSON false

Backend base URLs (app settings — fields on RerankSettings):

Var Default Notes
RERANK_VLLM_BASE_URL http://localhost:54201 Where the API wrapper reaches vLLM
RERANK_VLLM_API_KEY (empty) Optional API key for the vLLM server
RERANK_LLAMACPP_BASE_URL http://localhost:54210 Where the API wrapper reaches llama.cpp

Only *_BASE_URL / *_API_KEY are read by the API app (config.py). The model id, port, GPU and context knobs below are not RerankSettings fields.

Backend service variables (docker-compose only):

These are consumed by deploy/docker-compose.yml to launch the vLLM / llama.cpp containers — they configure the backend server, not the API app.

Var Default Description
RERANK_VLLM_MODEL BAAI/bge-reranker-v2-m3 HF model id loaded by the vLLM container
RERANK_VLLM_PORT 14201 Host port mapped to the vLLM container
RERANK_VLLM_GPU 0 CUDA_VISIBLE_DEVICES for the vLLM container
RERANK_VLLM_GPU_UTIL 0.50 vLLM --gpu-memory-utilization
RERANK_VLLM_MAX_LEN 8192 vLLM --max-model-len
RERANK_LLAMACPP_MODEL bge-reranker-v2-m3-q4_k_m.gguf GGUF filename inside MODELS_DIR
RERANK_LLAMACPP_PORT 14210 Host port mapped to the llama.cpp container
RERANK_LLAMACPP_CTX 8192 llama.cpp context size
RERANK_LLAMACPP_THREADS 4 llama.cpp threads
RERANK_LLAMACPP_PARALLEL 4 llama.cpp parallel slots
MODELS_DIR /cai2_ds_storage/models GGUF model directory for llama.cpp
HF_CACHE_DIR /cai2_ds_storage/hf_cache Mounted into the vLLM container
HF_TOKEN (unset) Forwarded as HUGGING_FACE_HUB_TOKEN

Full list with comments in .env.example.


Performance / model

  • Model: BAAI/bge-reranker-v2-m3 — multilingual cross-encoder, ~8K context window, ~568M params.
  • Output: single relevance score per (query, document) pair, normalized to [0..1].
  • Expected latency: roughly 100300 ms for ~20 candidates per query on a single mid-range GPU with vLLM batching; scales sublinearly with batch size.
  • Throughput: dominated by GPU memory and RERANK_VLLM_GPU_UTIL / RERANK_MAX_CONCURRENT_RERANKS. vLLM continuous batching helps a lot under concurrent load.
  • Memory: vLLM defaults to 50% GPU memory util (RERANK_VLLM_GPU_UTIL=0.50), so it can co-host on a GPU shared with other modules. llama.cpp Q4_K_M quantization fits comfortably on CPU.

Deployment

Docker Compose with three profiles in deploy/docker-compose.yml:

cd /home/admin365/didi_mono/ai_platform/modules/rerank/deploy

# Production (GPU + vLLM)
./deploy.sh --profile vllm -d

# Lightweight / no-GPU
./deploy.sh --profile llamacpp -d

# API only (external rerank servers running elsewhere)
./deploy.sh --profile api -d

# Logs / stop
./deploy.sh --profile vllm --logs
./deploy.sh --profile vllm --down

Containers:

  • didiAI-rerank-api — FastAPI wrapper, port RERANK_PORT (14200/54200)
  • didiAI-rerank-vllmvllm/vllm-openai:v0.8.5, GPU device pinned via RERANK_VLLM_GPU, internal port 14201
  • didiAI-rerank-llamacppghcr.io/ggml-org/llama.cpp:server-b4769, internal port 8080 → host 14210/54210

All services share the external didi-network Docker network so other DIDI modules (didi-brain, embeddings) can reach them by container name.

Health gates: API has a 30-s /health healthcheck; vLLM has a 5-min start period to allow model load on first boot.


  • Embeddings module (ai_platform/modules/embeddings) — separate service, dim 1024, used for the initial pgvector kNN stage that feeds this reranker.
  • didi-brain (ai_platform/modules/didi_brain/brain_api) — main consumer; services/gather.py calls rerank after kNN; shared/reranker_client.py is the helper.
  • GPU host: 10.11.10.17 / 10.11.10.12 (see project network architecture doc).