278 lines
12 KiB
Markdown
278 lines
12 KiB
Markdown
# Rerank Service — INDEX
|
||
|
||
Cross-encoder reranking API for DIDI semantic search precision. Serves the `BAAI/bge-reranker-v2-m3` model behind a Cohere/Jina-compatible HTTP API. Used after kNN retrieval to score `(query, candidate)` pairs and reorder evidence by true relevance, instead of relying solely on pgvector cosine similarity from embeddings.
|
||
|
||
- **Stack:** Python 3.10+, FastAPI, Uvicorn, Pydantic v2, httpx, vLLM (or llama.cpp) backend
|
||
- **URLs:**
|
||
- Dev API: `http://10.11.10.12:54200`
|
||
- Prod API: `http://10.11.10.12:14200`
|
||
- vLLM internal: `:54201` (Dev) / `:14201` (Prod)
|
||
- llama.cpp internal: `:54210` (Dev) / `:14210` (Prod)
|
||
- **Containers:** `didiAI-rerank-api` (FastAPI wrapper) + `didiAI-rerank-vllm` (or `didiAI-rerank-llamacpp`)
|
||
- **Model:** `BAAI/bge-reranker-v2-m3` (multilingual cross-encoder, ~8K context window)
|
||
- **Port schema:** `x42xx` family (Reranking)
|
||
|
||
---
|
||
|
||
## Ce face
|
||
|
||
Re-scores candidate documents against a query using a cross-encoder model — directly attending over both texts together — for higher precision than embedding cosine similarity alone. The cross-encoder gives a true relevance score in `[0..1]`.
|
||
|
||
**Pipeline role inside DIDI:**
|
||
|
||
1. didi-brain `/v1/gather` runs an initial kNN retrieval over pgvector embeddings (dim 1024, separate `embeddings` module).
|
||
2. The top-K candidates (typically 50–200) are passed to this rerank service.
|
||
3. The reranker scores each `(query, candidate)` pair and returns a relevance-sorted list.
|
||
4. didi-brain takes the top-N (e.g. top 10) reranked items as final evidence.
|
||
|
||
This two-stage retrieval (kNN → cross-encoder rerank) is the standard high-quality semantic search pattern: embeddings handle scale, the cross-encoder handles precision.
|
||
|
||
---
|
||
|
||
## API endpoints
|
||
|
||
Cohere/Jina-compatible. Auth is optional (Bearer token) — enabled if `RERANK_API_TOKENS` is set. Health endpoints stay public.
|
||
|
||
### `POST /v1/rerank` (alias `POST /v2/rerank`)
|
||
|
||
Rerank a list of documents against a query.
|
||
|
||
**Request body:**
|
||
```json
|
||
{
|
||
"model": "BAAI/bge-reranker-v2-m3",
|
||
"query": "What is machine learning?",
|
||
"documents": ["Machine learning is...", "Cats are pets", "Deep learning..."],
|
||
"top_n": 3,
|
||
"return_documents": false,
|
||
"backend": null
|
||
}
|
||
```
|
||
|
||
| Field | Type | Required | Description |
|
||
|-------|------|----------|-------------|
|
||
| `model` | string | yes | Model identifier (e.g. `BAAI/bge-reranker-v2-m3`) |
|
||
| `query` | string | yes | Search query |
|
||
| `documents` | string[] | yes | Documents to rerank (1–1000) |
|
||
| `top_n` | int | no | Return only top N results (default: all) |
|
||
| `return_documents` | bool | no | Echo document text in response |
|
||
| `backend` | string | no | Override default: `"vllm"` or `"llamacpp"` |
|
||
|
||
**Response:**
|
||
```json
|
||
{
|
||
"id": "rerank-abc123def456",
|
||
"model": "BAAI/bge-reranker-v2-m3",
|
||
"results": [
|
||
{"index": 2, "relevance_score": 0.9523, "document": null},
|
||
{"index": 0, "relevance_score": 0.8876, "document": null}
|
||
],
|
||
"usage": {"total_tokens": 150},
|
||
"backend": "vllm"
|
||
}
|
||
```
|
||
|
||
### `GET /v1/models`
|
||
|
||
List loaded models. Optional `?backend=vllm|llamacpp` filter.
|
||
|
||
### `GET /v1/backends`
|
||
|
||
List enabled backends, e.g. `{"backends": ["vllm", "llamacpp"]}`.
|
||
|
||
### `GET /health`
|
||
|
||
Detailed per-backend health. Status is `healthy` / `degraded` / `unhealthy`.
|
||
|
||
### `GET /ready`
|
||
|
||
Simple K8s-style readiness probe — returns `{"ready": true}`.
|
||
|
||
### Error codes
|
||
|
||
`400` bad request · `401` auth failed · `429` rate-limited (with `Retry-After` header) · `503` backend unavailable · `504` backend timeout. All errors return `{"detail": "..."}`. Every response carries an `X-Request-ID` header for tracing.
|
||
|
||
Full reference in `API.md`.
|
||
|
||
---
|
||
|
||
## Backends
|
||
|
||
The API is a thin FastAPI router in front of one of two cross-encoder servers. Both backends speak OpenAI-compatible HTTP, so the wrapper unifies them.
|
||
|
||
| Backend | When | Pros | Cons |
|
||
|---------|------|------|------|
|
||
| **vLLM** | Production, GPU host | High throughput, batched scoring, lowest latency under load | Requires NVIDIA GPU + CUDA 12.x |
|
||
| **llama.cpp** | Dev / CPU fallback | Runs on CPU or modest GPU, GGUF quantized models, low memory | Lower throughput |
|
||
|
||
The wrapper routes requests via a backend registry. Set `RERANK_DEFAULT_BACKEND` to choose, or override per-request with the `backend` field in the body. Both can be enabled simultaneously (`RERANK_ENABLE_VLLM=true`, `RERANK_ENABLE_LLAMACPP=true`).
|
||
|
||
vLLM is launched with `--task score` (vLLM 0.8.x cross-encoder mode). llama.cpp is launched with `--reranking`.
|
||
|
||
---
|
||
|
||
## How didi-brain uses it
|
||
|
||
didi-brain's gather/retrieval flow:
|
||
|
||
1. `services/gather.py` builds an initial candidate set via pgvector kNN over the `embeddings` module (1024-dim vectors).
|
||
2. Calls the reranker via `shared/reranker_client.py` — typically wrapping the `/v1/rerank` endpoint with the brain's HTTPS/auth config.
|
||
3. Picks the top-N reranked candidates as final evidence chunks.
|
||
4. Passes them to the LLM as grounded context.
|
||
|
||
The reranker is therefore in the critical path of every brain `/v1/gather` call. Its latency budget is small (a few hundred ms), so vLLM batching matters in production.
|
||
|
||
---
|
||
|
||
## Structura fisiere
|
||
|
||
```
|
||
rerank/
|
||
├── README.md # Quick start, install, env vars table
|
||
├── API.md # Full HTTP API reference + curl/Python examples
|
||
├── INDEX.md # This file
|
||
├── pyproject.toml # Package metadata, deps (fastapi, httpx, openai SDK)
|
||
├── .env.example # All RERANK_* env vars documented
|
||
├── deploy/
|
||
│ ├── Dockerfile # API wrapper image
|
||
│ ├── docker-compose.yml # api / vllm / llamacpp profiles
|
||
│ └── deploy.sh # Helper: ./deploy.sh --profile vllm -d
|
||
├── src/rerank/
|
||
│ ├── __init__.py # Public exports (RerankClient, ...)
|
||
│ ├── cli.py # `rerank` entrypoint — `python -m rerank.cli`
|
||
│ ├── client.py # RerankClient (Python SDK to call this API)
|
||
│ ├── config.py # Pydantic Settings, RERANK_ env prefix
|
||
│ ├── schemas.py # Pydantic request/response models
|
||
│ ├── exceptions.py # Custom error types
|
||
│ ├── logging.py # Structured logging setup
|
||
│ ├── types.py # Backend literal types
|
||
│ ├── api/
|
||
│ │ ├── app.py # FastAPI app factory
|
||
│ │ ├── dependencies.py # Auth + rate-limit DI
|
||
│ │ ├── middleware.py # X-Request-ID, rate limit, error handlers
|
||
│ │ └── routes/
|
||
│ │ ├── rerank.py # POST /v1/rerank, /v2/rerank
|
||
│ │ ├── models.py # GET /v1/models, /v1/backends
|
||
│ │ └── health.py # GET /health, /ready
|
||
│ └── backends/
|
||
│ ├── base.py # Abstract BackendBase (rerank, health, list_models)
|
||
│ ├── vllm_backend.py # vLLM via OpenAI SDK (--task score)
|
||
│ ├── llamacpp_backend.py # llama.cpp --reranking endpoint
|
||
│ └── registry.py # Backend registry + selector
|
||
└── tests/
|
||
├── conftest.py
|
||
├── test_config.py
|
||
└── test_schemas.py
|
||
```
|
||
|
||
---
|
||
|
||
## Configuration
|
||
|
||
All env vars use the `RERANK_` prefix and are read via Pydantic Settings (`config.py`).
|
||
|
||
**Required (no defaults):**
|
||
|
||
| Var | Example |
|
||
|-----|---------|
|
||
| `RERANK_DEFAULT_BACKEND` | `vllm` or `llamacpp` |
|
||
| `RERANK_ENABLE_VLLM` | `true` / `false` |
|
||
| `RERANK_ENABLE_LLAMACPP` | `true` / `false` |
|
||
| `RERANK_EXTERNAL_URL` | `http://localhost:14200` (used in OpenAPI spec) |
|
||
|
||
**API server:**
|
||
|
||
| Var | Default | Notes |
|
||
|-----|---------|-------|
|
||
| `RERANK_HOST` | `0.0.0.0` | |
|
||
| `RERANK_PORT` | `14200` | Prod 14200 / Dev 54200 |
|
||
| `RERANK_API_TOKENS` | (empty) | Comma-separated; if empty, auth is disabled |
|
||
| `RERANK_RATE_LIMIT_RPS` | `20.0` | |
|
||
| `RERANK_RATE_LIMIT_BURST` | `50` | |
|
||
| `RERANK_MAX_CONCURRENT_RERANKS` | `20` | |
|
||
| `RERANK_REQUEST_TIMEOUT` | `120.0` s | |
|
||
| `RERANK_CONNECT_TIMEOUT` | `10.0` s | |
|
||
| `RERANK_LOG_LEVEL` | `INFO` | |
|
||
| `RERANK_LOG_JSON` | `false` | |
|
||
|
||
**Backend base URLs (app settings — fields on `RerankSettings`):**
|
||
|
||
| Var | Default | Notes |
|
||
|-----|---------|-------|
|
||
| `RERANK_VLLM_BASE_URL` | `http://localhost:54201` | Where the API wrapper reaches vLLM |
|
||
| `RERANK_VLLM_API_KEY` | (empty) | Optional API key for the vLLM server |
|
||
| `RERANK_LLAMACPP_BASE_URL` | `http://localhost:54210` | Where the API wrapper reaches llama.cpp |
|
||
|
||
> Only `*_BASE_URL` / `*_API_KEY` are read by the API app (`config.py`). The model id, port, GPU and context knobs below are **not** `RerankSettings` fields.
|
||
|
||
**Backend service variables (docker-compose only):**
|
||
|
||
These are consumed by `deploy/docker-compose.yml` to launch the vLLM / llama.cpp containers — they configure the backend server, not the API app.
|
||
|
||
| Var | Default | Description |
|
||
|-----|---------|-------------|
|
||
| `RERANK_VLLM_MODEL` | `BAAI/bge-reranker-v2-m3` | HF model id loaded by the vLLM container |
|
||
| `RERANK_VLLM_PORT` | `14201` | Host port mapped to the vLLM container |
|
||
| `RERANK_VLLM_GPU` | `0` | `CUDA_VISIBLE_DEVICES` for the vLLM container |
|
||
| `RERANK_VLLM_GPU_UTIL` | `0.50` | vLLM `--gpu-memory-utilization` |
|
||
| `RERANK_VLLM_MAX_LEN` | `8192` | vLLM `--max-model-len` |
|
||
| `RERANK_LLAMACPP_MODEL` | `bge-reranker-v2-m3-q4_k_m.gguf` | GGUF filename inside `MODELS_DIR` |
|
||
| `RERANK_LLAMACPP_PORT` | `14210` | Host port mapped to the llama.cpp container |
|
||
| `RERANK_LLAMACPP_CTX` | `8192` | llama.cpp context size |
|
||
| `RERANK_LLAMACPP_THREADS` | `4` | llama.cpp threads |
|
||
| `RERANK_LLAMACPP_PARALLEL` | `4` | llama.cpp parallel slots |
|
||
| `MODELS_DIR` | `/cai2_ds_storage/models` | GGUF model directory for llama.cpp |
|
||
| `HF_CACHE_DIR` | `/cai2_ds_storage/hf_cache` | Mounted into the vLLM container |
|
||
| `HF_TOKEN` | (unset) | Forwarded as `HUGGING_FACE_HUB_TOKEN` |
|
||
|
||
Full list with comments in `.env.example`.
|
||
|
||
---
|
||
|
||
## Performance / model
|
||
|
||
- **Model:** `BAAI/bge-reranker-v2-m3` — multilingual cross-encoder, ~8K context window, ~568M params.
|
||
- **Output:** single relevance score per `(query, document)` pair, normalized to `[0..1]`.
|
||
- **Expected latency:** roughly 100–300 ms for ~20 candidates per query on a single mid-range GPU with vLLM batching; scales sublinearly with batch size.
|
||
- **Throughput:** dominated by GPU memory and `RERANK_VLLM_GPU_UTIL` / `RERANK_MAX_CONCURRENT_RERANKS`. vLLM continuous batching helps a lot under concurrent load.
|
||
- **Memory:** vLLM defaults to 50% GPU memory util (`RERANK_VLLM_GPU_UTIL=0.50`), so it can co-host on a GPU shared with other modules. llama.cpp Q4_K_M quantization fits comfortably on CPU.
|
||
|
||
---
|
||
|
||
## Deployment
|
||
|
||
Docker Compose with three profiles in `deploy/docker-compose.yml`:
|
||
|
||
```bash
|
||
cd /home/admin365/didi_mono/ai_platform/modules/rerank/deploy
|
||
|
||
# Production (GPU + vLLM)
|
||
./deploy.sh --profile vllm -d
|
||
|
||
# Lightweight / no-GPU
|
||
./deploy.sh --profile llamacpp -d
|
||
|
||
# API only (external rerank servers running elsewhere)
|
||
./deploy.sh --profile api -d
|
||
|
||
# Logs / stop
|
||
./deploy.sh --profile vllm --logs
|
||
./deploy.sh --profile vllm --down
|
||
```
|
||
|
||
Containers:
|
||
- `didiAI-rerank-api` — FastAPI wrapper, port `RERANK_PORT` (14200/54200)
|
||
- `didiAI-rerank-vllm` — `vllm/vllm-openai:v0.8.5`, GPU device pinned via `RERANK_VLLM_GPU`, internal port 14201
|
||
- `didiAI-rerank-llamacpp` — `ghcr.io/ggml-org/llama.cpp:server-b4769`, internal port 8080 → host 14210/54210
|
||
|
||
All services share the external `didi-network` Docker network so other DIDI modules (didi-brain, embeddings) can reach them by container name.
|
||
|
||
Health gates: API has a 30-s `/health` healthcheck; vLLM has a 5-min start period to allow model load on first boot.
|
||
|
||
---
|
||
|
||
## Related
|
||
|
||
- **Embeddings module** (`ai_platform/modules/embeddings`) — separate service, dim 1024, used for the initial pgvector kNN stage that feeds this reranker.
|
||
- **didi-brain** (`ai_platform/modules/didi_brain/brain_api`) — main consumer; `services/gather.py` calls rerank after kNN; `shared/reranker_client.py` is the helper.
|
||
- **GPU host:** `10.11.10.17` / `10.11.10.12` (see project network architecture doc).
|