didi-lot1-ai/ai_platform/modules/rerank/INDEX.md

278 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Rerank Service — INDEX
Cross-encoder reranking API for DIDI semantic search precision. Serves the `BAAI/bge-reranker-v2-m3` model behind a Cohere/Jina-compatible HTTP API. Used after kNN retrieval to score `(query, candidate)` pairs and reorder evidence by true relevance, instead of relying solely on pgvector cosine similarity from embeddings.
- **Stack:** Python 3.10+, FastAPI, Uvicorn, Pydantic v2, httpx, vLLM (or llama.cpp) backend
- **URLs:**
- Dev API: `http://10.11.10.12:54200`
- Prod API: `http://10.11.10.12:14200`
- vLLM internal: `:54201` (Dev) / `:14201` (Prod)
- llama.cpp internal: `:54210` (Dev) / `:14210` (Prod)
- **Containers:** `didiAI-rerank-api` (FastAPI wrapper) + `didiAI-rerank-vllm` (or `didiAI-rerank-llamacpp`)
- **Model:** `BAAI/bge-reranker-v2-m3` (multilingual cross-encoder, ~8K context window)
- **Port schema:** `x42xx` family (Reranking)
---
## Ce face
Re-scores candidate documents against a query using a cross-encoder model — directly attending over both texts together — for higher precision than embedding cosine similarity alone. The cross-encoder gives a true relevance score in `[0..1]`.
**Pipeline role inside DIDI:**
1. didi-brain `/v1/gather` runs an initial kNN retrieval over pgvector embeddings (dim 1024, separate `embeddings` module).
2. The top-K candidates (typically 50200) are passed to this rerank service.
3. The reranker scores each `(query, candidate)` pair and returns a relevance-sorted list.
4. didi-brain takes the top-N (e.g. top 10) reranked items as final evidence.
This two-stage retrieval (kNN → cross-encoder rerank) is the standard high-quality semantic search pattern: embeddings handle scale, the cross-encoder handles precision.
---
## API endpoints
Cohere/Jina-compatible. Auth is optional (Bearer token) — enabled if `RERANK_API_TOKENS` is set. Health endpoints stay public.
### `POST /v1/rerank` (alias `POST /v2/rerank`)
Rerank a list of documents against a query.
**Request body:**
```json
{
"model": "BAAI/bge-reranker-v2-m3",
"query": "What is machine learning?",
"documents": ["Machine learning is...", "Cats are pets", "Deep learning..."],
"top_n": 3,
"return_documents": false,
"backend": null
}
```
| Field | Type | Required | Description |
|-------|------|----------|-------------|
| `model` | string | yes | Model identifier (e.g. `BAAI/bge-reranker-v2-m3`) |
| `query` | string | yes | Search query |
| `documents` | string[] | yes | Documents to rerank (11000) |
| `top_n` | int | no | Return only top N results (default: all) |
| `return_documents` | bool | no | Echo document text in response |
| `backend` | string | no | Override default: `"vllm"` or `"llamacpp"` |
**Response:**
```json
{
"id": "rerank-abc123def456",
"model": "BAAI/bge-reranker-v2-m3",
"results": [
{"index": 2, "relevance_score": 0.9523, "document": null},
{"index": 0, "relevance_score": 0.8876, "document": null}
],
"usage": {"total_tokens": 150},
"backend": "vllm"
}
```
### `GET /v1/models`
List loaded models. Optional `?backend=vllm|llamacpp` filter.
### `GET /v1/backends`
List enabled backends, e.g. `{"backends": ["vllm", "llamacpp"]}`.
### `GET /health`
Detailed per-backend health. Status is `healthy` / `degraded` / `unhealthy`.
### `GET /ready`
Simple K8s-style readiness probe — returns `{"ready": true}`.
### Error codes
`400` bad request · `401` auth failed · `429` rate-limited (with `Retry-After` header) · `503` backend unavailable · `504` backend timeout. All errors return `{"detail": "..."}`. Every response carries an `X-Request-ID` header for tracing.
Full reference in `API.md`.
---
## Backends
The API is a thin FastAPI router in front of one of two cross-encoder servers. Both backends speak OpenAI-compatible HTTP, so the wrapper unifies them.
| Backend | When | Pros | Cons |
|---------|------|------|------|
| **vLLM** | Production, GPU host | High throughput, batched scoring, lowest latency under load | Requires NVIDIA GPU + CUDA 12.x |
| **llama.cpp** | Dev / CPU fallback | Runs on CPU or modest GPU, GGUF quantized models, low memory | Lower throughput |
The wrapper routes requests via a backend registry. Set `RERANK_DEFAULT_BACKEND` to choose, or override per-request with the `backend` field in the body. Both can be enabled simultaneously (`RERANK_ENABLE_VLLM=true`, `RERANK_ENABLE_LLAMACPP=true`).
vLLM is launched with `--task score` (vLLM 0.8.x cross-encoder mode). llama.cpp is launched with `--reranking`.
---
## How didi-brain uses it
didi-brain's gather/retrieval flow:
1. `services/gather.py` builds an initial candidate set via pgvector kNN over the `embeddings` module (1024-dim vectors).
2. Calls the reranker via `shared/reranker_client.py` — typically wrapping the `/v1/rerank` endpoint with the brain's HTTPS/auth config.
3. Picks the top-N reranked candidates as final evidence chunks.
4. Passes them to the LLM as grounded context.
The reranker is therefore in the critical path of every brain `/v1/gather` call. Its latency budget is small (a few hundred ms), so vLLM batching matters in production.
---
## Structura fisiere
```
rerank/
├── README.md # Quick start, install, env vars table
├── API.md # Full HTTP API reference + curl/Python examples
├── INDEX.md # This file
├── pyproject.toml # Package metadata, deps (fastapi, httpx, openai SDK)
├── .env.example # All RERANK_* env vars documented
├── deploy/
│ ├── Dockerfile # API wrapper image
│ ├── docker-compose.yml # api / vllm / llamacpp profiles
│ └── deploy.sh # Helper: ./deploy.sh --profile vllm -d
├── src/rerank/
│ ├── __init__.py # Public exports (RerankClient, ...)
│ ├── cli.py # `rerank` entrypoint — `python -m rerank.cli`
│ ├── client.py # RerankClient (Python SDK to call this API)
│ ├── config.py # Pydantic Settings, RERANK_ env prefix
│ ├── schemas.py # Pydantic request/response models
│ ├── exceptions.py # Custom error types
│ ├── logging.py # Structured logging setup
│ ├── types.py # Backend literal types
│ ├── api/
│ │ ├── app.py # FastAPI app factory
│ │ ├── dependencies.py # Auth + rate-limit DI
│ │ ├── middleware.py # X-Request-ID, rate limit, error handlers
│ │ └── routes/
│ │ ├── rerank.py # POST /v1/rerank, /v2/rerank
│ │ ├── models.py # GET /v1/models, /v1/backends
│ │ └── health.py # GET /health, /ready
│ └── backends/
│ ├── base.py # Abstract BackendBase (rerank, health, list_models)
│ ├── vllm_backend.py # vLLM via OpenAI SDK (--task score)
│ ├── llamacpp_backend.py # llama.cpp --reranking endpoint
│ └── registry.py # Backend registry + selector
└── tests/
├── conftest.py
├── test_config.py
└── test_schemas.py
```
---
## Configuration
All env vars use the `RERANK_` prefix and are read via Pydantic Settings (`config.py`).
**Required (no defaults):**
| Var | Example |
|-----|---------|
| `RERANK_DEFAULT_BACKEND` | `vllm` or `llamacpp` |
| `RERANK_ENABLE_VLLM` | `true` / `false` |
| `RERANK_ENABLE_LLAMACPP` | `true` / `false` |
| `RERANK_EXTERNAL_URL` | `http://localhost:14200` (used in OpenAPI spec) |
**API server:**
| Var | Default | Notes |
|-----|---------|-------|
| `RERANK_HOST` | `0.0.0.0` | |
| `RERANK_PORT` | `14200` | Prod 14200 / Dev 54200 |
| `RERANK_API_TOKENS` | (empty) | Comma-separated; if empty, auth is disabled |
| `RERANK_RATE_LIMIT_RPS` | `20.0` | |
| `RERANK_RATE_LIMIT_BURST` | `50` | |
| `RERANK_MAX_CONCURRENT_RERANKS` | `20` | |
| `RERANK_REQUEST_TIMEOUT` | `120.0` s | |
| `RERANK_CONNECT_TIMEOUT` | `10.0` s | |
| `RERANK_LOG_LEVEL` | `INFO` | |
| `RERANK_LOG_JSON` | `false` | |
**Backend base URLs (app settings — fields on `RerankSettings`):**
| Var | Default | Notes |
|-----|---------|-------|
| `RERANK_VLLM_BASE_URL` | `http://localhost:54201` | Where the API wrapper reaches vLLM |
| `RERANK_VLLM_API_KEY` | (empty) | Optional API key for the vLLM server |
| `RERANK_LLAMACPP_BASE_URL` | `http://localhost:54210` | Where the API wrapper reaches llama.cpp |
> Only `*_BASE_URL` / `*_API_KEY` are read by the API app (`config.py`). The model id, port, GPU and context knobs below are **not** `RerankSettings` fields.
**Backend service variables (docker-compose only):**
These are consumed by `deploy/docker-compose.yml` to launch the vLLM / llama.cpp containers — they configure the backend server, not the API app.
| Var | Default | Description |
|-----|---------|-------------|
| `RERANK_VLLM_MODEL` | `BAAI/bge-reranker-v2-m3` | HF model id loaded by the vLLM container |
| `RERANK_VLLM_PORT` | `14201` | Host port mapped to the vLLM container |
| `RERANK_VLLM_GPU` | `0` | `CUDA_VISIBLE_DEVICES` for the vLLM container |
| `RERANK_VLLM_GPU_UTIL` | `0.50` | vLLM `--gpu-memory-utilization` |
| `RERANK_VLLM_MAX_LEN` | `8192` | vLLM `--max-model-len` |
| `RERANK_LLAMACPP_MODEL` | `bge-reranker-v2-m3-q4_k_m.gguf` | GGUF filename inside `MODELS_DIR` |
| `RERANK_LLAMACPP_PORT` | `14210` | Host port mapped to the llama.cpp container |
| `RERANK_LLAMACPP_CTX` | `8192` | llama.cpp context size |
| `RERANK_LLAMACPP_THREADS` | `4` | llama.cpp threads |
| `RERANK_LLAMACPP_PARALLEL` | `4` | llama.cpp parallel slots |
| `MODELS_DIR` | `/cai2_ds_storage/models` | GGUF model directory for llama.cpp |
| `HF_CACHE_DIR` | `/cai2_ds_storage/hf_cache` | Mounted into the vLLM container |
| `HF_TOKEN` | (unset) | Forwarded as `HUGGING_FACE_HUB_TOKEN` |
Full list with comments in `.env.example`.
---
## Performance / model
- **Model:** `BAAI/bge-reranker-v2-m3` — multilingual cross-encoder, ~8K context window, ~568M params.
- **Output:** single relevance score per `(query, document)` pair, normalized to `[0..1]`.
- **Expected latency:** roughly 100300 ms for ~20 candidates per query on a single mid-range GPU with vLLM batching; scales sublinearly with batch size.
- **Throughput:** dominated by GPU memory and `RERANK_VLLM_GPU_UTIL` / `RERANK_MAX_CONCURRENT_RERANKS`. vLLM continuous batching helps a lot under concurrent load.
- **Memory:** vLLM defaults to 50% GPU memory util (`RERANK_VLLM_GPU_UTIL=0.50`), so it can co-host on a GPU shared with other modules. llama.cpp Q4_K_M quantization fits comfortably on CPU.
---
## Deployment
Docker Compose with three profiles in `deploy/docker-compose.yml`:
```bash
cd /home/admin365/didi_mono/ai_platform/modules/rerank/deploy
# Production (GPU + vLLM)
./deploy.sh --profile vllm -d
# Lightweight / no-GPU
./deploy.sh --profile llamacpp -d
# API only (external rerank servers running elsewhere)
./deploy.sh --profile api -d
# Logs / stop
./deploy.sh --profile vllm --logs
./deploy.sh --profile vllm --down
```
Containers:
- `didiAI-rerank-api` — FastAPI wrapper, port `RERANK_PORT` (14200/54200)
- `didiAI-rerank-vllm``vllm/vllm-openai:v0.8.5`, GPU device pinned via `RERANK_VLLM_GPU`, internal port 14201
- `didiAI-rerank-llamacpp``ghcr.io/ggml-org/llama.cpp:server-b4769`, internal port 8080 → host 14210/54210
All services share the external `didi-network` Docker network so other DIDI modules (didi-brain, embeddings) can reach them by container name.
Health gates: API has a 30-s `/health` healthcheck; vLLM has a 5-min start period to allow model load on first boot.
---
## Related
- **Embeddings module** (`ai_platform/modules/embeddings`) — separate service, dim 1024, used for the initial pgvector kNN stage that feeds this reranker.
- **didi-brain** (`ai_platform/modules/didi_brain/brain_api`) — main consumer; `services/gather.py` calls rerank after kNN; `shared/reranker_client.py` is the helper.
- **GPU host:** `10.11.10.17` / `10.11.10.12` (see project network architecture doc).