didi-lot1-ai/ai_platform/modules/embeddings/INDEX.md

223 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Embeddings Module — INDEX
Embedding generation API for DIDI knowledge graph + semantic search. Serves `BAAI/bge-m3` (1024-dim, multilingual, 8K context) via vLLM (or llama.cpp) behind an OpenAI-compatible FastAPI wrapper, used by `didi-brain` for atom storage in pgvector and for `/v1/gather` query embedding.
- **Stack**: Python 3.10+, FastAPI, uvicorn, httpx, Pydantic v2; backends via `vllm/vllm-openai:v0.8.5` or `ghcr.io/ggml-org/llama.cpp:server`
- **URLs**:
- API wrapper: `http://10.11.10.12:14100` (Prod) / `http://10.11.10.12:54100` (Dev)
- vLLM internal: `:14101` (Prod) / `:54101` (Dev)
- llama.cpp internal: `:14110` (Prod) / `:54110` (Dev)
- **Container layout (2-tier)**:
- `didiAI-embeddings-api` (FastAPI wrapper + auth + rate limiting)
- `didiAI-embeddings-vllm` (GPU embedding server) **OR** `didiAI-embeddings-llamacpp` (CPU/lightweight server)
- **Default model**: `BAAI/bge-m3` — 1024 dim, multilingual, 8K context
---
## Ce face
Generate text embeddings for downstream DIDI consumers:
- **didi-brain semantic atom storage** — text claims/atoms get embedded once at ingestion, stored as pgvector columns.
- **didi-brain `/v1/gather` query embedding** — incoming claim text gets embedded, used as the kNN probe vector.
- **Cross-encoder reranking** — separate `rerank` module (port 14200) handles re-scoring; this module only does dense embeddings.
- **Catalog / retrieval consumers** — anything that needs a vector representation hits this single endpoint.
The module is OpenAI-compatible: drop-in for `openai` SDK code that uses `client.embeddings.create()`.
---
## API endpoints
All routes are mounted by `src/embeddings/api/app.py` via `create_app()`. Health is unauthenticated and excluded from rate-limiting; everything else passes through `verify_bearer_token` if `EMB_API_TOKENS` is set.
| Method | Path | Description |
|--------|------|-------------|
| `POST` | `/v1/embeddings` | OpenAI-compatible. Body: `{input, model, encoding_format?, dimensions?, backend?}`. Returns `{object, data[], model, usage, backend}`. |
| `GET` | `/v1/models` | List available embedding models. Optional `?backend=vllm\|llamacpp` filter. |
| `GET` | `/v1/backends` | List enabled backends as `{backends: ["vllm", ...]}`. |
| `GET` | `/health` | Detailed health: per-backend status, overall `healthy/degraded/unhealthy`. |
| `GET` | `/ready` | Simple readiness probe (`{ready: true}`) — for k8s. |
Source files:
- `/home/admin365/didi_mono/ai_platform/modules/embeddings/src/embeddings/api/routes/embeddings.py``POST /v1/embeddings`
- `/home/admin365/didi_mono/ai_platform/modules/embeddings/src/embeddings/api/routes/models.py``/v1/models`, `/v1/backends`
- `/home/admin365/didi_mono/ai_platform/modules/embeddings/src/embeddings/api/routes/health.py``/health`, `/ready`
Errors: 400 (invalid params / backend disabled), 401 (auth), 429 (rate limit, sets `Retry-After`), 503 (backend connection), 504 (backend timeout), 500 (other backend failures).
---
## Backends
Two interchangeable engines, selected per-request via the `backend` field or globally via `EMB_DEFAULT_BACKEND`:
- **vLLM (production / GPU)** — `vllm/vllm-openai:v0.8.5` launched with `--task embed`, high-throughput batching, NVIDIA CUDA 12.x required. Listens internally on `14101` (Prod) / `54101` (Dev).
- **llama.cpp (dev / lightweight)** — `ghcr.io/ggml-org/llama.cpp:server-b4769` launched with `--embedding`, runs on CPU or GPU with GGUF-quantized models. Listens on `14110` / `54110`.
- **API wrapper** — own FastAPI app on `14100` (Prod) / `54100` (Dev) that proxies to whichever backend(s) are enabled.
Both backends speak OpenAI-compatible HTTP, so the wrapper just routes via `httpx`/`openai` SDK. Backend abstraction lives in `src/embeddings/backends/{base,vllm_backend,llamacpp_backend,registry}.py`.
---
## Structura fisiere
```
embeddings/
├── INDEX.md (this file)
├── README.md quick-start + config table
├── API.md full HTTP API reference + SDK examples
├── pyproject.toml package def, optional extras: vllm, llamacpp, all, dev
├── .env.example all EMB_* env vars documented
├── deploy/
│ ├── Dockerfile multi-stage, python 3.11-slim + uv
│ ├── docker-compose.yml 3 services + 3 profiles (api/vllm/llamacpp)
│ └── deploy.sh wrapper around `docker compose --profile`
├── src/embeddings/
│ ├── __init__.py re-exports EmbeddingClient
│ ├── cli.py `python -m embeddings.cli` entry point
│ ├── client.py EmbeddingClient (Python library API)
│ ├── config.py EmbeddingSettings (pydantic-settings, EMB_ prefix)
│ ├── schemas.py EmbeddingRequest/Response/Data, base64 encoder
│ ├── types.py BackendType enum, ModelInfo
│ ├── exceptions.py BackendNotAvailable/Enabled, RateLimit, Timeout, etc.
│ ├── logging.py structured logging (text or JSON)
│ ├── api/
│ │ ├── app.py create_app() factory + lifespan
│ │ ├── dependencies.py FastAPI deps (auth, concurrency limiter, client)
│ │ ├── middleware.py RateLimitMiddleware, RequestIdMiddleware
│ │ └── routes/
│ │ ├── embeddings.py POST /v1/embeddings
│ │ ├── models.py /v1/models, /v1/backends
│ │ └── health.py /health, /ready
│ └── backends/
│ ├── base.py abstract Backend interface
│ ├── vllm_backend.py vLLM HTTP client
│ ├── llamacpp_backend.py llama.cpp HTTP client
│ └── registry.py BackendRegistry (selects + caches backends)
└── tests/
├── conftest.py
├── test_config.py
├── test_schemas.py
└── test_types.py
```
---
## How didi-brain uses it
`didi-brain` consumes this service through `shared/embedding_client.py` (a thin wrapper around the `openai` SDK pointed at this API):
- **Ingestion path** — when atoms/claims are written, brain calls `EmbeddingClient.embed(text)` to get a 1024-dim vector and stores it in pgvector alongside the row. One round-trip per batch.
- **Query path** — for `/v1/gather`, the incoming claim text is embedded the same way, then the resulting vector is used as the probe in a pgvector `<=> ` (cosine) kNN search to retrieve candidate atoms.
- **Reranking** — top-k candidates from the dense search are forwarded to the separate `rerank` module (cross-encoder, port 14200) for fine-grained scoring. That module does not call this one — they're parallel concerns.
- **Catalog API** — also retrieves via the same embedding pipeline (text → vector → kNN), reusing this single endpoint.
Because the wrapper is OpenAI-compatible, `EmbeddingClient` can be a vanilla `openai.OpenAI(base_url=..., api_key=...)` instance — no DIDI-specific client code needed in brain.
---
## Configuration
All env vars use the `EMB_` prefix. Required vars have no defaults — the app refuses to start without them.
### Required
| Variable | Description |
|----------|-------------|
| `EMB_DEFAULT_BACKEND` | `vllm` or `llamacpp` |
| `EMB_ENABLE_VLLM` | `true` / `false` |
| `EMB_ENABLE_LLAMACPP` | `true` / `false` |
| `EMB_EXTERNAL_URL` | Public URL exposed in the OpenAPI spec |
### Common optional
| Variable | Default | Description |
|----------|---------|-------------|
| `EMB_PORT` | `14100` (Prod) / `54100` (Dev) | API server port |
| `EMB_HOST` | `0.0.0.0` | API bind address |
| `EMB_API_TOKENS` | unset | Comma-separated bearer tokens; auth disabled if unset |
| `EMB_VLLM_BASE_URL` | `http://localhost:54101` | Where the API wrapper reaches vLLM |
| `EMB_VLLM_API_KEY` | unset | Optional API key for the vLLM server |
| `EMB_LLAMACPP_BASE_URL` | `http://localhost:54110` | Where the API wrapper reaches llama.cpp |
| `EMB_REQUEST_TIMEOUT` | `120.0` | Per-request backend timeout (s) |
| `EMB_CONNECT_TIMEOUT` | `10.0` | TCP connect timeout (s) |
| `EMB_RATE_LIMIT_RPS` | `20.0` | Token-bucket rate (req/s) |
| `EMB_RATE_LIMIT_BURST` | `40` | Token-bucket burst capacity |
| `EMB_MAX_CONCURRENT_REQUESTS` | `20` | In-flight cap |
| `EMB_LOG_LEVEL` | `INFO` | DEBUG / INFO / WARNING / ERROR |
| `EMB_LOG_JSON` | `false` | Emit JSON-formatted log lines |
> The API app reads only the variables above (fields on `EmbeddingSettings` in `config.py`). `EMB_VLLM_BASE_URL` / `EMB_LLAMACPP_BASE_URL` just tell the wrapper where to reach the backend servers.
### Backend service variables (docker-compose only)
These are **not** fields on `EmbeddingSettings` — they are consumed by `deploy/docker-compose.yml` to launch the vLLM / llama.cpp containers (model id, GPU pinning, memory, context). They configure the backend server, not the API app.
| Variable | Default | Description |
|----------|---------|-------------|
| `EMB_VLLM_MODEL` | `BAAI/bge-m3` | HF model id loaded by the vLLM container |
| `EMB_VLLM_GPU` | `0` | `CUDA_VISIBLE_DEVICES` for the vLLM container |
| `EMB_VLLM_GPU_UTIL` | `0.50` | vLLM `--gpu-memory-utilization` |
| `EMB_VLLM_MAX_LEN` | `8192` | vLLM `--max-model-len` |
| `EMB_LLAMACPP_MODEL` | `bge-m3-q4_k_m.gguf` | GGUF filename inside `MODELS_DIR` |
| `EMB_LLAMACPP_CTX` | `8192` | llama.cpp context size |
| `EMB_LLAMACPP_THREADS` | `4` | llama.cpp threads |
| `EMB_LLAMACPP_PARALLEL` | `4` | llama.cpp parallel slots |
| `HF_CACHE_DIR` | `/cai2_ds_storage/hf_cache` | Mounted into the vLLM container |
| `HF_TOKEN` | unset | Forwarded as `HUGGING_FACE_HUB_TOKEN` |
| `MODELS_DIR` | `/cai2_ds_storage/models` | GGUF model directory for llama.cpp |
See `/home/admin365/didi_mono/ai_platform/modules/embeddings/.env.example` for the canonical, fully commented list.
---
## Deployment
GPU host required for the production profile (e.g. `10.11.10.17`). All three services live on the shared `didi-network` Docker network.
```bash
cd /home/admin365/didi_mono/ai_platform/modules/embeddings/deploy
# Configure
cp ../.env.example .env
$EDITOR .env
# Production: API + vLLM (GPU)
./deploy.sh --profile vllm -d
# Lightweight: API + llama.cpp
./deploy.sh --profile llamacpp -d
# API only (use external embedding servers via EMB_*_BASE_URL)
./deploy.sh --profile api -d
# Logs / shutdown
./deploy.sh --profile vllm --logs
./deploy.sh --profile vllm --down
```
Compose profiles:
- `api` — only the FastAPI wrapper
- `vllm` — wrapper + vLLM (GPU)
- `llamacpp` — wrapper + llama.cpp (CPU/GPU)
---
## Performance / model
- **Model**: `BAAI/bge-m3` — 1024 dim, multilingual (100+ languages), 8K token context, supports dense + sparse + multi-vector (this module uses dense only).
- **Throughput**: ~100200 embeddings/s on H100 with vLLM batching (depends on input length).
- **Latency P95**: ~50100 ms per request (single text, warm GPU). Larger batches amortize well — keep request batches at 1664 inputs for best throughput.
- **Concurrency knobs**: `EMB_MAX_CONCURRENT_REQUESTS` (in-flight at the wrapper) and `EMB_RATE_LIMIT_RPS` (token-bucket) bound the load reaching the GPU.
- **GPU memory**: bge-m3 fits comfortably in <4 GB; `EMB_VLLM_GPU_UTIL=0.50` is intentionally low to allow GPU sharing with reranker / other workloads.
---
## Related
- **`didi-brain`** consumes via `shared/embedding_client.py` for both ingestion (text pgvector storage) and `/v1/gather` query embedding.
- **`rerank` module** (separate, port 14200) handles cross-encoder scoring on top of dense kNN candidates from this module they're complementary, not chained inside this service.
- **Catalog API** uses the same endpoint for retrieval-side embeddings.
- **pgvector** in the brain Postgres stores the resulting 1024-dim vectors (cosine distance index).
- **Local Python use**: `from embeddings import EmbeddingClient` (works without the HTTP wrapper if you want in-process inference and have the `vllm`/`llamacpp` extras installed).