183 lines
4.1 KiB
Markdown
183 lines
4.1 KiB
Markdown
# Embeddings Module
|
|
|
|
OpenAI-compatible embeddings API with support for vLLM and llama.cpp backends.
|
|
|
|
## Features
|
|
|
|
- **OpenAI-compatible API**: Drop-in replacement for OpenAI's `/v1/embeddings` endpoint
|
|
- **Multiple backends**: Support for vLLM and llama.cpp
|
|
- **High performance**: Built on FastAPI with async support
|
|
- **Production-ready**: Rate limiting, authentication, health checks
|
|
|
|
## Prerequisites
|
|
|
|
### Required
|
|
|
|
- **Linux** - Ubuntu 22.04+ or similar
|
|
- **Python 3.10+** - Managed via `uv`
|
|
- **uv** - Fast Python package manager
|
|
|
|
### Optional (for backends)
|
|
|
|
- **vLLM** - Requires NVIDIA GPU with CUDA 12.x
|
|
- **llama.cpp** - Can run on CPU or GPU
|
|
|
|
## Quick Start
|
|
|
|
### 1. Install dependencies
|
|
|
|
```bash
|
|
cd modules/embeddings
|
|
uv sync --all-extras
|
|
```
|
|
|
|
### 2. Configure environment
|
|
|
|
```bash
|
|
cp .env.example .env
|
|
# Edit .env with your settings
|
|
```
|
|
|
|
### 3. Start backend server
|
|
|
|
**vLLM (GPU):**
|
|
```bash
|
|
python -m vllm.entrypoints.openai.api_server \
|
|
--model BAAI/bge-m3 \
|
|
--host 0.0.0.0 --port 54101 \
|
|
--task embed
|
|
```
|
|
|
|
**llama.cpp (CPU):**
|
|
```bash
|
|
llama-server \
|
|
--model /models/bge-m3-q4_k_m.gguf \
|
|
--host 0.0.0.0 --port 54110 \
|
|
--embedding
|
|
```
|
|
|
|
### 4. Start API server
|
|
|
|
```bash
|
|
# Set required environment variables
|
|
export EMB_DEFAULT_BACKEND=vllm
|
|
export EMB_ENABLE_VLLM=true
|
|
export EMB_ENABLE_LLAMACPP=false
|
|
export EMB_EXTERNAL_URL=http://localhost:54100
|
|
|
|
# Run the server
|
|
uv run python -m embeddings.cli --port 54100
|
|
```
|
|
|
|
### 5. Test the API
|
|
|
|
```bash
|
|
curl -X POST http://localhost:54100/v1/embeddings \
|
|
-H "Content-Type: application/json" \
|
|
-d '{
|
|
"input": "Hello, world!",
|
|
"model": "BAAI/bge-m3"
|
|
}'
|
|
```
|
|
|
|
## Docker Deployment
|
|
|
|
```bash
|
|
cd modules/embeddings/deploy
|
|
|
|
# Copy and configure .env
|
|
cp ../.env.example .env
|
|
# Edit .env with your settings
|
|
|
|
# Start with vLLM backend
|
|
./deploy.sh --profile vllm -d
|
|
|
|
# Or start with llama.cpp backend
|
|
./deploy.sh --profile llamacpp -d
|
|
|
|
# View logs
|
|
./deploy.sh --profile vllm --logs
|
|
|
|
# Stop
|
|
./deploy.sh --profile vllm --down
|
|
```
|
|
|
|
## Python Library Usage
|
|
|
|
```python
|
|
from embeddings import EmbeddingClient
|
|
|
|
# Initialize client (reads config from environment)
|
|
client = EmbeddingClient()
|
|
|
|
# Generate embeddings
|
|
response = await client.embed(
|
|
texts=["Hello, world!", "How are you?"],
|
|
model="BAAI/bge-m3",
|
|
)
|
|
|
|
# Access embeddings
|
|
for item in response.data:
|
|
print(f"Index {item.index}: {len(item.embedding)} dimensions")
|
|
|
|
# List available models
|
|
models = await client.list_models()
|
|
for model in models:
|
|
print(f"{model.id} on {model.backend}")
|
|
```
|
|
|
|
## Configuration
|
|
|
|
All configuration is via environment variables with the `EMB_` prefix:
|
|
|
|
| Variable | Required | Default | Description |
|
|
|----------|----------|---------|-------------|
|
|
| `EMB_DEFAULT_BACKEND` | Yes | - | Default backend: `vllm` or `llamacpp` |
|
|
| `EMB_ENABLE_VLLM` | Yes | - | Enable vLLM backend |
|
|
| `EMB_ENABLE_LLAMACPP` | Yes | - | Enable llama.cpp backend |
|
|
| `EMB_EXTERNAL_URL` | Yes | - | External URL for OpenAPI spec |
|
|
| `EMB_PORT` | No | 54100 | API server port |
|
|
| `EMB_VLLM_BASE_URL` | No | http://localhost:54101 | vLLM server URL |
|
|
| `EMB_LLAMACPP_BASE_URL` | No | http://localhost:54110 | llama.cpp server URL |
|
|
| `EMB_API_TOKENS` | No | - | Comma-separated API tokens |
|
|
| `EMB_RATE_LIMIT_RPS` | No | 20.0 | Requests per second limit |
|
|
| `EMB_MAX_CONCURRENT_REQUESTS` | No | 20 | Max concurrent requests |
|
|
|
|
See `.env.example` for the complete list.
|
|
|
|
## Port Allocation
|
|
|
|
Following the datacenter port schema (x41xx = Embeddings):
|
|
|
|
| Port | Service | Environment |
|
|
|------|---------|-------------|
|
|
| 14100 | Embeddings API | Production |
|
|
| 54100 | Embeddings API | Development |
|
|
| 14101 | vLLM Embed Server | Production |
|
|
| 54101 | vLLM Embed Server | Development |
|
|
| 14110 | llama.cpp Embed Server | Production |
|
|
| 54110 | llama.cpp Embed Server | Development |
|
|
|
|
## Development
|
|
|
|
```bash
|
|
# Install dev dependencies
|
|
uv sync --all-extras
|
|
|
|
# Run tests
|
|
uv run pytest
|
|
|
|
# Run tests with coverage
|
|
uv run pytest --cov=src/embeddings --cov-report=term-missing
|
|
|
|
# Lint and format
|
|
uv run ruff check .
|
|
uv run ruff format .
|
|
|
|
# Type check
|
|
uv run mypy src/
|
|
```
|
|
|
|
## API Reference
|
|
|
|
See [API.md](API.md) for the complete API documentation.
|