| .. | ||
| deploy | ||
| src/rerank | ||
| tests | ||
| .env.example | ||
| API.md | ||
| INDEX.md | ||
| pyproject.toml | ||
| README.md | ||
Rerank Module
Cohere/Jina-compatible reranking API with support for vLLM and llama.cpp backends.
Features
- Cohere/Jina-compatible API: Drop-in replacement for Cohere's
/v1/rerankendpoint - Multiple backends: Support for vLLM and llama.cpp
- High performance: Built on FastAPI with async support
- Production-ready: Rate limiting, authentication, health checks
Prerequisites
Required
- Linux - Ubuntu 22.04+ or similar
- Python 3.10+ - Managed via
uv - uv - Fast Python package manager
Optional (for backends)
- vLLM - Requires NVIDIA GPU with CUDA 12.x
- llama.cpp - Can run on CPU or GPU
Quick Start
1. Install dependencies
cd modules/rerank
uv sync --all-extras
2. Configure environment
cp .env.example .env
# Edit .env with your settings
3. Start backend server
vLLM (GPU):
python -m vllm.entrypoints.openai.api_server \
--model BAAI/bge-reranker-v2-m3 \
--host 0.0.0.0 --port 54201 \
--task score
llama.cpp (CPU):
llama-server \
--model /models/bge-reranker-v2-m3-q4_k_m.gguf \
--host 0.0.0.0 --port 54210 \
--reranking
4. Start API server
# Set required environment variables
export RERANK_DEFAULT_BACKEND=vllm
export RERANK_ENABLE_VLLM=true
export RERANK_ENABLE_LLAMACPP=false
export RERANK_EXTERNAL_URL=http://localhost:54200
# Run the server
uv run python -m rerank.cli --port 54200
5. Test the API
curl -X POST http://localhost:54200/v1/rerank \
-H "Content-Type: application/json" \
-d '{
"model": "BAAI/bge-reranker-v2-m3",
"query": "What is machine learning?",
"documents": [
"Machine learning is a subset of AI",
"Cats are pets",
"Deep learning uses neural networks"
]
}'
Docker Deployment
cd modules/rerank/deploy
# Copy and configure .env
cp ../.env.example .env
# Edit .env with your settings
# Start with vLLM backend
./deploy.sh --profile vllm -d
# Or start with llama.cpp backend
./deploy.sh --profile llamacpp -d
# View logs
./deploy.sh --profile vllm --logs
# Stop
./deploy.sh --profile vllm --down
Python Library Usage
from rerank import RerankClient
# Initialize client (reads config from environment)
client = RerankClient()
# Rerank documents
response = await client.rerank(
query="What is machine learning?",
documents=[
"ML is a subset of AI",
"Cats are pets",
"Deep learning uses neural networks",
],
model="BAAI/bge-reranker-v2-m3",
)
# Access results
for result in response.results:
print(f"Index {result.index}: {result.relevance_score:.4f}")
# Get top N results
response = await client.rerank(
query="programming",
documents=["Python", "Java", "Cooking", "C++"],
model="BAAI/bge-reranker-v2-m3",
top_n=2,
)
# Include documents in response
response = await client.rerank(
query="AI",
documents=["Machine learning", "Deep learning"],
model="BAAI/bge-reranker-v2-m3",
return_documents=True,
)
for result in response.results:
print(f"{result.document}: {result.relevance_score:.4f}")
Configuration
All configuration is via environment variables with the RERANK_ prefix:
| Variable | Required | Default | Description |
|---|---|---|---|
RERANK_DEFAULT_BACKEND |
Yes | - | Default backend: vllm or llamacpp |
RERANK_ENABLE_VLLM |
Yes | - | Enable vLLM backend |
RERANK_ENABLE_LLAMACPP |
Yes | - | Enable llama.cpp backend |
RERANK_EXTERNAL_URL |
Yes | - | External URL for OpenAPI spec |
RERANK_PORT |
No | 54200 | API server port |
RERANK_VLLM_BASE_URL |
No | http://localhost:54201 | vLLM server URL |
RERANK_LLAMACPP_BASE_URL |
No | http://localhost:54210 | llama.cpp server URL |
RERANK_API_TOKENS |
No | - | Comma-separated API tokens |
RERANK_RATE_LIMIT_RPS |
No | 20.0 | Requests per second limit |
RERANK_MAX_CONCURRENT_RERANKS |
No | 20 | Max concurrent requests |
See .env.example for the complete list.
Port Allocation
Following the datacenter port schema (x42xx = Reranking):
| Port | Service | Environment |
|---|---|---|
| 14200 | Rerank API | Production |
| 54200 | Rerank API | Development |
| 14201 | vLLM Rerank Server | Production |
| 54201 | vLLM Rerank Server | Development |
| 14210 | llama.cpp Rerank Server | Production |
| 54210 | llama.cpp Rerank Server | Development |
Development
# Install dev dependencies
uv sync --all-extras
# Run tests
uv run pytest
# Run tests with coverage
uv run pytest --cov=src/rerank --cov-report=term-missing
# Lint and format
uv run ruff check .
uv run ruff format .
# Type check
uv run mypy src/
API Reference
See API.md for the complete API documentation.