# ML Projects A monorepo for machine learning services that **expose ML models via REST APIs**. This is the core purpose of the repository - each module provides production-ready API endpoints for inference, allowing applications to consume ML capabilities over HTTP. Each module (LLM inference, RAG, etc.) is an independently deployable service with its own API, Docker configuration, and documentation. ## Quickstart (CPU-only host) Bootstrap the **web + dashboard** stack in one command. This is the recommended path for a host that doesn't have a GPU and points at an LLM endpoint running elsewhere. ```bash git clone cd ml-projects ./bootstrap.sh ``` The interactive script prompts for the GPU host URL, paid search provider keys, and generates a fresh DB password. See [`DEPLOYMENT.md`](./DEPLOYMENT.md) for architecture, manual steps, configuration reference, and troubleshooting. ### What gets deployed by `bootstrap.sh` | Service | Port | Role | |---------------|--------|-------------------------------------------------| | Web API | 51100 | Free (SearXNG) + premium (paid) search pipeline | | Dashboard | 51300 | Monitoring, runtime config, archive, cost | | Dashboard DB | 15432 | PostgreSQL — history + config + audit + archive | GPU-dependent services (`llm-inference`, `embeddings`, `rerank`, `audio`, `video-analysis`) remain on a separate GPU host and are called over HTTP. Their URLs are captured in `modules/*/deploy/.env`. ## Repository Structure ``` / ├── README.md # This file ├── CLAUDE.md # Guidelines for Claude AI assistant ├── .pre-commit-config.yaml # Pre-commit hooks ├── ruff.toml # Shared linting/formatting config ├── modules/ │ ├── / │ │ ├── pyproject.toml # Module dependencies and metadata │ │ ├── README.md # Module docs (incl. prerequisites) │ │ ├── .env.example # Required environment variables │ │ ├── src// # Source code (importable package) │ │ │ ├── __init__.py │ │ │ └── ... │ │ ├── tests/ # Tests │ │ │ ├── __init__.py │ │ │ └── ... │ │ └── deploy/ # REQUIRED: Deployment directory │ │ ├── deploy.sh # Main deployment script │ │ ├── docker-compose.yml │ │ ├── Dockerfile │ │ └── nginx.conf # Optional: reverse proxy config │ └── ... └── shared/ # Optional shared utilities └── ... ``` ## Getting Started ### Prerequisites **Required (all modules):** - **Linux** - Ubuntu 22.04+ or similar distribution (no Windows/macOS support) - **Docker Engine 24.0+** - With Docker Compose V2 - **Python 3.10+** - **[uv](https://docs.astral.sh/uv/)** - Fast Python package manager - **Git** - With pre-commit hooks support **Optional (for GPU workloads):** - **NVIDIA Driver 535+** - For CUDA 12.x support - **NVIDIA Container Toolkit** - For GPU access in Docker containers **Note:** Individual modules may have additional prerequisites (specific hardware, external services, etc.). Check each module's README for module-specific requirements. ### Initial Setup 1. Clone the repository: ```bash git clone cd ml-projects ``` 2. Install pre-commit hooks: ```bash uv tool install pre-commit pre-commit install ``` ### Working with a Module Each module is independent. Navigate to the module directory to work with it: ```bash cd modules/ # Create virtual environment and install dependencies uv sync # Run tests uv run pytest # Run linting uv run ruff check . uv run ruff format --check . # Start services (if docker-compose.yml exists) docker compose up -d ``` ## Module Guidelines ### Required Structure Every module MUST have: | File/Directory | Purpose | |----------------|---------| | `pyproject.toml` | Package metadata, dependencies, and build config | | `README.md` | Module documentation (purpose, prerequisites, usage) | | `API.md` | API documentation (required for modules with HTTP APIs) | | `.env.example` | Required environment variables with documentation | | `src//` | Source code as importable package | | `src//__init__.py` | Package init with version | | `tests/` | Test directory with pytest tests | | `deploy/` | Deployment directory (see below) | ### Deployment Directory (Required) Every module MUST have a `deploy/` directory containing: | File | Purpose | |------|---------| | `deploy.sh` | Main deployment script with `--help`, validation, up/down/logs | | `docker-compose.yml` | Docker Compose with profiles, health checks, pinned versions | | `Dockerfile` | Container image for the module | **Optional deployment files:** - `nginx.conf` - Reverse proxy configuration - `*.conf` - Other service configurations **deploy.sh requirements:** - Must validate all required environment variables (fail-fast) - Must support `--help` with documentation - Must handle `up`, `down`, `logs` actions - Must source `.env` file if present - Must NOT use fallback defaults ### pyproject.toml Requirements Every module's `pyproject.toml` must include: ```toml [project] name = "" version = "0.1.0" description = "Brief description" requires-python = ">=3.10" dependencies = [ # production dependencies ] [project.optional-dependencies] dev = [ "pytest>=8.0", "pytest-cov>=4.0", "ruff>=0.8", ] [build-system] requires = ["hatchling"] build-backend = "hatchling.build" [tool.hatch.build.targets.wheel] packages = ["src/"] # Use root ruff.toml - only override if necessary [tool.ruff] extend = "../../ruff.toml" ``` ### Docker/Docker Compose Standards - Use `docker-compose.yml` (not `docker-compose.yaml`) - Pin image versions (avoid `latest` tag) - Use environment files (`.env`) for configuration - Document all exposed ports in module README - Include health checks for services - **No fallback defaults** - use `${VAR}` not `${VAR:-default}` (fail-fast on missing config) ### Configuration Best Practices **Never use fallback defaults for configuration variables.** If a required variable isn't set, the application should fail immediately with a clear error message. ```yaml # Good - fails if OPENAI_API_KEY not set environment: - OPENAI_API_KEY=${OPENAI_API_KEY} # Bad - silently uses empty string environment: - OPENAI_API_KEY=${OPENAI_API_KEY:-} ``` This applies to: - Docker Compose environment variables - Python `os.getenv()` calls - Pydantic Settings defaults **Why?** Silent misconfiguration leads to hard-to-debug production issues. Fail-fast behavior catches problems at startup. ### API Documentation (Required for HTTP APIs) Modules that expose HTTP endpoints **MUST** have an `API.md` file documenting the API. **Required sections:** | Section | Description | |---------|-------------| | Base URL | Use `{BASE_URL}` variable with common configurations | | Authentication | Auth requirements (Bearer token, API key, or "none") | | Endpoints | All endpoints with method, path, request/response, examples | | Error Responses | HTTP status codes and error format | | Request Headers | Required and optional headers | **Optional sections:** Rate Limiting, SDK Examples, Supported Backends **Reference:** See `modules/llm-inference/API.md` for a complete example. ### Port Convention Port allocation follows datacenter schema (5-digit ports): | Prefix | Environment | Description | |--------|-------------|-------------| | 1xxxx | Production | Production services | | 5xxxx | Development | Development/testing services | | Second Digit | Category | Example | |--------------|----------|---------| | x0xxx | Web / Frontend | 10000, 50000 | | x1xxx | API / Gateway | 11000, 51100 | | x4xxx | LLM / AI | 14001 | **Current Port Allocation:** | Port | Service | Environment | |------|---------|-------------| | 11000 | Catalog API | Production | | 14001 | Qwen3.5-35B-A3B | Production | | 14011 | LLM API Gateway | Production | | 51100 | Web API | Development | | 54300 | Audio API | Development | | 54500 | BusterX | Development | | 54600 | Video API | Development | **Guidelines:** - Use 1xxxx for production services - Use 5xxxx for development/testing services - Document your port in ENDPOINTS.md and docker-compose.yml ## Code Quality Standards ### Linting and Formatting This repo uses **[Ruff](https://docs.astral.sh/ruff/)** for linting and formatting. Configuration is in the root `ruff.toml`. ```bash # Check for issues uv run ruff check . # Auto-fix issues uv run ruff check --fix . # Format code uv run ruff format . ``` ### Type Hints Type hints are **required** for all public functions and methods: ```python # Good def process_document(text: str, max_length: int = 512) -> list[str]: ... # Bad - missing type hints def process_document(text, max_length=512): ... ``` ### Testing Tests are **required** for all modules. Use pytest: ```bash # Run all tests uv run pytest # Run with coverage uv run pytest --cov=src/ --cov-report=term-missing # Run specific test file uv run pytest tests/test_specific.py ``` **Minimum requirements:** - All public functions must have tests - Aim for >80% code coverage - Include both unit tests and integration tests where applicable ### Pre-commit Hooks Pre-commit hooks run automatically on `git commit`. They enforce: - Ruff linting and formatting - YAML/TOML/JSON validation - No trailing whitespace - No large files (>1MB) - No private keys committed To run manually: ```bash pre-commit run --all-files ``` ## Contributing ### Adding a New Module 1. Create the module directory structure: ```bash mkdir -p modules//src/ mkdir -p modules//tests mkdir -p modules//deploy touch modules//src//__init__.py touch modules//tests/__init__.py ``` 2. Create `pyproject.toml` following the template above 3. Create `README.md` documenting: - What the module does - Installation instructions - Usage examples 4. Create `API.md` if module has HTTP endpoints (see API Documentation section) 5. Add initial tests 6. Submit a merge request ### Merge Request Process 1. Create a feature branch: `git checkout -b feature/` 2. Make changes and ensure all checks pass: ```bash pre-commit run --all-files uv run pytest ``` 3. Push and create a merge request 4. Request review from at least one team member 5. Address review comments 6. Squash and merge when approved ### Code Review Expectations Reviewers will check for: - [ ] Code follows repo conventions - [ ] Type hints present on public interfaces - [ ] Tests cover new functionality - [ ] Documentation updated if needed - [ ] API.md included/updated for HTTP endpoints - [ ] No security issues introduced - [ ] Changes are focused (no unrelated modifications) ## Module Index | Module | Port | Description | Status | |--------|------|-------------|--------| | [catalog-api](modules/catalog-api/) | 11000 | Service catalog & discovery gateway | Active | | [llm-inference](modules/llm-inference/) | 14011 | Unified LLM inference with multiple backends | Active | | [audio](modules/audio/) | 54300 | Speech-to-text (Whisper) | Active | | [video-analysis](modules/video-analysis/) | 54600 | Deepfake detection & semantic video analysis | Active | | [web](modules/web/) | 51100 | Web scraping & fact-checking evidence gathering | Active | ## Port Allocation Summary Port allocation follows datacenter schema (5-digit ports): | Prefix | Environment | |--------|-------------| | 1xxxx | Production | | 5xxxx | Development | **Current Services:** | Port | Service | Module | Environment | |------|---------|--------|-------------| | 11000 | Catalog API | catalog-api | Production | | 14001 | vLLM Qwen3.5-35B-A3B | llm-inference | Production | | 14011 | LLM API Gateway | llm-inference | Production | | 51100 | Web API | web | Development | | 54300 | Audio API (Whisper) | audio | Development | | 54500 | BusterX vLLM | video-analysis | Development | | 54600 | Video Analysis API | video-analysis | Development | For detailed endpoint documentation, see [ENDPOINTS.md](ENDPOINTS.md).