This repository contains a local proof of concept that:
- reads all PDFs from
sources/ - renders every page to an image
- extracts native PDF text with PyMuPDF
- stores page renders as real image vectors
- stores embedded PDF images as separate real image vectors when the PDF exposes them
- embeds native text chunks with the same multimodal model
- writes everything into a local Qdrant database inside this repo
- supports retrieval-only
searchandaskcommands against the joint text/image vector space
Qwen3-VL-Embedding-2Bfor true multimodal embeddings- PyMuPDF for PDF text extraction and rendering
- Qdrant local mode for persistent vector storage
- local repo-scoped
.venv/,.hf/, andmodels/
- Virtual environment:
.venv/ - Hugging Face cache:
.hf/ - Embedding model:
models/Qwen3-VL-Embedding-2B/ - Vector DB:
data/qdrant/ - Page renders and extracted assets:
artifacts/page_renders/ - Run reports:
artifacts/reports/ - Future-agent docs:
.codex/
python -m pip --python .\.venv\Scripts\python.exe install torch torchvision --index-url https://download.pytorch.org/whl/cu128 --extra-index-url https://pypi.org/simple transformers safetensors qwen-vl-utils huggingface-hub accelerate pymupdf qdrant-client pillowNew-Item -ItemType Directory -Force .hf, .hf\hub, models\Qwen3-VL-Embedding-2B | Out-Null
.\.venv\Scripts\hf.exe download Qwen/Qwen3-VL-Embedding-2B --cache-dir .\.hf\hub --local-dir .\models\Qwen3-VL-Embedding-2B --max-workers 4.\.venv\Scripts\python.exe .\main.py ingest --recreate --embed-batch-size 1Optional flags:
--limit-pages 3limits ingest for quick tests--render-dpi 96changes page render resolution--embed-batch-size 2can improve throughput if VRAM allows it--collection my_collectionchanges the local Qdrant collection name
.\.venv\Scripts\python.exe .\main.py search "Welche Designregeln fuer Perception werden vorgestellt?" --limit 6The search operates in one shared embedding space across:
- text chunk records
- full page-image records
- embedded PDF image records
.\.venv\Scripts\python.exe .\main.py ask "Welche Designregeln fuer Perception werden vorgestellt?" --limit 8This retrieves the most relevant multimodal contexts for a natural-language question and returns the matching chunks/pages directly.
.\.venv\Scripts\python.exe .\main.py statsThe ingest report tracks:
- pages per second
- vector records per second
- embedding items per second
- embedding input tokens per second
- text-vs-image embedding item counts
Important note:
- page and extracted-image records are embedded as images, not by converting them to text first