Multimodal product search · Qdrant HNSW
Find products by what they look like,
not what they're called.
Browse the catalog
Speed vs accuracy, this query
Same query at each search-time ef. Median of 7 runs. Higher ef walks more of the graph: better recall, more time.
Missed by the approximate search
Index tuning
Speed, accuracy and memory across HNSW configs
Each build indexes every product photo; queries are shopper-style text searched cross-modally. Ground truth is a NumPy brute-force scan.
No benchmark yet
Run python benchmark.py (about 10 minutes on the full catalog), then reload.
Recall@k vs latency
Each line is one graph degree m; points along it are search-time ef. Up and left is better.
Memory per build
Estimated resident memory: float32 vectors + graph, or int8 copy + graph with originals on disk.
Build time
Bulk upload, then one HNSW build per segment.
All runs
Click a column to sort.
Architecture
How a query becomes twelve products
Encode
Text or photo goes through Marqo-FashionSigLIP (ViT-B/16, ONNX on CPU) to a 768-d unit vector. Text and images share one space, so words can find pictures.
Walk the graph
Qdrant's HNSW index enters at a sparse top layer and greedily descends, keeping the best ef candidates. It touches a few thousand vectors, not all of them.
Score cheaply
Candidates are scored against an int8 copy of the vectors (4× smaller, in RAM), then the finalists are rescored with the float32 originals.
Check the work
Every search is also answered exactly with NumPy brute force. The UI shows recall against that, and which products the approximation dropped.
The knobs
- m
- Links per node in the graph. Higher m means better recall at a given ef, but more memory (≈ 8·m bytes per vector) and slower builds.
- ef_construct
- Candidate list size while building. Higher gives a better-connected graph, with build-time cost only.
- ef
- Candidate list size while searching. The per-query dial between latency and recall; try the slider on the Search tab.
- int8
- Scalar quantization: each float becomes one byte (99th-percentile clipping). It cuts vector RAM by 4×; rescoring recovers most of the lost precision.
- filters
- Keyword payload indexes on gender, category, type, colour, usage and season. Qdrant applies them inside the graph walk, not after it.