Perplexity open-sources pplx-embed-v2-late, MIT-licensed ColBERT embeddings that search PDF pages without OCR
Released Oct 7 in 0.6B and 9B sizes on Hugging Face, the Qwen3.5-based models keep one 128-dim vector per token for text, images, and rendered document pages, scoring 62.3% and 65.2% nDCG@10 on ViDoRe v3 (image). They share one embedding space, so you can index with the 9B and query cheaply with the 0.6B, and they load through sentence-transformers>=6.0.0's MultiVectorEncoder with no custom code.