How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
02:00 · August 21, 2026 · Hugging Face Blog

Summary
The hybrid search system for Papers with Code combines PostgreSQL full-text search with dense vector retrieval over more than 110,000 arXiv papers. Lexical matching supplies exact titles, identifiers and rare terms, while semantic search, built on pgvector, surfaces conceptual matches even when query phrasing diverges from paper text. Reciprocal rank fusion merges the two ranked lists, giving equal weight to each branch and preserving deterministic identity lookups on top of the fused scores.
Corpus construction runs as an offline Hugging Face Job. A repeatable-read snapshot exports the current paper catalog into versioned JSONL shards stored in a private Bucket. The same pinned Qwen3-Embedding-0.6B revision, configured for 256-dimensional L2-normalized output, encodes every document on an L4 GPU instance. Checksums, model revision, prompt template and dimensionality are recorded alongside each vector so that any later mismatch can be detected before import. Completed shards are imported atomically: a new HNSW index is built beside the active one and only swapped in once coverage and content hashes are verified.
Query-time embedding uses a single-replica Inference Endpoint backed by Text Embeddings Inference. The endpoint applies the model’s query prompt and returns a normalized vector for cosine search. When the endpoint is cold, saturated or returns an invalid result, the application immediately drops the semantic branch and returns lexical results only. This design treats scale-to-zero as a normal operating mode rather than an exception.
Incremental updates run hourly. Changed or missing papers are batched through the same endpoint using the document prompt, with row-level locking and hash checks to avoid stale vectors. Related-paper recommendations reuse the stored document embeddings, requiring no additional model call at request time; citation-graph fallbacks from Semantic Scholar cover any temporary gaps.
Benchmarks on a 5,000-paper pilot showed that 256-dimensional vectors retain 0.9955 Recall@20 against exact search while using roughly one-quarter of the storage required by 1024-dimensional vectors. Median HNSW lookup latency on L4 hardware stayed below 1.5 ms. The architecture therefore separates throughput-oriented batch work from latency-sensitive online inference, keeps an explicit versioning contract across every stage, and supplies a robust lexical fallback whenever the semantic path is unavailable.
Why it matters
Provides actionable production patterns for embedding pipelines, vector search, and reliable inference that Dutch ML teams can directly apply with HF tooling. Addresses latency, VRAM, versioning, and cost concerns relevant to EU practitioners.








