MemReranker is a reasoning-aware reranking model family (0.6B / 4B) purpose-built for agent memory retrieval. It is fine-tuned from Qwen3-Reranker-4B through multi-stage LLM knowledge distillation.
In agent memory systems, the reranking model serves as the critical bridge connecting user queries with long-term memory. Most systems adopt the “retrieve-then-rerank” two-stage paradigm, but generic reranking models rely on semantic similarity matching and lack genuine reasoning capabilities. This leads to recalled results that are semantically relevant yet do not contain the key information needed to answer the question.
MemReranker addresses three specific problems in memory scenarios:
- Score Miscalibration — Relevance scores from generic models are poorly calibrated, making threshold-based filtering difficult.
- Complex Query Degradation — Ranking degrades when facing temporal constraints, causal reasoning, and other complex queries.
- Context Disambiguation — The model cannot leverage dialogue context for semantic disambiguation.
https://arxiv.org/html/2605.06132v2
As illustrated in Figure 2, MemReranker utilizes Qwen3-Reranker as its foundation. We employ Binary Cross-Entropy (BCE) loss for the training process—a design choice informed by the empirical evidence from BiXSE [18]. Their findings demonstrate that at this specific parameter scale, BCE-trained models consistently yield superior performance compared to those trained with InfoNCE loss, effectively establishing BCE as the optimal maximum-likelihood estimator for sigmoid-activated relevance scoring.
3.1.2 Instruction-Aware Design
Inspired by the instruction-following capabilities of Qwen3-Reranker and the task-aware approach of Jina Reranker v3 [23], MemReranker supports three categories of retrieval instructions:
Intent-Focusing Instructions.
These extract the core retrieval intent from history-heavy long queries. For example, when dialogue history discusses mobile phone preferences and the current query is “I want to look at Apple,” the instruction guides the model to interpret this as a smartphone query rather than a fruit query.
Entity/Keyword Augmentation Instructions.
These bridge the vocabulary gap between colloquial user queries and professional document terminology, mapping informal descriptions to domain-specific terms.
Aspect-Constraint Instructions.
When a query contains multiple needs but a document satisfies only one aspect, the instruction guides the model to focus on the relevant portion, enabling partial-match scoring.

