Pith. sign in

REVIEW 3 major objections 5 minor 61 references

SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read SPEAR's central claim: training query-rewrite selection on click and retrieval outcomes lifts click recall@10 by +99.5% and rewrite similarity@10 by +18.2% over the production baseline.

desk verdict Solid industrial paper with a believable failure mode and coherent fix; the intent-faithfulness headline rests on an under-specified metric and needs a review push. read the letter →

arxiv 2608.01738 v1 pith:PR2OVMAL submitted 2026-08-03 cs.IR

classification cs.IR
keywords queryreformulationpersonalizedretrievalembedding-basedrewriteselectionpath-basedgeneric-worddominanceend-to-endoptimizatione-commercecommunitysearch
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that the long-standing split between query rewrite quality and retrieval effectiveness in e-commerce search can be closed by making rewrite selection part of the retrieval objective. It names the failure mode that blocks naive end-to-end transplanting of path-based recommendation architectures: generic-word dominance, where rewrites such as "phone case" or "women's dress" win by frequency and co-occurrence while drifting from the user's stated intent. SPEAR counteracts this with three structural components: a dual-embedding backbone with gradient isolation, a multiplicative gating aggregator, and a Dynamic Rewrite Selector that emits per-request weights and calibration. On 100K held-out industrial search sessions the paper reports click recall@10 improving by +99.5% and rewrite semantic similarity@10 by +18.2% over the production PDN baseline, and a deployed A/B test shows simultaneous gains in query-view CTR and reading depth. A sympathetic reader would care because it claims a practical path to rewrites that are both more faithful and more engaging, not one at the expense of the other.

What carries the argument

The load-bearing mechanism is the multiplicative rewrite-path score $s_{rw}(q,d,u)=\sum_{i=1}^{K}\alpha_i\,\mathrm{softplus}\big(\beta(u)\,\mathrm{sim}(q_i^{\mathrm{rank}},d^{\mathrm{rank}})+b(u)\big)$, fused with a residual direct-path score from the original query. Because the selector's $\alpha_i$ values form a masked probability distribution and the softplus keeps each item-relevance term non-negative, the product $\alpha_i r_i$ is large only when both selector confidence and item relevance are strong, which eliminates the generic-word shortcut. The second mechanism is gradient isolation, expressed as $\nabla_{\theta_{\mathrm{rec}}}\mathcal{L}_{\mathrm{main}}=0$, which shields recall-branch parameters from CTR-driven updates while the InfoNCE loss $\mathcal{L}_{\mathrm{NCE}}$ preserves semantic alignment. The third is the Dynamic Rewrite Selector, which jointly predicts request-level candidate weights and user-query-conditioned scale $\beta(u)$ and bias $b(u)$, letting both rewrite preference and relevance calibration adapt to each request.

What would settle it

Recompute the offline results with a similarity model trained only on human relevance judgments that never saw the new system's rewrite outputs or model weights, and with item pools produced independently by the old and new systems; if the +99.5% click-recall and +18.2% similarity gains shrink sharply under this independent measurement, the central claim would not be supported.

Watch

Extended reading notes

Core claim

The core claim is that the generic-word dominance effect is an architectural artifact rather than an unavoidable trade-off, and that three coordinated changes to a path-based rewrite-retrieval network remove it. First, the paper stops CTR gradients from reaching the recall branch, so the semantic geometry that retrieval depends on is not eroded by ranking feedback. Second, it replaces additive path scoring with multiplicative gating, so a rewrite can lift the final score only when its selection weight and its item relevance are both high. Third, it makes the rewrite selector dynamic: it predicts a distribution over rewrite candidates and per-request scale and bias terms conditioned on user and original-query representations, with binary cross-entropy on clicks supervising the whole path. The paper reports that the full system improves rewrite semantic similarity@10 by +18.2% and click recall@10 by +99.5% over Baseline, with the Dynamic Rewrite Selector contributing the largest single gain in recall and Multiplicative Gating contributing the largest single gain in similarity. Online A/B testing adds +0.259 in query-view CTR and +0.733 in average reading depth, which the paper reads as evidence that the engagement gains come from relevance rather than clickbait.

Load-bearing premise

The load-bearing premise is that the model used to score whether a rewrite matches the original query is an independent measure of intent faithfulness, and that the lists of items users were shown or clicked under the old system are fair measures of retrieval quality.

Editorial extensions

If this is right

  • Production rewrite selection can be moved out of the detached pre-processing stage: SPEAR is deployed on the Dewu community search platform, and its selector adds only about 25 ms per request within the existing 30-ms query-understanding stage.
  • Any additive path-based retrieval system that aggregates selection confidence with item relevance is exposed to the same generic-word dominance effect, since Multiplicative Gating alone raises click recall@10 by +16.4% over Baseline.
  • Separating recall and rank embeddings with stop-gradient changes the item-space geometry: Recall Purity@10 is 0.740 versus 0.229 for the collapsed shared space, and the Inter/Intra Ratio is 1.753 versus 1.039.
  • LLM-generated rewrites can plug into the same selector and gating machinery because SPEAR's components are encoder-agnostic; replacing the shared encoder and projection heads is a stated future direction.
  • Simultaneous gains in query-view CTR (+0.259) and reading depth (+0.733) support the paper's interpretation that retrieval quality, not surface-level attraction, drives the engagement lift, while retention gains remain non-significant.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit: decompose the +99.5% click recall@10 gain into coverage expansion versus better matching of already-exposed items; the Exposure Recall gain of +110.2% suggests coverage is a large component, but the paper does not isolate the two.
  • The Dynamic Rewrite Selector alone costs only -0.5% top-10 similarity while Multiplicative Gating adds +11.9%, suggesting the two components optimize nearly orthogonal axes; a variant that adds a small semantic-fidelity auxiliary to the selector's click objective might improve both metrics together.
  • The generic-word dominance mechanism is not e-commerce-specific: any retrieval system with a high-prior path or trigger term entering an additive score could test the multiplicative gating fix, for example in news or app search where broad category terms dominate.
  • Because the semantic similarity metric is produced by a production relevance model and the paper reports no confidence intervals or alternative splits, an external replication with an independently trained similarity model is the cleanest check on the +18.2% faithfulness claim.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper presents SPEAR, an end-to-end framework for query rewriting and retrieval in e-commerce community search, deployed on Dewu's platform. SPEAR targets a failure mode the authors call generic-word dominance, which arises when PDN-style path-based models, optimized for engagement, assign high selection scores to generic rewrites at the expense of semantic fidelity. The framework introduces three components: a dual-embedding backbone with a stop-gradient between rank and recall branches, a multiplicative gating aggregator that replaces additive path scoring, and a dynamic rewrite selector that generates request-specific rewrite weights and calibration parameters. Training combines a click-prediction loss with an InfoNCE semantic-alignment loss. Offline experiments on 100K held-out sessions report click recall@10 improving by +99.5%, exposure recall@10 by +110.2%, and semantic similarity@10 by +18.2% over the production PDN baseline. Online A/B testing reports significant gains in query-view CTR and reading depth, and a GSB human evaluation reports a +4.6% satisfaction lift. The code is released.

Significance. If the reported results are trustworthy, SPEAR is a substantial applied contribution: it identifies a concrete and plausible failure mode of path-based rewrite-retrieval models in search, proposes a principled architecture whose components are individually ablated, and validates the system with live traffic and human evaluation. The paper also makes reproducibility gestures in the right direction: public code, explicit hyperparameters, and a deployed-artifact evaluation. The main scientific value is in demonstrating that rewrite selection can be supervised by end-task retrieval outcomes while maintaining intent faithfulness, and in quantifying the generic-word dominance effect. That said, the current evidence for the intent-faithfulness claim rests on a semantic similarity metric whose independence from the candidate-generation process is not established, and the flagship gradient-isolation mechanism is incompletely specified as implemented. These issues are fixable, but until addressed they materially reduce confidence in the central differentiator.

major comments (3)
  1. [3.1.3, Eq. (7); 3.5, Eq. (21)] The claimed gradient isolation is incomplete. Eq. (7) sets the gradient of the main loss with respect to the recall-specific parameters to zero, but Eq. (21) states that the CTR objective updates the shared encoder parameters and the semantic-alignment objective also updates them. Because the recall-domain embedding is computed as a function of the shared encoder output, CTR gradients flowing into the shared encoder change the input to the recall branch, so ranking signals can still distort the recall-side geometry. The t-SNE and Purity analysis in Section 4.3.1 compares the trained Recall, Rank, and Shared spaces but does not isolate this shared-encoder gradient path. Please add an ablation with stop-gradient applied to the shared-encoder input of the recall branch (or a separate recall encoder) and show that recall-space purity and inter/intra-class ratios are unaffected, or revise the text to state precisely which gradient paths are blocked and why the residual path through the shared encoder does not undermine the 'shields recall-side semantics' claim.
  2. [4.1.1, 4.1.3, Table 2] The Semantic Similarity@K metric may be partly circular. The candidate rewrite pool is built using swing-based similarity mining and semantic ANN retrieval (108M query-query pairs per day), while Semantic Similarity@K is computed with a 'production-grade relevance model' trained on query-item pairs. The paper does not disclose whether this relevance model is the same model, or shares an encoder, with the semantic ANN index that generated the rewrite candidates. If the same representation geometry both proposes and scores the rewrites, the +18.2% gain measures proximity in the very space that generated the candidates rather than an independent assessment of intent faithfulness. Please specify the relevance model's architecture and training data and its overlap with the candidate-mining encoder, and report the correlation between Semantic Similarity@K and the human GSB judgments from Section 4.2.3 as a validity check.
  3. [4.1.1, Tables 1-2, Figure 3] No uncertainty quantification is provided for any offline metric. All numbers are point estimates from a single held-out day (2026-03-15), with no confidence intervals, bootstrap resampling, multiple evaluation days, or repeated training runs. Given that the headline claims are relative gains of +99.5% and +18.2%, and the ablation ordering is reported only as single values, it is hard to assess whether the differences are systematic. Figure 4 similarly reports significance stars but no effect-size confidence intervals for the online metrics. Please add confidence intervals over at least several held-out days or training seeds for Tables 1-2 and Figure 3, and report confidence intervals or exact p-values for the online metrics.
minor comments (5)
  1. [Abstract vs 4.2.2] The online gains are reported inconsistently: the abstract writes '+0.259 in query-view CTR and +0.733 in average reading depth' without percent signs, while Section 4.2.2 writes '+0.259%' and '+0.733%' and Figure 4 labels them relative improvements. Please clarify whether these are relative or absolute percentage-point changes.
  2. [4.1.1] The computation of Exposure Recall@K and Click Recall@K is underspecified; please state how the top-K rewrites are selected, which retriever and item pool are used, and whether the logged exposure and click sets come solely from the Baseline system.
  3. [4.2.3, Table 3] The GSB study is based on 150 queries with no inter-annotator agreement reported; please add agreement measures such as Cohen's kappa and describe how the 150 diff queries were sampled to avoid selection bias.
  4. [4.1.2] The three component ablations are described as independent additions, but the tuning protocol and training budget for each variant are not stated; reporting these details would make the large component-level differences more interpretable.
  5. [4.1.1 and 4.2.1] The offline and online evaluations compare only against the authors' own PDN production baseline; adding at least one non-PDN baseline, such as a standard two-stage rewrite-then-retrieve pipeline or an LLM-based rewriter, would strengthen the external validity of the claims.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: SPEAR's headline results are empirical evaluations against an external production baseline and an independently trained human-labeled relevance model, not consequences of the paper's definitions.

full rationale

SPEAR's central claims are empirical: +18.2% Semantic Similarity@10, +99.5% Click Recall@10, and online A/B gains. None of these numbers is derived from an equation in the paper; they are measurements on a held-out day (2026-03-15) and in a live A/B test. The semantic similarity metric is computed by an external 'production-grade relevance model trained offline on query–item pairs with human-labeled relevance' (Section 4.1.3), which is not shown to share weights or training data with SPEAR's recall branch. The Click/Exposure Recall metrics are computed against logged exposures and clicks from the production baseline, a standard if imperfect evaluation protocol; the comparison is not forced by construction because SPEAR's selector weights and ranking scores are trained with click labels and InfoNCE, not optimized to maximize these specific recall metrics at evaluation time. The component ablations show internal consistency (e.g., Dynamic Rewrite Selector alone slightly lowers semantic similarity while full SPEAR raises it), which would be unlikely if the semantic metric were merely a renamed training objective. The paper does contain self-citations ([49], [54]) by author Xiaobin Hu, but these relate to activation replay and latent-space computation in the Related Work section and are not load-bearing for SPEAR's claims. The reviewer's concern that the 'semantic ANN retrieval' used to mine 108M query–query pairs per day might share an encoder with the evaluation relevance model is a legitimate validity question, but the paper does not state this identity, and the metric is anchored to human relevance labels; without a quoted reduction, this remains a hypothesis rather than demonstrated circularity. Verdict: no circular derivation; score 0.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the validity of the industrial evaluation setup (metric definitions, data split, baseline fairness) and on the hand-chosen loss weights and temperatures. No new physical or formal entities are introduced.

free parameters (5)
  • lambda_align = 0.1
    Weight on InfoNCE auxiliary loss; chosen by hand, no sensitivity analysis reported (Section 4.1.1).
  • tau_NCE = 0.07
    InfoNCE temperature; set to standard value, no tuning reported (Section 4.1.1).
  • lambda_reg = 1e-6
    Weight decay; set in Section 4.1.1.
  • tau_init = 0.8
    Selector softmax temperature initialization (Section 4.1.1).
  • tau_floor = 0.6
    Selector softmax temperature lower bound (Section 4.1.1).
assumptions (4)
  • domain assumption The production-grade relevance model's embeddings are a valid measure of query intent faithfulness (Semantic Similarity@K).
    Section 4.1.1 defines the metric using this model; if the model is biased or shares training data with SPEAR's encoders, the +18.2% similarity claim is weakened.
  • domain assumption Click labels and exposure logs from the 180-day window are representative supervision for rewrite quality and retrieval coverage.
    Training and evaluation rely entirely on Dewu's click/exposure logs (Section 4.1.1); no external dataset is used.
  • domain assumption The single-week train/eval split (2026-03-08 to 03-14 train, 03-15 test) is representative of the traffic distribution.
    Reported in Section 4.1.1; no cross-validation or multiple splits.
  • domain assumption Stop-gradient between rank and recall branches prevents representation erosion without introducing optimization instability.
    Core design assumption in Section 3.1.3; not proven theoretically or with sensitivity analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search." pith.science (2026). https://pith.science/paper/PR2OVMAL

@misc{pith2026260801738,
  author       = {Pith},
  title        = {Pith review of: SPEAR: Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval for Community Search},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PR2OVMAL}},
  note         = {Machine review of arXiv:2608.01738}
}
read the original abstract

Query reformulation bridges user intent and retrieval in e-commerce search, yet production systems optimize rewrite quality and retrieval effectiveness separately, leaving the two stages structurally misaligned. Path-based architectures unify them end-to-end but were designed for personalization, where relevance is not an explicit constraint-search additionally requires the rewrite to remain faithful to the user's stated query intent. Transplanted directly, these models learn a shortcut we term the generic-word dominance effect: they favor generic rewrites that score well on paths but drift from query intent. To address this, we propose SPEAR (Selection-aware Personalized End-to-end Adaptive Rewriting and Retrieval), which integrates three components that each target one failure mode: (1) a dual-embedding backbone with auxiliary loss and gradient isolation that shields recall-side semantics from being eroded by CTR-driven ranking signals; (2) a multiplicative gating aggregator that lets a rewrite score high only when both its confidence and item relevance are strong, eliminating the generic-word shortcut; (3) a Dynamic Rewrite Selector that jointly generates request-specific rewrite weights and user-query-conditioned scale and bias terms, allowing both rewrite preference and relevance calibration to adapt to each request. Offline evaluation on 100K held-out industrial search sessions shows that the proposed framework improves rewrite semantic similarity@10 by +18.2 and click recall@10 by +99.5 over the production baseline. In online A/B testing, SPEAR achieves +0.259 in query-view CTR and +0.733 in average reading depth, confirming that improved rewrite selection translates into stronger retrieval and deeper user engagement. The proposed SPEAR system has been fully deployed in Dewu's community search platform since 2025. Our code is available at https://github.com/mallocagi1-cell/spear.

Figures

Figures reproduced from arXiv: 2608.01738 by the authors.

Figure 1
Figure 1. Community search results page in the Dewu mobile [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Overview of SPEAR. The framework integrates a Dual-Encoder backbone with gradient-isolated Rank/Recall branching [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Exposure Recall@5/10/30 across model vari [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Online A/B test: relative improvements of SPEAR [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: t-SNE visualization of item embeddings from the Recall, Rank, and Shared spaces ( [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

61 extracted references · 11 canonical work pages

  1. [1]

    Qingyao Ai, Daniel N. Hill, S. V. N. Vishwanathan, and W. Bruce Croft. 2019. A Zero Attention Model for Personalized Product Search. InProceedings of the 28th ACM International Conference on Information and Knowledge Management (CIKM ’19). ACM, New York, NY, USA, 379–388. doi:10.1145/3357384.3357980

  2. [2]

    Olivier Chapelle, Donald Metlzer, Ya Zhang, and Pierre Grinspan. 2009. Expected Reciprocal Rank for Graded Relevance. InProceedings of the 18th ACM Conference on Information and Knowledge Management (CIKM ’09). ACM, New York, NY, USA, 621–630. doi:10.1145/1645953.1646033

  3. [3]

    Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich. 2018. GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Mul- titask Networks. InProceedings of the 35th International Conference on Machine Learning (ICML ’18) (Proceedings of Machine Learning Research, Vol. 80). PMLR, Brookline, MA, USA, 794–803

  4. [4]

    Jaekeol Choi, Euna Jung, Jangwon Suh, and Wonjong Rhee. 2021. Improving Bi-Encoder Document Ranking Models with Two Rankers and Multi-Teacher Distillation. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’21). ACM, New York, NY, USA, 2192–2196. arXiv:2103.06523

  5. [5]

    Paul Covington, Jay Adams, and Emre Sargin. 2016. Deep Neural Networks for YouTube Recommendations. InProceedings of the 10th ACM Conference on Recommender Systems (RecSys ’16). ACM, New York, NY, USA, 191–198. doi:10. 1145/2959100.2959190

  6. [6]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-Training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Associ- ation for Computational Linguistics: Human Language Technologies (NAACL ’19). Association for Computational Linguistics, ...

  7. [7]

    Luyu Gao and Jamie Callan. 2021. Condenser: a Pre-Training Architecture for Dense Retrieval. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP ’21). Association for Computational Linguistics, Stroudsburg, PA, USA, 981–993. doi:10.18653/v1/2021.emnlp-main.75

  8. [8]

    Luyu Gao, Xueguang Ma, Jimmy Lin, and Jamie Callan. 2023. Precise Zero-Shot Dense Retrieval without Relevance Labels. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL ’23). Association for Computational Linguistics, Toronto, Canada, 1762–1777. doi:10.18653/v1/ 2023.acl-long.99

Show all 61 references
  1. [9]

    Weihao Gao, Xiangjun Fan, Jiankai Sun, Kai Jia, Wenzhi Xiao, Chong Ding, and Bo Long. 2021. Deep Retrieval: Learning A Retrievable Structure for Large-Scale Recommendations. InProceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM ’2...

  2. [10]

    Jiafeng Guo, Gu Xu, Xueqi Cheng, and Hang Li. 2009. Named Entity Recognition in Query. InProceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’09). ACM, New York, NY, USA, 267–274. doi:10.1145/1571941.1571989

  3. [11]

    Yunlong He, Jiliang Tang, Hua Ouyang, Changsung Kang, Dawei Yin, and Yi Chang. 2016. Learning to Rewrite Queries. InProceedings of the 25th ACM International Conference on Information and Knowledge Management (CIKM ’16). ACM, Indianapolis, IN, USA, 1443–1452. doi:10.1145/29833...

  4. [12]

    Jui-Ting Huang, Ashish Sharma, Shuying Sun, Li Xia, David Zhang, Philip Pronin, Janani Padmanabhan, Giuseppe Ottaviano, and Linjun Yang. 2020. Embedding- Based Retrieval in Facebook Search. InProceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & ...

  5. [13]

    Po-Sen Huang, Xiaodong He, Jianfeng Gao, Li Deng, Alex Acero, and Larry Heck. 2013. Learning Deep Structured Semantic Models for Web Search using Clickthrough Data. InProceedings of the 22nd ACM International Conference on Information & Knowledge Management (CIKM ’13). ACM, Ne...

  6. [14]

    Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bo- janowski, Armand Joulin, and Edouard Grave. 2022. Unsupervised Dense Infor- mation Retrieval with Contrastive Learning. Transactions on Machine Learning Research. arXiv:2112.09118

  7. [15]

    Rolf Jagerman, Honglei Zhuang, Zhen Qin, Xuanhui Wang, and Michael Ben- dersky. 2023. Query Expansion by Prompting Large Language Models. arXiv preprint arXiv:2305.03653. arXiv:2305.03653

  8. [16]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated Gain-Based Evaluation of IR Techniques.ACM Transactions on Information Systems20, 4 (2002), 422–446. doi:10.1145/582415.582418

  9. [17]

    Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019. Billion-Scale Similarity Search with GPUs.IEEE Transactions on Big Data7, 3 (2019), 535–547. doi:10. 1109/TBDATA.2019.2921572

  10. [18]

    Vladimir Karpukhin, Barlas Oğuz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP ...

  11. [19]

    Alex Kendall, Yarin Gal, and Roberto Cipolla. 2018. Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR ’18). IEEE Computer Society, Los Alamitos, CA, USA,...

  12. [20]

    Houyi Li, Zhihong Chen, Chenliang Li, Rong Xiao, Hongbo Deng, Peng Zhang, Yongchao Liu, and Haihong Tang. 2021. Path-Based Deep Network for Candidate Item Matching in Recommenders. InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Info...

  13. [21]

    Kwok, and Qianli Ma

    Sen Li, Fuyu Lv, Taiwei Jin, Guiyang Li, Yukun Zheng, Tao Zhuang, Qingwen Liu, Xiaoyi Zeng, James T. Kwok, and Qianli Ma. 2022. Query Rewriting in TaoBao Search. InProceedings of the 31st ACM International Conference on Information and Knowledge Management (CIKM ’22). ACM, Atl...

  14. [22]

    Sen Li, Fuyu Lv, Taiwei Jin, Guli Lin, Keping Yang, Xiaoyi Zeng, Xiao-Ming Wu, and Qianli Ma. 2021. Embedding-Based Product Retrieval in Taobao Search. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD ’21). ACM, New York, NY, USA, 3181...

  15. [23]

    Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. In-Batch Negatives for Knowledge Distillation with Tightly-Coupled Teachers for Dense Retrieval. InProceedings of the 6th Workshop on Representation Learning for NLP (RepL4NLP- 2021). Association for Computational Linguist...

  16. [24]

    Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. Multi-Stage Conver- sational Passage Retrieval: An Approach to Fusing Term Importance Estimation and Neural Query Rewriting.ACM Transactions on Information Systems39, 4 (2021), 1–29. doi:10.1145/3446426

  17. [25]

    Yiqun Liu, Kaushik Rangadurai, Yunzhong He, Siddarth Malreddy, Xunlong Gui, Xiaoyi Liu, and Fedor Borisyuk. 2021. Que2Search: Fast and Accurate Query and Document Understanding for Search at Facebook. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Dat...

  18. [26]

    Zheng Liu, Chaofan Li, Shitao Xiao, Yingxia Shao, and Defu Lian. 2024. Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL ’24). Association for Computa...

  19. [27]

    Jiaqi Ma, Zhe Zhao, Xinyang Yi, Jilin Chen, Lichan Hong, and Ed H. Chi. 2018. Modeling Task Relationships in Multi-Task Learning with Multi-Gate Mixture- of-Experts. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’18). A...

  20. [28]

    Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine- Tuning LLaMA for Multi-Stage Text Retrieval. InProceedings of the 47th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’24). ACM, New York, NY, USA, 2421–24...

  21. [29]

    Alessandro Magnani, Feng Liu, Suthee Chaidaroon, Sachin Yadav, Praveen Reddy Suram, Ajit Puthenputhussery, Sijie Chen, Min Xie, Anirudh Kashi, Tony Lee, and Ciya Liao. 2022. Semantic Retrieval at Walmart. InProceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery &...

  22. [30]

    Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Seungyeon Kim, Sashank Reddi, and Sanjiv Kumar. 2022. In Defense of Dual-Encoders for Neural Ranking. InProceedings of the 39th International Conference on Machine Learning (ICML ’22) (Proceedings of Machine Learning ...

  23. [31]

    Akash Kumar Mohankumar, Nikit Begwani, and Amit Singh. 2021. Diversity Driven Query Rewriting in Search Advertising. InProceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery & Data Mining (KDD ’21). ACM, New York, NY, USA, 3423–3431. doi:10.1145/3447548.3467204

  24. [32]

    Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernandez Abrego, Ji Ma, Vincent Zhao, Yi Luan, Keith Hall, Ming-Wei Chang, and Yinfei Yang. 2022. Large Dual Encoders Are Generalizable Retrievers. InProceedings of the 2022 Conference on Empirical Methods in Natural Language P...

  25. [33]

    Priyanka Nigam, Yiwei Song, Vijai Mohan, Vihan Lakshman, Weitian Ding, Ankit Shingavi, Choon Hui Teo, Hao Gu, and Bing Yin. 2019. Semantic Product Search. InProceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’19). ACM, New Yor...

  26. [34]

    Tong Niu, Shafiq Joty, Ye Liu, Caiming Xiong, Yingbo Zhou, and Semih Yavuz

  27. [35]

    Rodrigo Nogueira and Kyunghyun Cho. 2017. Task-Oriented Query Reformula- tion with Reinforcement Learning. InProceedings of the 2017 Conference on Empir- ical Methods in Natural Language Processing (EMNLP ’17). Association for Compu- tational Linguistics, Copenhagen, Denmark, ...

  28. [36]

    Wenjun Peng, Guiyang Li, Yue Jiang, Zilong Wang, Dan Ou, Xiaoyi Zeng, and Enhong Chen. 2024. Large Language Model Based Long-Tail Query Rewriting in Taobao Search. InCompanion Proceedings of the ACM Web Conference 2024 (WWW ’24 Companion). ACM, New York, NY, USA, 20–28. arXiv:...

  29. [37]

    Qi Pi, Guorui Zhou, Yujing Zhang, Zhe Wang, Lejian Ren, Ying Fan, Xiaoqiang Zhu, and Kun Gai. 2020. Search-Based User Interest Modeling with Lifelong Sequential Behavior Data for Click-Through Rate Prediction. InProceedings of the 29th ACM International Conference on Informati...

  30. [38]

    Yiming Qiu, Kang Zhang, Han Zhang, Songlin Wang, Sulong Xu, Yun Xiao, Bo Long, and Wen-Yun Yang. 2021. Query Rewriting via Cycle-Consistent Transla- tion for E-Commerce Search. In2021 IEEE 37th International Conference on Data Engineering (ICDE). IEEE, Piscataway, NJ, USA, 2435–2446

  31. [39]

    Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxi- ang Dong, Hua Wu, and Haifeng Wang. 2021. RocketQA: An Optimized Training Approach to Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2021 Conference of the North Am...

  32. [40]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP ’19). Association for Computa- tional Linguistics, Hong Kong, China, 3982–3992...

  33. [41]

    Stefan Riezler and Yi Liu. 2010. Query Rewriting Using Monolingual Statistical Machine Translation.Computational Linguistics36, 3 (2010), 569–582. doi:10. 1162/coli_a_00010

  34. [42]

    Stephen Robertson and Hugo Zaragoza. 2009. The Probabilistic Relevance Frame- work: BM25 and Beyond.Foundations and Trends in Information Retrieval3, 4 (2009), 333–389. doi:10.1561/1500000019

  35. [43]

    Parikshit Sondhi, Mohit Sharma, Pranam Kolari, and ChengXiang Zhai. 2018. A Taxonomy of Queries for E-Commerce Search. InProceedings of the 41st Interna- tional ACM SIGIR Conference on Research & Development in Information Retrieval (SIGIR ’18). ACM, New York, NY, USA, 1245–12...

  36. [44]

    Hongyan Tang, Junning Liu, Ming Zhao, and Xudong Gong. 2020. Progres- sive Layered Extraction (PLE): A Novel Multi-Task Learning (MTL) Model for Personalized Recommendations. InProceedings of the 14th ACM Conference on Recommender Systems (RecSys ’20). ACM, New York, NY, USA, ...

  37. [45]

    Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, and Furu Wei. 2024. Improving Text Embeddings with Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL ’24). Association for Computational Lin...

  38. [46]

    Liang Wang, Nan Yang, and Furu Wei. 2023. Query2doc: Query Expansion with Large Language Models. InProceedings of the 2023 Conference on Empirical Meth- ods in Natural Language Processing (EMNLP ’23). Association for Computational Linguistics, Singapore, 9414–9423. doi:10.1865...

  39. [47]

    Xiao Wang, Sean MacAvaney, Craig Macdonald, and Iadh Ounis. 2023. Genera- tive Query Reformulation for Effective Adhoc Search. The First Workshop on Generative Information Retrieval (Gen-IR@SIGIR ’23). arXiv:2308.00415

  40. [48]

    Rong Xiao, Jianhui Ji, Baoliang Cui, Haihong Tang, Wenwu Ou, Yanghua Xiao, Jiwei Tan, and Xuan Ju. 2019. Weakly Supervised Co-Training of Query Rewrit- ing and Semantic Matching for E-Commerce. InProceedings of the 12th ACM International Conference on Web Search and Data Minin...

  41. [49]

    Yun Xing, Xiaobin Hu, Qingdong He, Jiangning Zhang, Shuicheng Yan, Shijian Lu, and Yu-Gang Jiang. 2026. Boosting Reasoning in Large Multimodal Models via Activation Replay. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE Compute...

  42. [50]

    Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. International Conference on Learning Representations (ICLR ’21). arXiv:2007.00808

  43. [52]

    Shi Yu, Jiahua Liu, Jingqin Yang, Chenyan Xiong, Paul Bennett, Jianfeng Gao, and Zhiyuan Liu. 2020. Few-Shot Generative Conversational Query Rewriting. InProceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR ’20)...

  44. [53]

    Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn. 2020. Gradient Surgery for Multi-Task Learning. InAdvances in Neural Information Processing Systems 33 (NeurIPS ’20). Curran Associates, Inc., Red Hook, NY, USA, 5824–5836

  45. [54]

    Xinlei Yu, Zhangquan Chen, Yongbo He, Tianyu Fu, Cheng Yang, Chengming Xu, Yue Ma, Xiaobin Hu, Zhe Cao, Jie Xu, Guibin Zhang, Jiale Tao, Jiayi Zhang, Siyuan Ma, Kaituo Feng, Haojie Huang, Youxing Li, Ronghao Chen, Huacan Wang, Chenglin Wu, Zikun Su, Xiaogang Xu, Kelu Yao, Kun ...

  46. [55]

    Han Zhang, Songlin Wang, Kang Zhang, Zhiling Tang, Yunjiang Jiang, Yun Xiao, Weipeng Yan, and Wen-Yun Yang. 2020. Towards Personalized and Semantic Retrieval: An End-to-End Solution for E-Commerce Search via Embedding Learn- ing. InProceedings of the 43rd International ACM SIG...

  47. [56]

    Haonan Zhang, Yanzhao Zhang, Dingkun Long, and Pengjun Xie. 2024. TriSam- pler: A Better Negative Sampling Principle for Dense Retrieval. InProceedings of the 38th AAAI Conference on Artificial Intelligence (AAAI ’24). AAAI Press, Palo Alto, CA, USA, 9269–9277. doi:10.1609/aaa...

  48. [57]

    Yukun Zheng, Jiang Fan, Yitian Zhang, Xiao Jin, Yue Liu, Shuchang Xu, Keping Yang, Hao Xu, and Xiaoyi Zeng. 2022. Multi-Objective Personalized Product Retrieval in Taobao Search. arXiv preprint arXiv:2210.04170. arXiv:2210.04170

  49. [58]

    Guorui Zhou, Xiaoqiang Zhu, Chenru Song, Ying Fan, Han Zhu, Xiao Ma, Yanghui Yan, Junqi Jin, Han Li, and Kun Gai. 2018. Deep Interest Network for Click- Through Rate Prediction. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining ...

  50. [59]

    Chuanqi Zhu, Ming Gao, Bin Zhu, and Jindong Chen. 2021. More Robust Dense Retrieval with Contrastive Dual Learning. arXiv preprint arXiv:2107.07773. arXiv:2107.07773

  51. [60]

    Han Zhu, Xiang Li, Pengye Zhang, Guozheng Li, Jie He, Han Li, and Kun Gai

  52. [2018]

    InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’18)

    Learning Tree-Based Deep Model for Recommender Systems. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (KDD ’18). ACM, New York, NY, USA, 1079–1088. doi:10.1145/3219819. 3219826

  53. [2024]

    arXiv preprint arXiv:2411.00142

    JudgeRank: Leveraging Large Language Models for Reasoning-Intensive SPEAR RecSys ’26, September 27-October 02, 2026, Minneapolis, MN, USA Reranking. arXiv preprint arXiv:2411.00142. arXiv:2411.00142

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.