Pith. sign in

REVIEW 3 major objections 6 minor 69 references

CoRA converts a frozen encoder into a task-conditioned retriever using only closed-form linear algebra, so demonstration selection never needs training.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 01:45 UTC pith:VDNKJTRG

load-bearing objection Solid, incremental gradient-free retrieval work whose on-device framing outruns its evaluation: homogeneous pools only, and the release is missing. the 3 major comments →

arxiv 2607.27766 v1 pith:VDNKJTRG submitted 2026-07-30 cs.CL cs.IRcs.LG

Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning

classification cs.CL cs.IRcs.LG
keywords in-context learningexemplar retrievalon-device retrievaltask-conditioned retrievalclosed-form ridge regressionlow-rank projectionfrozen encodermultimodal retrieval
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper sets out to resolve the tension in on-device in-context learning between task-aware demonstration selection and lightweight, local retrieval. It claims that a frozen text encoder can be turned into a task-conditioned retriever by fitting candidate input representations to a conditioning space built from paired candidate outputs, using ridge regression and a low-rank projection. If correct, output-aware exemplar selection becomes possible without retriever fine-tuning, backpropagation, or calls to the downstream model, and the whole index can be built in two streaming passes. The authors extend the same alignment to multimodal retrieval and report end-to-end feasibility on a Raspberry Pi 5.

Core claim

The central claim is that the components of a frozen encoder's layerwise input representations that are linearly explained by candidate outputs define a useful retrieval subspace. CoRA builds a conditioning matrix from mean-pooled embeddings of candidate outputs, fits the selected-layer input features to it by closed-form ridge regression, and takes the top-r right singular vectors of the fitted matrix as a compact retrieval basis. Proposition 3.1 shows this basis maximizes retained fitted energy among all rank-r orthogonal projections, i.e., it is the optimal low-rank compression of the output-conditioned fit. At query time only the input is encoded and projected through the precomputed bas

What carries the argument

The conditioning matrix C (a column of ones plus standardized mean-pooled frozen-encoder embeddings of candidate outputs, optionally with visual embeddings in the multimodal variant) and the ridge-fitted projection P_C = C(C^T C + lambda I)^{-1} C^T. CoRA applies P_C to layerwise input representations, concatenates the fitted blocks, and takes the top-r right singular vectors V_{1:r} of the fitted matrix as the retrieval basis. This object carries the argument: it converts output-side regularities into a fixed, compact projection that transfers task information to unseen queries using only their inputs, and its optimality is characterized by Proposition 3.1.

Load-bearing premise

The load-bearing premise is that mean-pooled frozen-encoder embeddings of candidate outputs (plus visual features in the multimodal variant) carry a stable, task-relevant signal about which demonstrations will help an unseen query; if output surface form is only weakly tied to demonstration utility, the fitted subspace encodes irrelevant structure and query projections degrade.

What would settle it

Swap or randomly permute the output embeddings used in the conditioning matrix while keeping inputs, index size, and query pipeline fixed; if downstream ICL accuracy does not drop below CoRA's, the output-derived conditioning is not what carries the gains. The paper runs this check only for shuffled visual features in CoRA-M, not for text outputs.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Demonstration selection can become output-aware on-device without training a retriever or querying the target LLM, closing the gap between task-agnostic similarity search and learning-based selection.
  • Index construction is a two-pass streaming computation whose working memory depends on chunk and feature dimensions, not on the number of candidates, so pools can grow without materializing large matrices.
  • Since query-time retrieval is a single projection plus nearest-neighbor search in r dimensions, standard ANN indexing and CPU/edge hardware apply directly.
  • The same closed-form alignment extends to multimodal ICL by appending visual features to conditioning and retrieval spaces, giving a unified retrieval formulation across text and vision-language tasks.
  • Because CoRA is gradient-free, it runs where training-based retrievers run out of memory, including the reported Raspberry Pi end-to-end pipeline.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The optimality result is about the fitted matrix, not about downstream ICL accuracy; a natural extension would test whether the same basis remains optimal when the retrieval objective is the end-to-end language-model score, and if not, what closed-form correction restores it.
  • The paper's own NL2Bash result suggests the method's load-bearing assumption can be stressed: tasks with many-to-many input-output mappings, where surface-form outputs are weak proxies for demonstration utility, are exactly where output conditioning may need richer or learned output representations.
  • Because the streaming construction maintains sufficient statistics G and T, an incremental update variant that absorbs new candidate pairs without a full rebuild seems directly within reach, though the paper leaves evolving pools as future work.
  • The conditioning mechanism is not tied to a particular encoder; applying the same alignment to other frozen encoders, or using task labels instead of outputs as the conditioning signal, would test how general the principle is.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper proposes CoRA, a gradient-free, output-conditioned exemplar retriever for on-device in-context learning. CoRA selects representative encoder layers via CKA, constructs a conditioning matrix from pooled candidate outputs, fits candidate input representations to that matrix through closed-form ridge regression, and extracts a low-rank retrieval basis by SVD. Query time requires only the query input and a precomputed index. The authors prove that the retained basis optimally preserves the fitted representation (Prop. 3.1), derive an exact two-pass streaming construction (Sec. 4.1), and extend the method to multimodal retrieval (CoRA-M). Experiments cover ten textual datasets and four multimodal benchmarks with several downstream LLMs, plus a Raspberry Pi 5 deployment. The paper claims consistent gains over static retrieval baselines and the prior MLSM method.

Significance. If the empirical claims are supported, CoRA would be a useful contribution: it provides output-informed, task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls, and the streaming index construction with memory independent of n is a genuine practical advantage. Proposition 3.1 is correctly proved, and the exact two-pass derivation in Sec. 4.1 is a clear strength. The Raspberry Pi 5 end-to-end evaluation is also a positive feature. However, several load-bearing validation gaps remain: the default rank is selected on the evaluation datasets themselves, the motivating heterogeneous on-device memory regime is never tested, and the reported error bars are undefined. The significance of the headline claims is therefore not yet fully established.

major comments (3)
  1. [§5.3, Table 5] The default subspace rank r=256 is selected from a grid r∈{32,64,128,256,512} by examining CoRA's performance on the same evaluation datasets (SST-5, QNLI, WebQs, MTOP) that later appear in the headline tables (Tables 2–3). No held-out validation split or nested hyperparameter procedure is described. Since r is a free parameter and is tuned on the test tasks, the comparison against baselines is optimistically biased by test-set model selection. Please report results with validation-based r, or show that the conclusions are stable across r on a separate split.
  2. [§5.2, Tables 2–3; §6] The motivating use case is a single local, heterogeneous memory, yet all retrieval experiments use per-dataset homogeneous pools (Table 1). The paper's own NL2Bash results show that output-surface variability weakens the conditioning signal: CoRA is below BM25 or SBERT on NL2Bash under every backbone (e.g., 0.3248 vs. BM25 0.3346 for Llama-3.2-1B; 0.2871 vs. SBERT 0.3498 for MobileLLM-Pro). A mixed-task pool would combine heterogeneous output spaces and plausibly amplify this failure mode. Without a mixed-pool experiment, or an explicit restriction of the claim to single-task pools, the abstract's assertion that CoRA converts a frozen encoder into a task-conditioned retriever for on-device local memories is not empirically supported.
  3. [§5.2, Tables 2–4] The ± values attached to CoRA results (e.g., 0.6486±0.02, 0.3951±0.02, 39.68±0.01) are never defined: there is no statement of the number of runs, the set of seeds, or whether the intervals are standard deviations or standard errors. Baselines are reported without intervals, so the reader cannot judge whether the headline gains over MLSM or BERT, which are often only 0.01–0.03, are meaningful. Please define the interval and report comparable variability for all methods.
minor comments (6)
  1. [§3.3, Table 8] The retained-fitted-energy comparison in Table 8 is, by construction, maximized by the SVD basis: Eq. (14)–(18) define V_{1:r} as the optimizer of exactly the quantity reported in the table. This does not invalidate the downstream metric column, but the fitted-energy column is not independent evidence for the method. Please present it as a sanity check rather than as empirical validation.
  2. [§4.1, Eq. (21)] The 'two-pass' streaming description does not state where the Z-score standardization statistics (mean and standard deviation) are computed. If they require a full pass over the pool, this is effectively an extra pre-pass or must be accumulated as sufficient statistics during Pass 1. Please clarify so that the pass count is exact.
  3. [Table 4] The table caption says 'Avg. denotes the average score across all five tasks,' but only four multimodal benchmarks are listed. Please correct the caption.
  4. [Figure 2] Figure 2 appears corrupted in the manuscript as rendered: the text contains long '/uni000...' placeholder sequences, making the figure unreadable. Please regenerate the figure.
  5. [§5.3, Table 8] The downstream metric column in Table 8 does not specify the backbone, prompt configuration, or k. The MRPC value 0.7353 appears to coincide with Table 10 (Raspberry Pi deployment with Qwen3.5-0.8B), which is confusing. Please specify the evaluation setting.
  6. [General] No code or data availability statement is provided. To support reproducibility, please indicate whether the implementation and processed datasets will be released.

Circularity Check

1 steps flagged

CoRA's 'optimal low-rank basis' claim is tautological (V is defined as the SVD of the fitted matrix), but the downstream ICL-accuracy claims are externally evaluated and not circular.

specific steps
  1. self definitional [Sec. 3.2 Eqs. (6)-(7); Sec. 3.3 Proposition 3.1; Sec. 5.3 Table 8]
    "Finally, to obtain a compact retrieval space suitable for on-device deployment, we perform a spectral decomposition ˆH𝑥 = UΣV⊤, and retain the top-r right singular vectors V 1:r ∈ R dcat×r. ... Proposition 3.1 (Optimal Rank-r Compression of the Output-Conditioned Fit). For 1 ≤ r ≤ dcat, V 1:r is an optimal solution to max W∈R dcat×r, W⊤W=Ir ||ˆH𝑥 W|| 2 F ... Table 8: 'All values of ρr(W) are evaluated with respect to the same output-conditioned fitted matrix ˆH𝑥.'"

    V 1:r is defined in Eq. (7) as the top-r right singular vectors of ˆH𝑥, and Proposition 3.1 then proves that this same V 1:r maximizes the retained fitted energy ||ˆH𝑥 W||_F^2. This is not a derived consequence but the definition of the SVD (Eckart–Young); the 'optimality' claim is true by construction. Table 8 then compares bases on exactly ρr(W) = ||ˆH𝑥 W||_F^2 / ||ˆH𝑥||_F^2 with respect to the same fitted matrix, so the ordering CoRA > unconditioned > random is forced by the choice of V rather than being an empirical discovery. The downstream ICL metrics in Table 8 are external and do break the circularity for the main empirical claim.

full rationale

The paper's derivational core is largely self-contained linear algebra: the conditioning matrix C, the ridge-fitted representation ˆH𝑥 = PC ˜H𝑥, and the streaming construction are all closed-form and do not presuppose the retrieval results. The main circular element is the formal optimality proposition: because V 1:r is chosen as the SVD of ˆH𝑥, the claim that it is the optimal low-rank compressor of ˆH𝑥 is a restatement of the construction, and Table 8's retained-energy comparison is therefore not independent evidence. However, the paper's central empirical claim—that CoRA improves ICL accuracy over BM25, SBERT, and MLSM on ten textual and four multimodal benchmarks—is evaluated with external downstream models, disjoint query sets, and withheld evaluation targets, so it does not reduce to the fitted inputs. The self-citations to the authors' prior MLSM work [36] supply the representative-layer selection procedure and hypotheses, but that component is re-described and ablated in this paper (Tables 5 and 6), and the prior work is an independent publication, so it is not load-bearing in a circular sense. The paper also acknowledges the NL2Bash failure mode, which is a limitation rather than a circularity. Overall, the circularity is partial and confined to a formal claim; the main empirical contribution stands independently.

Axiom & Free-Parameter Ledger

5 free parameters · 6 axioms · 0 invented entities

The method contributes closed-form linear algebra, but its retrieval quality rests on several domain assumptions (frozen-layer cues, output embeddings as task signal, transfer from candidate pool to queries) and on hyperparameters inherited from prior work or tuned on evaluation data. No new entities are postulated.

free parameters (5)
  • r (retained subspace dimension) = 256
    Selected from ablations on SST-5/QNLI/WebQs/MTOP (Table 5), i.e., tuned on the evaluation datasets; main results use this value.
  • lambda (ridge regularization) = 1e-6
    Hand-set small constant; the paper does not study sensitivity.
  • n_l (number of representative layers) = 3
    Inherited from the authors' MLSM configuration; fixed across all experiments.
  • k (number of demonstrations) = 20
    Fixed protocol value from prior work; task-dependent optimum shown in Fig. 2 but not used.
  • n_s (calibration subset size for CKA) = 1000
    Hand-set; no sensitivity analysis provided.
axioms (6)
  • domain assumption Mean-pooled layerwise token representations (Eq. 1) capture complementary lexical, syntactic, and semantic cues useful for retrieval.
    Adopted from probing literature and the authors' prior MLSM analysis; no proof that these cues suffice.
  • domain assumption CKA-based layer clustering with k-means selects a complementary, low-redundancy layer set (Eqs. 2-3).
    Inherited from reference [36]; the paper does not validate this selection independently.
  • domain assumption Frozen-encoder representations of candidate outputs provide a valid conditioning signal for retrieval (Eqs. 4-5).
    Core premise of the method; supported only indirectly by downstream ICL accuracy.
  • domain assumption The dominant right singular subspace of the fitted matrix Hhat_x generalizes from candidate pool to disjoint test queries (Eqs. 7-8).
    Assumes distributional alignment between the local candidate memory and the queries; not validated under pool shift.
  • standard math Ridge regression with lambda*I regularization and the Eckart-Young theorem justify Eqs. (5) and (14)-(16).
    Standard linear algebra; the proof of Prop. 3.1 restates Eckart-Young without citing it.
  • domain assumption For multimodal CoRA-M, textual content dominates and final-layer target text plus visual features are sufficient conditioning (Eqs. 10-12).
    Justified by prior multimodal ICL observations [6, 11]; not independently verified here.

pith-pipeline@v1.3.0-daily-deepseek · 25065 in / 14092 out tokens · 142456 ms · 2026-08-01T01:45:15.134757+00:00 · methodology

0 comments
read the original abstract

On-device in-context learning (ICL) relies on pre-inference retrieval to select demonstrations for useful context before downstream model inference. This retrieval must exploit task-specific information while operating over local memories under limited computation, memory, and data-exposure budgets. We propose Conditional Retrieval Alignment (CoRA), a gradient-free framework that converts a frozen encoder into a task-conditioned retriever using paired candidate inputs and outputs. CoRA selects complementary encoder layers, constructs an output-derived conditioning space from candidate memory, and aligns candidate input representations to this space through closed-form ridge regression. Low-rank factorization then produces a compact retrieval basis where candidate outputs are used only during offline index construction, whereas query-time retrieval requires only the query input and precomputed index. We show that CoRA's rank-constrained basis is the optimal low-rank compression of the output-conditioned fitted representation, and derive an exact two-pass streaming construction that avoids materializing the full fitted matrix. We further extend the framework to multimodal exemplar retrieval by incorporating visual representations into the conditioning and retrieval spaces. Experiments across ten textual datasets and four multimodal benchmarks with Llama-3.2-1B, MobileLLM-Pro, OpenFlamingo-3B, and Qwen3.5-2B, as well as end-to-end Raspberry Pi~5 deployment demonstrate that CoRA supports effective task-conditioned retrieval without retriever fine-tuning, backpropagation, or target-model calls.

Figures

Figures reproduced from arXiv: 2607.27766 by Arindam Basu, Haoliang Li, Hui Liu, Junyi Yang, Xinyu Luo, Yihua Shao.

Figure 1
Figure 1. Figure 1: Left: Shortcomings of existing retrieval methods. Right: Overview of CoRA framework, which consists of three main stages. By projecting representations into a task-conditioned subspace without gradient-based parameter updates, it enables effective textual and multimodal retrieval on edge devices. Method Desiderata. The above setup and preliminaries suggest three desiderata for our retrieval alignment: 𝑖) t… view at source ↗
Figure 2
Figure 2. Figure 2: Performance comparison on four datasets with various number of exemplars [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Peak CPU memory (MB) versus corpus size during the pre-inference retrieval workflow, with and without chunked streaming. The shuffled-visual variant performs below CoRA-T on both datasets. By preserving visual feature dimensionality while disrupting their correspondence with candidate examples, this result indicates that CoRA-M gains depend on correctly aligned visual information rather than additional fea… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

69 extracted references · 3 canonical work pages

  1. [1]

    Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katherine Millican, Malcolm Reynolds, et al. 2022. Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems35 (2022), 23716–23736

  2. [2]

    Shengnan An, Bo Zhou, Zeqi Lin, Qiang Fu, Bei Chen, Nanning Zheng, Weizhu Chen, and Jian-Guang Lou. 2023. Skill-Based Few-Shot Selection for In-Context Learning. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 13472–13492

  3. [3]

    Anderson, Z

    E. Anderson, Z. Bai, C. Bischof, S. Blackford, J. Demmel, J. Dongarra, J. Du Croz, A. Greenbaum, S. Hammarling, A. McKenney, and D. Sorensen. 1999.LAPACK Users’ Guide(third ed.). Society for Industrial and Applied Mathematics, Philadelphia, PA. 22 Luo et al

  4. [4]

    Jacob Andreas, John Bufe, David Burkett, Charles Chen, Josh Clausman, Jean Crawford, Kate Crim, Jordan DeLoach, Leah Dorner, Jason Eisner, et al. 2020. Task-oriented dialogue as dataflow synthesis.Transactions of the Association for Computational Linguistics8 (2020), 556–571

  5. [5]

    Anas Awadalla, Irena Gao, Josh Gardner, Jack Hessel, Yusuf Hanafy, Wanrong Zhu, Kalyani Marathe, Yonatan Bitton, Samir Gadre, Shiori Sagawa, et al. 2023. Openflamingo: An open-source framework for training large autoregressive vision-language models.arXiv preprint arXiv:2308.01390(2023)

  6. [6]

    Folco Bertini Baldassini, Mustafa Shukor, Matthieu Cord, Laure Soulier, and Benjamin Piwowarski. 2024. What makes multimodal in-context learning work?. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1539–1550

  7. [7]

    Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on freebase from question-answer pairs. InProceedings of the 2013 conference on empirical methods in natural language processing. 1533–1544

  8. [8]

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in neural information processing systems33 (2020), 1877–1901

  9. [9]

    Cheng Chen, Yunpeng Zhai, Yifan Zhao, Jinyang Gao, Bolin Ding, and Jia Li. 2025. Provoking multi-modal few-shot lvlm via exploration-exploitation in-context learning. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 3826–3835

  10. [10]

    Laming Chen, Guoxin Zhang, and Eric Zhou. 2018. Fast greedy map inference for determinantal point process to improve recommendation diversity.Advances in neural information processing systems31 (2018)

  11. [11]

    Shuo Chen, Zhen Han, Bailan He, Jianzhe Liu, Mark Buckley, Yao Qin, Philip Torr, Volker Tresp, and Jindong Gu. 2025. Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?. In2025 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). IEEE, 6000–6010

  12. [12]

    Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick

  13. [13]

    Fahim Dalvi, Hassan Sajjad, Nadir Durrani, and Yonatan Belinkov. 2020. Analyzing redundancy in pretrained transformer models. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). 4908–4926

  14. [14]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186

  15. [15]

    Yucheng Ding, Chaoyue Niu, Fan Wu, Shaojie Tang, Chengfei Lyu, and Guihai Chen. 2024. Enhancing On-Device LLM Inference with Historical Cloud-Based LLM Interactions. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining(Barcelona, Spain)(KDD ’24). Association for Computing Machinery, New York, NY, USA, 597–608. doi:10.1145...

  16. [16]

    William B Dolan, Chris Quirk, and Chris Brockett. 2004. Unsupervised construction of large paraphrase corpora: Exploiting massively parallel news sources. InCOLING 2004: Proceedings of the 20th international conference on computational linguistics. 350–356

  17. [17]

    Sivan Doveh, Shaked Perek, M Jehanzeb Mirza, Wei Lin, Amit Alfassy, Assaf Arbelle, Shimon Ullman, and Leonid Karlinsky. 2024. Towards multimodal in-context learning for vision and language models. InEuropean Conference on Computer Vision. Springer, 250–267

  18. [18]

    Ge Gao, Alexey Taymanov, Eduardo Salinas, Paul Mineiro, and Dipendra Misra. 2024. Aligning llm agents by learning latent preference from user edits.Advances in neural information processing systems37 (2024), 136873–136896

  19. [19]

    Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. InProceedings of the IEEE conference on computer vision and pattern recognition. 6904–6913

  20. [20]

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al . 2024. The llama 3 herd of models.arXiv preprint arXiv:2407.21783(2024)

  21. [21]

    Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham

  22. [22]

    Brandon Huang, Chancharik Mitra, Assaf Arbelle, Leonid Karlinsky, Trevor Darrell, and Roei Herzig. 2024. Multimodal task vectors enable many-shot multimodal in-context learning.Advances in Neural Information Processing Systems37 (2024), 22124–22153

  23. [23]

    Patrick Huber, Ernie Chang, Wei Wen, Igor Fedorov, Tarek Elgamal, Hanxian Huang, Naveen Suda, Chinnadhurai Sankar, Vish Vogeti, Yanghan Wang, et al. 2025. MobileLLM-Pro Technical Report.arXiv preprint arXiv:2511.06719 (2025). Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning 23

  24. [24]

    Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019. What Does BERT Learn about the Structure of Language?. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 3651–3657

  25. [25]

    Julian Killingback, Ofer Meshi, Henry Li, Hamed Zamani, and Maryam Karimzadehgan. 2026. A Unified Model and Document Representation for On-Device Retrieval-Augmented Generation.arXiv preprint arXiv:2604.14403(2026)

  26. [26]

    Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019. Similarity of neural network representations revisited. InInternational conference on machine learning. PMlR, 3519–3529

  27. [27]

    Sneha Kudugunta, Aditya Kusupati, Tim Dettmers, Kaifeng Chen, Inderjit Dhillon, Yulia Tsvetkov, Hannaneh Hajishirzi, Sham Kakade, Ali Farhadi, and Prateek Jain. 2024. Matformer: Nested transformer for elastic inference.Advances in Neural Information Processing Systems37 (2024), 140535–140564

  28. [28]

    Jurek Leonhardt, Henrik Müller, Koustav Rudra, Megha Khosla, Abhijit Anand, and Avishek Anand. 2024. Efficient Neural Ranking Using Forward Indexes and Lightweight Encoders.ACM Trans. Inf. Syst.42, 5, Article 117 (April 2024), 34 pages. doi:10.1145/3631939

  29. [29]

    Haoran Li, Abhinav Arora, Shuohui Chen, Anchit Gupta, Sonal Gupta, and Yashar Mehdad. 2021. MTOP: A compre- hensive multilingual task-oriented semantic parsing benchmark. InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume. 2950–2962

  30. [30]

    Xiaonan Li, Kai Lv, Hang Yan, Tianyang Lin, Wei Zhu, Yuan Ni, Guotong Xie, Xiaoling Wang, and Xipeng Qiu. 2023. Unified Demonstration Retriever for In-Context Learning. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 4644–4668

  31. [31]

    Yanshu Li, Jianjiang Yang, Tian Yun, Pinyuan Feng, Jinfa Huang, and Ruixiang Tang. 2025. Taco: Enhancing multimodal in-context learning via task mapping-guided sequence configuration. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 736–763

  32. [32]

    Ji Lin, Wei-Ming Chen, Yujun Lin, Chuang Gan, Song Han, et al. 2020. Mcunet: Tiny deep learning on iot devices. Advances in neural information processing systems33 (2020), 11711–11722

  33. [33]

    Sheng-Chieh Lin and Jimmy Lin. 2023. A Dense Representation Framework for Lexical and Semantic Matching.ACM Trans. Inf. Syst.41, 4, Article 110 (April 2023), 29 pages. doi:10.1145/3582426

  34. [34]

    Xi Victoria Lin, Chenglong Wang, Luke Zettlemoyer, and Michael D Ernst. 2018. NL2Bash: A Corpus and Semantic Parser for Natural Language Interface to the Linux Operating System. InProceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018)

  35. [35]

    Chaoqiang Liu, Dan Chen, Yu Huang, Wenjing Xiao, Haifeng Liu, Yi Zhang, Huize Li, Xiaofei Liao, and Hai Jin. 2025. SeIM: In-Memory Acceleration for Approximate Nearest Neighbor Search. In2025 62nd ACM/IEEE Design Automation Conference (DAC). IEEE, 1–7

  36. [36]

    Hui Liu, Wenya Wang, Hao Sun, Chris Xing Tian, Chenqi Kong, Xin Dong, and Haoliang Li. 2025. Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context Learning. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2623–2641

  37. [37]

    Jiachang Liu, Dinghan Shen, Yizhe Zhang, William B Dolan, Lawrence Carin, and Weizhu Chen. 2022. What Makes Good In-Context Examples for GPT-3?. InProceedings of Deep Learning Inside Out (DeeLIO 2022): The 3rd Workshop on Knowledge Extraction and Integration for Deep Learning Architectures. 100–114

  38. [38]

    Yao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel, and Pontus Stenetorp. 2022. Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity. InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 8086–8098

  39. [39]

    Xiaofei Ma, Zhiguo Wang, Patrick Ng, Ramesh Nallapati, and Bing Xiang. 2019. Universal text representation from bert: An empirical study.arXiv preprint arXiv:1910.07973(2019)

  40. [40]

    Piergiulio Mannocci, Elisabetta Giannone, and Daniele Ielmini. 2023. In-Memory Principal Component Analysis by Analogue Closed-Loop Eigendecomposition.IEEE Transactions on Circuits and Systems II: Express Briefs71, 4 (2023), 1839–1843

  41. [41]

    Zeping Min and Xinshang Wang. 2025. DOCS: Quantifying weight similarity for deeper insights into large language models. InThe Thirteenth International Conference on Learning Representations

  42. [42]

    Chen Nie, Chao Jiang, Liming Xiao, Weifeng Zhang, and Zhezhi He. 2025. PICK: An SRAM-based Processing-in- Memory Accelerator for K-Nearest-Neighbor Search in Point Clouds. In2025 62nd ACM/IEEE Design Automation Conference (DAC). IEEE, 1–7

  43. [43]

    OpenClaw contributors. 2026. OpenClaw. https://github.com/openclaw/openclaw. Open-source personal AI assistant, version 2026.3.13, accessed 2026-03-15

  44. [44]

    Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5

  45. [45]

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. 2021. Learning transferable visual models from natural language supervision. InInternational conference on machine learning. PmLR, 8748–8763. 24 Luo et al

  46. [46]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 3982–3992

  47. [47]

    1995.Okapi at TREC-3

    Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995.Okapi at TREC-3. British Library Research and Development Department

  48. [48]

    Ohad Rubin, Jonathan Herzig, and Jonathan Berant. 2022. Learning To Retrieve Prompts for In-Context Learning. InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2655–2671

  49. [49]

    Dustin Schwenk, Apoorv Khandelwal, Christopher Clark, Kenneth Marino, and Roozbeh Mottaghi. 2022. A-okvqa: A benchmark for visual question answering using world knowledge. InEuropean conference on computer vision. Springer, 146–162

  50. [50]

    Yijia Shao, Tianshi Li, Weiyan Shi, Yanchen Liu, and Diyi Yang. 2024. Privacylens: Evaluating privacy norm awareness of language models in action.Advances in Neural Information Processing Systems37 (2024), 89373–89407

  51. [51]

    Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts. 2013. Recursive deep models for semantic compositionality over a sentiment treebank. InProceedings of the 2013 conference on empirical methods in natural language processing. 1631–1642

  52. [52]

    Kangkang Sun, Jun Wu, Ali Kashif Bashir, Jianhua Li, Hansong Xu, Qianqian Pan, and Yasser D Al-Otaibi. 2024. Personalized privacy-preserving distributed artificial intelligence for digital-twin-driven vehicle road cooperation. IEEE Internet of Things Journal11, 22 (2024), 35902–35916

  53. [53]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers). 4149–4158

  54. [54]

    Anton Voronov, Lena Wolf, and Max Ryabinin. 2024. Mind your format: Towards consistent evaluation of in-context learning improvements. InFindings of the Association for Computational Linguistics: ACL 2024. 6287–6310

  55. [55]

    Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018. GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding. InProceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. 353–355

  56. [56]

    Cheng Wang, Zenghui Yuan, Pan Zhou, Zichuan Xu, Ruixuan Li, and Dapeng Oliver Wu. 2023. The security and privacy of mobile-edge computing: An artificial intelligence perspective.IEEE Internet of Things Journal10, 24 (2023), 22008–22032

  57. [57]

    Jiacheng Ye, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. 2023. Compositional exemplars for in-context learning. InInternational Conference on Machine Learning. PMLR, 39818–39833

  58. [58]

    John M Zelle and Raymond J Mooney. 1996. Learning to parse database queries using inductive logic programming. In Proceedings of the national conference on artificial intelligence. 1050–1055

  59. [59]

    Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence?. InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 4791–4800

  60. [60]

    Peng-Fei Zhang, Guangdong Bai, Hongzhi Yin, and Zi Huang. 2023. Proactive Privacy-preserving Learning for Cross-modal Retrieval.ACM Trans. Inf. Syst.41, 2, Article 35 (Jan. 2023), 23 pages. doi:10.1145/3545799

  61. [61]

    Tianyi Zhang, Jonah Yi, Bowen Yao, Zhaozhuo Xu, and Anshumali Shrivastava. 2024. Nomad-attention: Efficient llm inference on cpus through multiply-add-free attention.Advances in Neural Information Processing Systems37 (2024), 112706–112730

  62. [62]

    Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.arXiv preprint arXiv:2506.05176(2025)

  63. [63]

    Haozhe Zhao, Zefan Cai, Shuzheng Si, Xiaojian Ma, Kaikai An, Liang Chen, Zixuan Liu, Sheng Wang, Wenjuan Han, and Baobao Chang. 2024. MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning. In The Twelfth International Conference on Learning Representations

  64. [64]

    Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, Jie Jiang, and Bin Cui. 2026. Retrieval-augmented generation for ai-generated content: A survey.Data Science and Engineering(2026), 1–29

  65. [65]

    Wayne Xin Zhao, Jing Liu, Ruiyang Ren, and Ji-Rong Wen. 2024. Dense Text Retrieval Based on Pretrained Language Models: A Survey.ACM Trans. Inf. Syst.42, 4, Article 89 (Feb. 2024), 60 pages. doi:10.1145/3637870

  66. [66]

    Han Zhou, Xingchen Wan, Lev Proleev, Diana Mincu, Jilin Chen, Katherine A Heller, and Subhrajit Roy. 2024. Batch Calibration: Rethinking Calibration for In-Context Learning and Prompt Engineering. InThe Twelfth International Conference on Learning Representations. Gradient-free Task-Conditioned Retrieval for On-Device In-Context Learning 25

  67. [67]

    Pushen Zuo, Qishen Wang, Yubiao Luo, Ruiqing Xie, Shiqing Wang, Zezhi Cheng, Lin Bao, Zongwei Wang, Yimao Cai, Ru Huang, et al. 2025. Precise and scalable analogue matrix equation solving using resistive random-access memory chips.Nature Electronics(2025), 1–12

  68. [2015]

    Microsoft coco captions: Data collection and evaluation server.arXiv preprint arXiv:1504.00325(2015)

  69. [2018]

    InProceedings of the IEEE conference on computer vision and pattern recognition

    Vizwiz grand challenge: Answering visual questions from blind people. InProceedings of the IEEE conference on computer vision and pattern recognition. 3608–3617