Pith. sign in

Paper Citation Record · LEDGER

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2509.22854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22854 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:50:23.244358Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T05:32:01.059706Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T05:33:58.600956Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c25d9971-bf1e-4d93-a035-a836d7cc835e · outbound

This paper cites Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.033438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.033438Z digest=sha256:b82999dbfdfcd1342fbb836b7319830dd941609e4bdeb54810c24b6cca0756d2

Observation 9af70504-b30d-4139-af1c-afc375105dad · outbound

This paper cites More importantly, as the input length increases, the inference time of few-shot grows much faster than that of ICR.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention More importantly, as the input length increases, the inference time of few-shot grows much faster than that of ICR

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.239578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.239578Z digest=sha256:4139927c697fb7c3f4689176363893736a37961bd485ec64b78bebf9de837b10

Observation a5af7271-d133-4aca-b2d9-32e2bec09075 · outbound

This paper cites cross-dataset.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention cross-dataset

Reference 7

Resolution
malformed identifier
no resolver link, observed 2026-08-04T14:50:23.244358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.244358Z digest=sha256:ec4eb86ec1a5afcab1e7a4f4ace6235351e4403630741d13366f63eb0cae98ed

Observation 82aa3593-2484-48e3-888a-c38ca0010efa · outbound

This paper cites The Llama 3 Herd of Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.731921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.731921Z digest=sha256:8a167ebb2daa6d96e16dc1a8229c7f925c9f26a124e046e61498964e1978a285

Observation 44a17e93-b972-45ac-9a3c-0beca0d9a932 · outbound

This paper cites In-Context Learning Creates Task Vectors.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention In-Context Learning Creates Task Vectors

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.808650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.808650Z digest=sha256:c13beac9240ebea448e237b63b732cdb10cd33a8c5a7e96adaf96e41480a12f1

Observation 0a9b24fd-0d51-4178-8f14-e357930dbe82 · outbound

This paper cites Language Models Implement Simple Word2Vec-style Vector Arithmetic.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Language Models Implement Simple Word2Vec-style Vector Arithmetic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.166106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.166106Z digest=sha256:202cf1e481ada8824c8bc2334e7c87e1e2681abd8ca17ea4cd91481c318ce8b1

Observation c3bce674-6557-48fd-b743-f82b92e071a9 · outbound

This paper cites MetaICL: Learning to Learn In Context.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention MetaICL: Learning to Learn In Context

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.170294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.170294Z digest=sha256:444e410c84a73fc9858f385a108e696a427b7863338a86f70dfd2ff19e4465d3

Observation f400c71f-1e22-4677-8f4e-747d308f8f76 · outbound

This paper cites In-context Learning and Induction Heads.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention In-context Learning and Induction Heads

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.174869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.174869Z digest=sha256:1ca808a540feaee567fecff12f70d81a55e6cd636f46e068e2af399dc22d850d

Observation 4a7fe9d9-a096-4e72-9fd2-9f96448617ae · outbound

This paper cites doi: 10.3115/1219840.1219855.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention doi: 10.3115/1219840.1219855

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.184071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.184071Z digest=sha256:5676fc05e766d0debbd3d7cc69462029b06369f30283e0bc9a3286e6073853d5

Observation 994a7e63-778f-4647-b355-323f51efd39b · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.192290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.192290Z digest=sha256:ad98a8e29b0a039414eef5b2b9ebcbdaf628c667a4aa331635b3a9b132f12344

Observation 05319dbf-edd4-4a00-a543-cd34d99d4e15 · outbound

This paper cites ELICIT: LLM Augmentation via External In-Context Capability.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention ELICIT: LLM Augmentation via External In-Context Capability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.196943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.196943Z digest=sha256:ca06e713e7cdf91cc0fdd5253cca57dfacaa82791b48d872d48e003885d21904

Observation bc499890-326d-4133-bc1a-28337bec954b · outbound

This paper cites Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.201430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.201430Z digest=sha256:7e372efc651819edfdcaaef1b4cad43cb40e576fd7f46db11bfd7709eea560ab

Observation d09bab47-963e-40f8-9968-d96c6e3338be · outbound

This paper cites Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.206160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.206160Z digest=sha256:d59e2528e1d8263d2015b4f31c8d58120cdbefe53b93cd165fbd7e37b9da8593

Observation 4cd13016-bd0e-4f67-9b63-926595056edc · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.211091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.211091Z digest=sha256:bb6289601f04e71b63266ebff4f7cf9ffcf6cee7c8a163b4be6eb22a58291e38

Observation 6d911264-1561-445f-a98e-36f6d29df409 · outbound

This paper cites Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.215388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.215388Z digest=sha256:6c187ba3551de53654280c7676af389f210b7063ffe8b0d745d2558c19f44dad

Observation 9d9ffb5d-df01-4b9c-b496-f5e54429b280 · outbound

This paper cites Qwen2.5 Technical Report.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Qwen2.5 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.219645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.219645Z digest=sha256:bb3c984730601f2f3971c6ddddfda0385b06021f97f0be435f5b031d4d782442

Observation d4b2d54c-1f05-4460-88aa-b56786235d86 · outbound

This paper cites an unresolved cited work.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.224062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.224062Z digest=sha256:9918eed455ad9c7624c4315a217d1bf4bb04e0eaa67de324263f0ffd40fc35ee

Observation b7157842-139f-4c28-8750-e890e6040551 · outbound

This paper cites An identical argument applies toU k.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention An identical argument applies toU k

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.229151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.229151Z digest=sha256:0af8d79d4783f6c7afd38b98586f831aefc72221fa415f14336b28665830fbfd

Observation ca6c0cdf-c5b0-45e4-9c7a-f855583512b4 · outbound

This paper cites For training, we use the same number of few-shot examples as those contained in an ICL prompt during the construction of ICL bases, drawn from five in-domain datasets.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention For training, we use the same number of few-shot examples as those contained in an ICL prompt during the construction of ICL bases, drawn from five in-domain datasets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.234829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.234829Z digest=sha256:8b6de3c9141bca802bbf5dbb4060f33486b38bda0cb7ff28dbab45ed34bd608c

Observation 783df9d7-7fa4-4631-b747-30e336762a43 · outbound

This paper cites URLhttps://doi.org/10.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention URLhttps://doi.org/10

Reference 1970

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.552599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.552599Z digest=sha256:76d37ab6a34d924d427221eb54a361e741af358cceb113141a04444235ab0531

Observation add3b2d0-56e7-4ae6-9388-8f9344101000 · outbound

This paper cites Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Is attention required for ICL? Exploring the Relationship Between Model Architecture and In-Context Learning Ability

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.031609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.031609Z digest=sha256:72e05f71eb6ba05b4403a2244a280e746ffac12a670e922ed9e4b0c106f8ff2f

Observation cba2eb2f-9aa9-493e-b8e0-c66fbeda1485 · outbound

This paper cites M2iv: Towards efficient and fine-grained multimodal in-context learning via representation engineering.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention M2iv: Towards efficient and fine-grained multimodal in-context learning via representation engineering

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.156936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.156936Z digest=sha256:1dc0ef95629262c59a7840b2be76fbb35f15cd7000c22fd131e0b1117eee5f06

Observation 32e6b99f-c986-4658-a64f-414c0e48b27f · outbound

This paper cites A Survey on In-context Learning.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention A Survey on In-context Learning

Reference 2005

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.627647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.627647Z digest=sha256:afc60b96526eef669f971534619bf1163254e87b52f6bf12e48370328f1c381c

Observation ba35a10a-481a-42e2-9508-fdf5e60b80a3 · outbound

This paper cites Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.436932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.436932Z digest=sha256:ec2fde018ae8f0634fc76605d5ed0e6bcee8e28d8a9677d70a9d23a17de6cbeb

Observation 16181b88-35f6-415c-9a9e-ef0583fae94f · outbound

This paper cites Function Vectors in Large Language Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Function Vectors in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.187923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.187923Z digest=sha256:97d2a83660fbe52a4c34992b4f875de43ffd9998eb27e6ca301b58f259c25771

Observation 3b1a1d2a-a493-4ba9-9c45-de2bea18e80b · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:21.969963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:21.969963Z digest=sha256:aabb74ee3d32a4a79ab23bb5f8b3f73b05b187c4f65035e6af19f054917ad3b6

Observation 52ee78c3-2997-409b-9924-78f6d5982f6c · outbound

This paper cites In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.161687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.161687Z digest=sha256:6e916f8dfb3483ac085f634e53c474e53c01f37c6e32d880806ad98b0c3e0163

Observation 092222b2-7b0b-421f-86fc-33828e0c044d · outbound

This paper cites CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention CREAK: A Dataset for Commonsense Reasoning over Entity Knowledge

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:23.179507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:23.179507Z digest=sha256:65cf595cba5dcd9ee0c915842a6c61dc351a8904588866ae13d7d99109f6ea9b

Observation a5a7667d-935e-4e89-9372-e419cd2f311a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention LoRA: Low-Rank Adaptation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.913368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.913368Z digest=sha256:010d35fab04ae88eda743f9ca5fa564f65c4e6dd117ab581f5ca44420a5ef93a

Observation c7acb8e5-97fb-4cb9-93c2-2060b522dffa · outbound

This paper cites On the Relation between Sensitivity and Accuracy in In-context Learning.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention On the Relation between Sensitivity and Accuracy in In-context Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.163974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.163974Z digest=sha256:8a94b0e793c802e0f33e7a10e8c51505bb78bb15cb488b81f5b5c14ee3ead473

Observation 096fbc21-f5ac-4cd6-bbd0-3a79b1fb0c85 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.311948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.311948Z digest=sha256:3694ab45753741af1a23e7f44592bdba27bc2940566ce09e8d775b460455444d

Pith citing papers

Observation 966e88eb-be64-4ced-9e10-ee70b5369dd1 · inbound

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning cites this paper.

Distributional Alignment as a Criterion for Designing Task Vectors in In-Context Learning Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-03T13:05:38.627651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T05:32:01.059706Z digest=sha256:ddede50b6955794fda71d1f95f64c089f41e3f112f6f04571ca01a4e3381f49b