Pith. sign in

Paper Citation Record · LEDGER

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2502.02406.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02406 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:23:40.585910Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23cbddc7-6441-4224-8f25-fc14a084dbc9 · outbound

This paper cites write newline.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.403541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.403541Z digest=sha256:883653add35cfd5b053e77b312186339cd4be71548da6386f63f3fd5fc173ea0

Observation e68b4704-9172-4ff5-a8f0-b34c1ac43dd8 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.294271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.409366Z digest=sha256:ec31943983c8fbed42e6112fc4705d07727bd48c724a18e8fa737ebf341e48ce

Observation b4c60257-454e-4756-866d-3a71467c3e66 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.414321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.414321Z digest=sha256:de5aefec1a167688b69777e315407cc5cbaec4876c03c420153b4a542c612bd4

Observation aebbd7d9-714b-4fd8-9d13-9663f9ba0898 · outbound

This paper cites Longformer: The Long-Document Transformer.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.420888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.420888Z digest=sha256:d05e97477d79128206285e1952819b91b1d43f30dac44b9e55049a3c2bffe5b7

Observation 5318c903-e261-442d-ae39-3add2963c798 · outbound

This paper cites an unresolved cited work.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:23:41.283500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.425814Z digest=sha256:5d0b117d45cbd0fb3899ef5068bde993af8726ea14f3bc7b5312f8dbaed9ccfd

Observation 441df307-bd25-4c85-acb1-14b589bc797c · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Striped Attention: Faster Ring Attention for Causal Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.430129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.430129Z digest=sha256:7aef090ff2c70eb2a8c8424ac4e4e7e72dfc61dba5c9f49c53cee656696a4bfd

Observation 2e4acc3f-d938-4512-af51-bd0dea97e9e4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.434769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.434769Z digest=sha256:96cda66c737e3a318ec5f0c89ec1d5640a4e9c7c283d6c4c518f69e77954c365

Observation 9df26055-1cd8-4052-94ac-2ce8b2a4c3f7 · outbound

This paper cites Adapting language models to compress contexts.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Adapting language models to compress contexts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.266075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.439586Z digest=sha256:1b14b3c3d7aa398729a253dfc10eaa51954d15f06d29b3e473d85291c8950848

Observation 7c55aa77-e64d-4ed6-afcb-cd2196b2a9dd · outbound

This paper cites M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models M., Likhosherstov, V., Dohan, D., Song, X., Gane, A., Sarlos, T., Hawkins, P., Davis, J

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.443931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.443931Z digest=sha256:506ccf0000cf7eb329b779bcf52b819b1b3716c5a5852be2814890702b0e7c46

Observation be44c55a-50ac-41c7-b535-6ee346f5d340 · outbound

This paper cites Nvlm: Open frontier-class multimodal llms.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Nvlm: Open frontier-class multimodal llms

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.248285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.448162Z digest=sha256:d7d6acec2171c24871ec6e6a7a0005e4db645c9de3977e62f3c35d58d25e29a7

Observation 55554418-4064-4083-ad50-9ffaf6b3814e · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.452814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.452814Z digest=sha256:2d2f8b18da7aa5c6fcf9404e93225434964ad40b638ed7c31a6eb8782ae19cbb

Observation cb45b3b9-ddf3-4fa4-8e70-4bd743cb4629 · outbound

This paper cites Y., Ermon, S., Rudra, A., and R \'e , C.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Y., Ermon, S., Rudra, A., and R \'e , C

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.456933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.456933Z digest=sha256:e1f5abda59e180c9fccc87ee736bd4b9a0eff38f7e944af583787b669644f73f

Observation 94261df9-98b0-48d8-bc7a-1e5e3ce7da21 · outbound

This paper cites S., Monga, R., Chen, K., Devin, M., Le, Q.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models S., Monga, R., Chen, K., Devin, M., Le, Q

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.224522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.460640Z digest=sha256:2fe451b05e993ee89cb5909e0246d57ab424bc2187534be19050dfeac512a794

Observation aedfadff-b42b-4e90-9a3d-fa558079e64d · outbound

This paper cites LongNet: Scaling Transformers to 1,000,000,000 Tokens.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models LongNet: Scaling Transformers to 1,000,000,000 Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.464868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.464868Z digest=sha256:be500256276fd1a37c4f1a7823db9e55684b58d508d3d25395bbdf803caa7920

Observation d60cc123-5f52-4e57-800f-b4b14048891c · outbound

This paper cites The design and operation of CloudLab.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models The design and operation of CloudLab

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.213838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.469166Z digest=sha256:24a5cadc952db0849ac15211ecf036384ca6bb4a005cc792f1fbda24987019fe

Observation 3224933a-38ff-426d-b9eb-2025c44d3273 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.472983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.472983Z digest=sha256:538a5488d7675a6e125f6f13a17be36df19d1b13323a8470f5ece6f6514735a8

Observation c7b5788d-0ecb-4cae-808c-301145995424 · outbound

This paper cites The Llama 3 Herd of Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models The Llama 3 Herd of Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.477312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.477312Z digest=sha256:e02eaaf06f278ac284c206ff2a37621a0e91194738577494fcc6206425ed70f1

Observation eeeeecb0-6ee5-478b-bb91-1eac35ee2559 · outbound

This paper cites Llava-uhd: An lmm perceiving any aspect ratio and high-resolution images.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Llava-uhd: An lmm perceiving any aspect ratio and high-resolution images

Reference 18

Resolution
verified exact
doi, observed 2026-08-09T12:23:40.634688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.481945Z digest=sha256:caa4c317feddea7c37ee7b796035d526a0d7ecc30a2fd6729c5134e8d95b20a5

Observation 1d2b86f3-3ce3-49ee-bac1-a38a8925cbb3 · outbound

This paper cites K., Jia, M., Cao, X., Shah, A., Shrivastava, A., and Lim, S.-N.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models K., Jia, M., Cao, X., Shah, A., Shrivastava, A., and Lim, S.-N

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.203438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.485838Z digest=sha256:8a7ccdf0deddeb8ad76f0e7eca8a05e8942feeae9efde6945be9db024304151b

Observation e1ddd6bc-bfa5-4fc8-9098-56725910e111 · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.491200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.491200Z digest=sha256:568b660db06ad546e09319a48cdaca6709a8d8f06ec2316809b332dd0aad371e

Observation 63d306a6-146e-47f2-b971-f1a38d53bb57 · outbound

This paper cites A., Tanaka, M., Zhang, C., Zhang, M., Aminadabi, R.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models A., Tanaka, M., Zhang, C., Zhang, M., Aminadabi, R

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.495082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.495082Z digest=sha256:faf512c7d082106e610a01ae5b01a1baf5c66103e6720308b7b795455cbc6086

Observation 9ee31d37-2d85-4086-9cd2-7a58eb10fbae · outbound

This paper cites Reformer: The efficient transformer.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Reformer: The efficient transformer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.498688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.498688Z digest=sha256:66941259479741c566a429650657314820dc970ea9f8ba844287e80ceb58edc2

Observation dd60d2e3-059a-452a-a70d-ae25d0daed60 · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.186213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.502235Z digest=sha256:690e07b2a9a723571070e07f2381f586a6d901409d368052d2b148afa2789c9b

Observation 7f6189eb-792e-44ad-a149-da35816a6ff4 · outbound

This paper cites A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models A., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.175383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.505912Z digest=sha256:08bc72491264c6a3190ca8cbde10be2dd7867249c799a109e4107cb400568e4c

Observation 83f57cc7-d32f-477a-9dda-d00ed98de030 · outbound

This paper cites M., Kiela, D., Cord, M., and Sanh, V.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models M., Kiela, D., Cord, M., and Sanh, V

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.164568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.509643Z digest=sha256:f3d62f6a60a20318a75c3c544f886b82f97794b23951396e14f77addb2f6e81d

Observation 2460e5a7-2fbb-46b3-a919-30a1280278e6 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.513418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.513418Z digest=sha256:db7db6233e6f4a2936a711a2030a80784be7481b2cb384d5c71b2703266ebf17

Observation d60026c8-84b5-428f-9b58-d1c9ec52e47d · outbound

This paper cites P., Ma, X., Stoica, I., Gonzalez, J.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models P., Ma, X., Stoica, I., Gonzalez, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.153717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.517456Z digest=sha256:74578a6d4f0506343df0f9cc5ca3a67c5a4d436b7eae6996b21e463c9a78f475

Observation be1f82a0-b4b0-4d61-bfd9-1e27e0cb9e6d · outbound

This paper cites Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Blip-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.142932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.521288Z digest=sha256:391775219719e5bc85788a78182463948ce54525c16caf967dfa797709112521

Observation fd1e80ac-e63a-4310-a9a9-39bf3e74e556 · outbound

This paper cites an unresolved cited work.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:23:41.132089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.525030Z digest=sha256:fad9d16e6c4e9e1bd3b59664f580abde9e201104b3733da5075f05e9a888f637

Observation d7881f19-2f4d-4c98-a88c-a836db4bedd2 · outbound

This paper cites Ringattention with blockwise transformers for near-infinite context.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Ringattention with blockwise transformers for near-infinite context

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.121090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.528809Z digest=sha256:b0f03f58d71c155ec10e7d9ceb9fb774abad5876115f72baa6e409a15e210e60

Observation e6297378-2b6d-428a-b06c-488d05e77eee · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.532423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.532423Z digest=sha256:0dfd0581c721afbdfd60aff72790abf01f68d813e67f4ebbc0a25021c4f4e983

Observation a8989cd5-06b8-40c3-b946-9947e56522ca · outbound

This paper cites Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.536435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.536435Z digest=sha256:f90db7d9c3c9a69e47f5483e18f4aa07927fa94c90ca0c7d58b1f2372adca346

Observation b5554898-8713-4103-9a92-9742951063dc · outbound

This paper cites R., Ganger, G.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models R., Ganger, G

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.540566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.540566Z digest=sha256:4f77a5ee98b9cbacce924610535a8e05c0b968f5409b22c171d2e7e6b365a60e

Observation 74ddb37f-4cc6-4e0e-9864-635d472ea920 · outbound

This paper cites Momentor: advancing video large language model with fine-grained temporal reasoning.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Momentor: advancing video large language model with fine-grained temporal reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.110012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.544509Z digest=sha256:0c3d2272044332f6a8d71ab1db5ad11a554fcda0a53a938517be2b0594271d85

Observation db78e0b2-276a-4709-99f1-bd534e6608c5 · outbound

This paper cites ModServe : Scalable and resource-efficient large multimodal model serving, 2025.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models ModServe : Scalable and resource-efficient large multimodal model serving, 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.548348Z digest=sha256:44c50a0023e9617d3863eb0a2b38e573828701c7acc105d3904ad07eef3c8e71

Observation 06db48e3-66b4-4590-8a0d-9129585b9bbc · outbound

This paper cites Zero: memory optimizations toward training trillion parameter models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Zero: memory optimizations toward training trillion parameter models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.098690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.552540Z digest=sha256:a94803034e046ffa0c6eda77fc39d466ee8f983364dff4ae0c304760ff5eb582

Observation 379d8cdf-8b39-4175-8596-df4581f53c59 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.556139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.556139Z digest=sha256:2460187231ff9ad3cf2ba1b1b1b1f2cfd2b893c310217659f24e32d31f659beb

Observation ba04c592-e6b8-4068-9537-dc810042bf83 · outbound

This paper cites Repository-level prompt generation for large language models of code.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Repository-level prompt generation for large language models of code

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.087584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.560141Z digest=sha256:4b4d468209a3b2b11c1b539c9006b2855cd5f463824c61b57a61b16f85f91cf3

Observation 6895e989-0ca5-4de7-a58c-b953c8279411 · outbound

This paper cites PEARL : Prompting large language models to plan and execute actions over long documents.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models PEARL : Prompting large language models to plan and execute actions over long documents

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.075061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.563745Z digest=sha256:50165d5c3d43fca50643801aa44e89441b013dc85717359db48a2dfcdd2ecd41

Observation a06ddf6a-335b-4569-b1f9-33d6e3a8227d · outbound

This paper cites T., and Cox, D.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models T., and Cox, D

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.567345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.567345Z digest=sha256:add6acee7e07f09c9acbb06f7fb789f40a052364375d71daf2255b3eef70d03f

Observation 835f1694-a240-4085-bdad-6310f2cb8dbf · outbound

This paper cites N., Kaiser, L., and Polosukhin, I.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models N., Kaiser, L., and Polosukhin, I

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.570913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.570913Z digest=sha256:3958de6d6883ceee45b1e33695f45070ff51ad9a2ecf00874ca9aef502ba8ba0

Observation c1dc54ac-ab78-4c1f-acee-58b39e7581c2 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.574541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.574541Z digest=sha256:76e3d6c118f0555c923814b53e2145f69eb4aaa3679de4951943698e840c486c

Observation 5c79bfde-886f-459b-bfcf-e69046163793 · outbound

This paper cites Big bird: transformers for longer sequences.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Big bird: transformers for longer sequences

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:23:41.057941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T12:23:40.578333Z digest=sha256:c7a22fc902877dfe7fde4311f460a44246d6ce69f8a924c3e36ba473930131c1

Observation 5bd7e58e-7aff-40c9-ac3c-fb150b07ea85 · outbound

This paper cites R epo C oder: Repository-level code completion through iterative retrieval and generation.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models R epo C oder: Repository-level code completion through iterative retrieval and generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.582047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.582047Z digest=sha256:7dc4e5ae86ec29dd243e43d202b4a7d44f6831adfbbca659258689648f126da0

Observation da7a0ffc-6d1b-4c9c-9f65-a080c283e80c · outbound

This paper cites Mini GPT -4: Enhancing vision-language understanding with advanced large language models.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models Mini GPT -4: Enhancing vision-language understanding with advanced large language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.585910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.585910Z digest=sha256:216340c3f745a07e75f9ddb7117dad207b18dea44848ccc41046bd130d1596cf

Pith citing papers

No inbound Pith citation observations are available.