Pith. sign in

Paper Citation Record · LEDGER

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 7 inbound Pith citation observations for arXiv:2603.01455.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.01455 v3

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T18:09:59.236030Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T15:27:46.664630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T10:58:03.480373Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact21
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cc156a03-de17-4bd0-8e80-bba526553996 · outbound

This paper cites GPT-4 Technical Report.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.155913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:f12083dda276f09a14b555fac5db1a6f92d14b8af21fc6b5122911c11e7b80b7

Observation 664a6953-e737-4c58-a9e9-b3faf37c7571 · outbound

This paper cites Qwen2.5-VL Technical Report.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.138587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:7fbf7a52f4ea1d615849ab5b94bb64ab4536e63edb2333bfab1bcd0af084583b

Observation de650f8b-18bc-46f8-bf4b-94dfd41637c4 · outbound

This paper cites Videominer: Iteratively grounding key frames of hour-long videos via tree- based group relative policy optimization.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Videominer: Iteratively grounding key frames of hour-long videos via tree- based group relative policy optimization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.495780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:68ebe6239e7b1ccfe4a0cbdc7af5566085231dd1d7c754aa4813797e7e08d003

Observation 2a28a8ff-da50-4c82-87a3-e1e5b4e4fcf7 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.150161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:497f5f01f7e0a13c9dae00dd6ef2936200ca5a4f1c40b88e6a56dc8e68880912

Observation 24ec224f-bd4c-47f2-b2fd-670a6ac9458a · outbound

This paper cites Towards large language models with human-like episodic memory.Trends in Cognitive Sciences.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Towards large language models with human-like episodic memory.Trends in Cognitive Sciences

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.512146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:70a52359f3c08552810815d891f52f20c8f0c559f69ffb30a57f42014dcf189c

Observation 65645373-9018-4f07-9609-564a85ed5cf2 · outbound

This paper cites Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.503934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:82e2ac21bc8e2fe7facbafef153619e6cfb18bb0dfdbb0a19531bedd531a03cc

Observation c9b382f0-5f33-433b-b3b0-7ed2fe5477cf · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.955297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:d59556b3e9385594244b7eb187f39f6215f36fc7ba953cc36c4b8f8988a849c8

Observation 8fbf239e-910b-4428-b273-4da57db7d6ba · outbound

This paper cites Inducing high energy-latencyoflargevision-languagemodelswith verbose images.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Inducing high energy-latencyoflargevision-languagemodelswith verbose images

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.499850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:ead3320f847206f94fb2edc283740498f4a03e8013bed998c398b23697c99af6

Observation 5e5a71de-842e-4006-8d89-b03d77d2a21a · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video under- standing.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Ma-lmm: Memory-augmented large multimodal model for long-term video under- standing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.516297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:62b7632b5d4c344b45f09e4d5beef054d1070e6d69fa5636de70c3d5107e1d75

Observation 1212c5af-ecda-4b89-bcd8-ef90209d5865 · outbound

This paper cites View from the top: Hierarchies and reverse hierarchies in the visual system.Neuron, 36(5):791–804.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents View from the top: Hierarchies and reverse hierarchies in the visual system.Neuron, 36(5):791–804

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.507897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:6ba36f907a651285f464499d61d82fa9c86a3a7fc758d1519c957ae2e4824433

Observation e2605757-1e64-4e5c-a073-994223232c67 · outbound

This paper cites Lightweight and cognitive agentic memory for efficient long- term interaction.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Lightweight and cognitive agentic memory for efficient long- term interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.113991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:66e4c84d53fb97169e92f26c5505952b342e83ce5bfd5d0e953f7f84c5341d33

Observation 24f5bf25-7579-4845-9bfb-692322018610 · outbound

This paper cites GPT-4o System Card.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents GPT-4o System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.071050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:4520814a984310969a0b139adc8a37f034997e2b681e6f1a15cce772d9aa3f35

Observation f92fd66a-635c-463a-acec-78abb64e8fa4 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Chat-univi: Unified visual representation empowers large language models with image and video understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.418998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:87903c7ed4592424ecca2a2bff4cbec70131cea237be9fd17641f30b40297aec

Observation 9c12cb24-b5e0-4cff-88cd-342167ecb798 · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.486822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:4263f58e49b7dafa361cb96f086057e1045029d0575403530ea173872b090e35

Observation 247fea39-5a72-4225-bbb3-d20fd3775268 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large lan- guage models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.504286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:0dde10243b614233ef340755e2236fc56342df37ea47d38a5308e3f0cdb11afd

Observation 39938b33-4bf6-4e1a-8b85-7a39d71390ab · outbound

This paper cites Llama- vid: An image is worth 2 tokens in large language models.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Llama- vid: An image is worth 2 tokens in large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.499076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:5e833e91df21954e1d4ff04bb148fa6cd0e434decb90a81e8d8e97c66d48e4fa

Observation 8d334c04-4072-44ba-a6e6-53b6f39117b7 · outbound

This paper cites Video-llava: Learning united visual representation by alignment before projec- tion.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Video-llava: Learning united visual representation by alignment before projec- tion

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.515404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:f33e40168b1eac4a672e1920a0446752f0c73fc410e51903ad0fe8ed7d155222

Observation a4ec8e0f-80e4-48f5-aefc-54e53d20ed84 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892– 34916.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Visual instruction tuning.Advances in neural information processing systems, 36:34892– 34916

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.446404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:acac5455197bfc28b1f2d4aa036102919e49e57f6db978956709f54c016d5c91

Observation 72c29613-599b-4000-a140-8b60cfd7b9eb · outbound

This paper cites Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Seeing, listening, remembering, and reasoning: A multi- modal agent with long-term memory

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.084528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:78e38a6fab114f1ff4f7dde621c361f2dbdba52b61dc020f9ae3963a2df6a6fc

Observation f0a909e8-7e1e-47a4-bdf8-799ecaa1a65a · outbound

This paper cites arXiv preprint arXiv:2411.13093 , year=.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents arXiv preprint arXiv:2411.13093 , year=

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.146835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:aed0d271db0c36c7a233d7648000a78c941cd7dede842c7d7616e8b051f0251a

Observation a3d7dae4-f278-46bc-ad08-e2d3849bc656 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.461893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:a9670c3fdb0b1ec122175a7129b445b73fe1ba50e0afebd605397577597bbb87

Observation a97c1720-f839-40f2-924c-2f4121c7565e · outbound

This paper cites M3-embedding: Multi-linguality, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents M3-embedding: Multi-linguality, multi-functionality, multi-granularity text em- beddings through self-knowledge distillation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.477074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:2d3f686bc43bfaadb401ad2ed88a764d4cfd931589e366f268ef43399e611f47

Observation 771debdc-c122-4e95-a4a7-a4e8c4991b56 · outbound

This paper cites Memgpt: Towards llms as operating systems.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Memgpt: Towards llms as operating systems

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.509851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:15ca79b0851728cdee3520f2fde62c0d214658820f4865ee59c34a290a857289

Observation 49364b86-735e-4834-baf8-9e4ec88862df · outbound

This paper cites Hd-epic: A highly-detailed egocen- tric video dataset.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Hd-epic: A highly-detailed egocen- tric video dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.451276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:858d29a9bcc50cbc0a015bfc3709c2f1fe28f0bc3c397819411760737e971ded

Observation 3bd30bc5-7c1e-4930-b98d-59144df894e2 · outbound

This paper cites Streaming long video understanding with large lan- guage models.Advances in Neural Information Pro- cessing Systems, 37:119336–119360.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Streaming long video understanding with large lan- guage models.Advances in Neural Information Pro- cessing Systems, 37:119336–119360

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.461129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:ea5b72cf2a3d7adcc21541e3495aaec7f279f435837a4f7604b2985729ea8713

Observation 6a39f7bc-7546-4152-80e0-510cde6ab1da · outbound

This paper cites Fuzzy- trace theory: An interim synthesis.Learning and Individual Differences, 7(1):1–75.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Fuzzy- trace theory: An interim synthesis.Learning and Individual Differences, 7(1):1–75

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.481345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:a49656eed69ec47f3cdc5771b141e8b7eccc31879a7f9a402d97a8a8275ab4a7

Observation 6935991b-0208-474b-9d5b-a6bcf70e3bbe · outbound

This paper cites Vgent: Graph-based retrieval- reasoning-augmented generation for long video understanding.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Vgent: Graph-based retrieval- reasoning-augmented generation for long video understanding

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.140006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:a3b933cf565c26c149c3ef402457e9616424bff7fe1896662acffde3053ee9a3

Observation 0148f92c-dc6a-4ec8-9ff4-e9345b7b3c4c · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Moviechat: From dense token to sparse memory for long video understanding

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.491010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:05fdbd827c73338c7be0ba07bf9f8e108d8877d33406fd366f3570ffbd8492d9

Observation b3b003ad-ebba-447e-ab18-7af2c0aad524 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.087757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:bf5b9b653d2c0b774842c930134d2b5e59e54a2835b083870d6c0dbe23fa8c3c

Observation cc9c67ac-c6d2-42bf-bf95-e3816e82a5b4 · outbound

This paper cites The information bottleneck method.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents The information bottleneck method

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.125526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:25d2e0cea01d25c888f2e15731cda87575cd5fbfa7929160af8827d4669cd166

Observation 1f8f433f-4812-4bc4-8563-c4b538c5550d · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.132631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:a470d3b1626fd97f4c80d8a677736bbe5d9848ea88a106ecbc15bd6d2aa85ade

Observation 5e6b9451-beb0-4ed6-8a1f-14600b7a3fea · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.074747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:35fdb3965024ac8cff2fe976f2aa8688fb1e38ab6a83501e806ed1a54f5a0a4f

Observation 3d023001-1c3d-4ded-b914-42f940f2a27b · outbound

This paper cites Videoagent: Long-form video under- standing with large language model as agent.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Videoagent: Long-form video under- standing with large language model as agent

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.412290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:90acd1c07cb5cfff4c10f1fda3f343e94ceb235d813524237ae85ace5caf8cec

Observation 849b962e-fa03-47a8-a029-0688a4c5f44b · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.094597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:cfe46e9f424bcc799ed43b8d6fbd96fa30d33819d55f9889a7455dd4546cca9d

Observation 7cdf5f1f-6e6f-439f-a89b-abbc27e4c6a9 · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.465839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:b60048b81e4cced1b6a960f534d30c625afd3d1570cd490cc031908954022f5c

Observation 6530a7df-536e-4c8b-a1f4-e3f07abfd18a · outbound

This paper cites C-pack: Packed resources for general chinese embeddings.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents C-pack: Packed resources for general chinese embeddings

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.451710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:55e26528b177104a5fe79ee8a0aff46e2b6060aaa004784348910b2c110739c1

Observation 2a99e6f8-c524-448c-96a2-ce1119528069 · outbound

This paper cites Large Multimodal Agents: A Survey.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Large Multimodal Agents: A Survey

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.100889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:0fb565eaecad6ac88219c6e4a8a2773abac67151484ad05a9f4b4e4c2abf9abd

Observation 1ba2968c-dcaa-4206-a256-cb8e7eee4cc9 · outbound

This paper cites A-MEM: Agentic Memory for LLM Agents.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents A-MEM: Agentic Memory for LLM Agents

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.119772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:97479bb9c1fd539cccd78af6d43197e509a4659c615b3e48ddd83feb5aaab809

Observation 5cab4dc2-71de-45ad-b4fd-c21a3270fe6e · outbound

This paper cites Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.078382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:0c1382e9dc111b3b71349b9bdad74e959a196b7e43577c21079a76c866231b5e

Observation 172309e8-90f8-42e5-8eea-cb96806110f7 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.120845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:53f70f66d3eed613cde80e5ea3309d6d2619245878077fbe4ed392a2acf507f8

Observation 5bc8b24f-ac61-4f4b-a48e-e571ca4389d6 · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.107496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:8c17a73900bfe6bf5d87295c0c18493839f97ffbf82e6a03c8f1731e887f8238

Observation 1862eb9f-6e6d-4eac-9e5b-43713013f906 · outbound

This paper cites Long Context Transfer from Language to Vision.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Long Context Transfer from Language to Vision

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.040656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:798c06962bf9ba7130e30f16e6a706021eecb5af7fca8868a3ada491553d88a7

Observation fc7ef0be-1d40-4e91-b4a3-82756e7975b4 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-15T18:10:13.114723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:b5acb9f5ab7f95e7960ee9cc79a050a30a26ba6f3e04409787967e0e5a471e7c

Observation 0abd9925-a65c-4c13-b7d5-788292e8eb20 · outbound

This paper cites Memorybank: Enhancing large language models with long-term memory.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Memorybank: Enhancing large language models with long-term memory

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.424386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:97f9ea0fd1288bd88c6fb3fe4cc7be8d1c5d9226aaee6f3353933d4f26a9bc9a

Observation 3435fd34-5232-4c8e-ad6e-a0f0247dce24 · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Mlvu: Benchmarking multi-task long video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.398127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:4739808d615d197c040fbe88ca57a77a001d8c1f4aa35f9132f2b1e6af48ff14

Observation 1f98cdf6-5093-482e-9e24-ce409a1251b8 · outbound

This paper cites Select the best answer to the following multiple-choice question based on the video.\n.

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents Select the best answer to the following multiple-choice question based on the video.\n

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T18:10:13.402824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T18:09:59.236030Z digest=sha256:ea2ebbe685f8bacc942544b227076e49f5aeef3215a0cc7533e48d9ac34f97a4

Pith citing papers

Observation 3b0736f9-0c87-4c2f-8e15-64cb654a8117 · inbound

Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement cites this paper.

Audio-Oscar: A Multi-Agent System for Complex Audio Scene Generation, Orchestration, and Refinement From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T20:17:21.356027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T20:48:31.993369Z digest=sha256:b7a376e884c0654920e719801295d116dc2f5b165c8319ffd9193a0f1a2786c6

Observation ecb2deb6-c7e5-4163-9018-55b689de3521 · inbound

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning cites this paper.

MODF-SIR: A Multi-agent Omni-modal Distilled Framework for Social Intelligence Reasoning From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:58:03.481588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:43:56.299302Z digest=sha256:e0f36c900c26d589bef2a5d29190093bed9c63d5ee31fe826aee03cdcd6f9dff

Observation 2f4d6e20-97a5-4bdd-a9ae-d16db14bc916 · inbound

Xiaomi-GUI-0 Technical Report cites this paper.

Xiaomi-GUI-0 Technical Report From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:55:40.349060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-01T06:07:48.137351Z digest=sha256:707fcf27749ccebac226f268846b2035a127d658804e86634550d9d03a69881a

Observation cfa67b57-089a-469f-9130-3974647e9a7b · inbound

Xiaomi-GUI-0 Technical Report cites this paper.

Xiaomi-GUI-0 Technical Report From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:18.999836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-02T19:37:49.661596Z digest=sha256:821cc92e97dd288c86f4d7a6efabb06c538b968fa84e00d224e672a28a422c82

Observation baf01ff1-873c-48ed-839b-1b9b7c85d152 · inbound

UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents cites this paper.

UI-MOPD: Multi-Platform On-Policy Distillation for Unified GUI Agents From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:37.302399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:37.302399Z digest=sha256:032df23065e8320b9fdd663ed03a3525555475afb786d309ed6eba36925fa088

Observation 17f89242-abae-4e70-a51f-d88a8812ddd7 · inbound

FOLIO: Focused Semantic Memory for Streaming Video Understanding cites this paper.

FOLIO: Focused Semantic Memory for Streaming Video Understanding From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:40:44.305707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:40:44.305707Z digest=sha256:e17bf40f9a5f90fcff3e335a9771145ef44ab0dbd372940fb6f16bd8a9b8d866

Observation 1f54c517-b46d-4551-8608-2ad739758dec · inbound

DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding cites this paper.

DocMemo: Dynamic Evidence Discovery via Probabilistic Memory-Guided Retrieval for Multi-Modal Document Understanding From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T15:27:46.664630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:27:46.664630Z digest=sha256:d07b654924d79609cde40520b86690fcbeeef143f151a134c9f831164c987d48