Pith. sign in

Paper Citation Record · LEDGER

Clapper: Compact Learning and Video Representation in VLMs

As of 22 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2505.15529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15529 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:49.012698Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f111aa1-f365-4e08-a008-587d8d0b50db · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan.

Clapper: Compact Learning and Video Representation in VLMs Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:20:51.732946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.435871Z digest=sha256:9c3e3492f26f66c7eff4d7f8704765b399f7a77db8de5af112240aef7a475f54

Observation 4757746a-bce9-4cbc-b8e7-4def18f07d21 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.581798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.511993Z digest=sha256:3a81d5616b20046187faea85a404da51a80b222dfb78deed36751d8909c3f38d

Observation 7e5aaeaf-d987-421e-b41d-72332f633ec6 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.407479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.586369Z digest=sha256:63191ea68694d331da0833c1683c3f658e3d12b6a50074f4fe134c40a801ddd0

Observation aeb95aec-7c5b-48a8-9143-5c9c25861d5b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Clapper: Compact Learning and Video Representation in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.667515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.667515Z digest=sha256:7fd8fb4dc866eb38191bb324cb1925aa9ad41cc7a3601ab3d01171cc8335bfea

Observation 89c9b623-186a-4330-be78-cabcf333b42b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Clapper: Compact Learning and Video Representation in VLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.764951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.764951Z digest=sha256:fd846bbe4495c58c5eabaf664fd16bfb72d2b58e26f4269f3418042e0a8ccf9b

Observation 042da8ab-4e83-4132-bc45-8a49da6f3094 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Clapper: Compact Learning and Video Representation in VLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.871031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.871031Z digest=sha256:b8f7f4a6ba7480795e473fa61f77ebd1ee6d93c8050961f20fe0ec6c0367839c

Observation 9df7712b-7efb-4210-95f3-c02909292127 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.997198Z digest=sha256:414811204975e2fcda6d309a13c93617bfdc2b34b06a5f9b431de5ecb4f6ed06

Observation 6537dfb7-4d62-42b2-8b9e-152561038d2e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Clapper: Compact Learning and Video Representation in VLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.092112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.092112Z digest=sha256:fc8ce06384fc12548e366b8002dc8660b6adceba0429f06d051d43ed41d946c7

Observation 32001848-f00c-4b73-b278-eb58c1f57253 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.191691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.191691Z digest=sha256:3dd9cace0fdb6773598e968cc3d2e4eb87ecd1ee85359f9c09fa3451d76b5327

Observation fc5946e2-dd72-48a1-a1bc-98da3fea1e84 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.283481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.283481Z digest=sha256:140415f6d4cbb2981a0a672c4fd105954d3327537e85ca31ca6fe961540b472a

Observation 568b2f86-13e7-4fb3-bedd-e721c409f31d · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.011985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:45.401124Z digest=sha256:3a903642c0a1fef045e03837483db485e1e9362d39dcab810ab59376d8307d48

Observation 6dcc8ae6-4df9-4b87-a376-5dbf66e45257 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.816531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:45.516975Z digest=sha256:217213e8ae9e329832e0056e13735808db55710e7f364556cc1527bfff59afc3

Observation 0833c6ca-2656-44aa-93b4-173aa01292d6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Clapper: Compact Learning and Video Representation in VLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.618643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.618643Z digest=sha256:68596f6a77b714db6e9e1318d5dc64ea254315327cb3458f53e27809014a7b50

Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.733144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.733144Z digest=sha256:de808a704a8efa0cf0cd5161e5adbe7ba70bfd00ba863d15248e7ea0c29faf7e

Observation 69f0435a-ef32-44d7-b5ee-a8290edf9e0f · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Clapper: Compact Learning and Video Representation in VLMs TempCompass: Do Video LLMs Really Understand Videos?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.814866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.814866Z digest=sha256:c8a792fe6369025df2d86f60dbf5b71973f6973ff9756be8ecffa20b7f7d0e54

Observation 68cd9ea3-10f6-419d-b479-3cc699219b56 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.926076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.926076Z digest=sha256:c251c095a690b709cc336f537a0875dd92b3256287063dbf2f9f4f2a64da4764

Observation 5e32be1b-96f0-47c5-a829-3e799dc63adc · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Clapper: Compact Learning and Video Representation in VLMs OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.025304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.025304Z digest=sha256:4fc02ef8ba463ad7a8487686d570ba5038e3553fc8fe7a6940b7fd9e7d2d3c22

Observation 2c42a8a7-e8b0-4b89-a580-674d9dd00723 · outbound

This paper cites GPT-4 Technical Report.

Clapper: Compact Learning and Video Representation in VLMs GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.130255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.130255Z digest=sha256:f5163f5d4d7323f176ba4d115a5152db6d45a803a5ba4d73b24c4c27075cec51

Observation 29148629-9ec8-41a5-8c18-a42e3e65cfc3 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.240845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.240845Z digest=sha256:045c252281bf5fe3603eebc25a037ae45b11a8bbc7130350f983365a4fa5dee3

Observation ccdf8af0-55ec-4984-87a4-2d1ccafd23d7 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.356809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.356809Z digest=sha256:4ca7287187fb9df50d245ad9206912063236e0c9450a470c8cc30c0007248862

Observation 11bfe300-e9c2-47e6-91a2-4e1d3cdb879a · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.438533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.438533Z digest=sha256:c91e7417e7ec00eb974e6e342c10970ba90b0de5b570ee7a5c45ec8a3f51965f

Observation 53be6b26-fba9-45f5-9341-c07df4c0678b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Clapper: Compact Learning and Video Representation in VLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.538312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.538312Z digest=sha256:31b971c66b9d2d09d9f9b1ee762fd05800fcd83ae3f93467e802dbe209268f40

Observation df10b4bf-19b9-4cc1-9cda-ea7ba31abd8d · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Clapper: Compact Learning and Video Representation in VLMs LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.656613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.656613Z digest=sha256:36246fe81ca95ad3d596725490af4d212171ea497d6836f66f4e43a67b01a339

Observation 8e97a8a8-ca31-4697-95fb-48e4305a115e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Clapper: Compact Learning and Video Representation in VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.774925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.774925Z digest=sha256:5851464b251d124efb07c160ad04c6914fe023b4619a0689d9ba130ffe10c44f

Observation 4e8687f3-97cb-4618-809d-b22cb0a964e5 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.882268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.882268Z digest=sha256:d06a93a238d65b1625754f65d1c26995553c7080ef7fd25a676525fb59dbd6cf

Observation 52373ded-3d67-44cb-b8ae-820b898f4e2f · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.489650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:46.992993Z digest=sha256:b46c3ebbd419d85d3650c05d270e46962efe19d8ea902ae4206876d1ac25a22b

Observation e80577a3-53a0-4e6a-b701-074b0ce0ca31 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.314362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:47.097954Z digest=sha256:b81a92bf5ef9f15f1d8d71584d48144f9b1ae2bb10599a6e309997f47e1a4288

Observation 6d1c5864-6482-4522-bde2-8406b87655e1 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.061845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:47.209607Z digest=sha256:c1986a3b4180de3e85e96e9ec3e2ff4aa399dd3e1747b75f17581a8b50815f77

Observation 65d26525-3b22-4b76-8c52-335ada6560d3 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.360957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.360957Z digest=sha256:0e4ad2dcbe321d0ae8cd435bcf65dfe5fe361922d7d46309dcaeb6db3c063586

Observation 4e19bcb9-4faa-4af7-8a25-9f01f45e34d9 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Clapper: Compact Learning and Video Representation in VLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.442019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.442019Z digest=sha256:41b498ac1727e4e665eb4af3b2d1a7852691f6f3aa8433c834f9db5fcd2f1e9c

Observation 0a170583-a703-4bcf-9e43-4138824d1160 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Clapper: Compact Learning and Video Representation in VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.549466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.549466Z digest=sha256:691cfaaa037106442037de6e742fac389d6513218e7f15640d0a8b4ce4e2b713

Observation 504d6f4d-54f0-4a77-8703-264b2636ec34 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Clapper: Compact Learning and Video Representation in VLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.582468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.582468Z digest=sha256:de7a69991a82478ec5423c5f3630e4bf648314fa75b51d72b86b1c0e63717c07

Observation 3b7b5785-9727-433c-8acf-043e29bc57cb · outbound

This paper cites Qwen2 Technical Report.

Clapper: Compact Learning and Video Representation in VLMs Qwen2 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.666280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.666280Z digest=sha256:84cead20abcbadb84874c8b307a30587d7443b070d9f9baf00a28ccceb2706c7

Observation dcf3998a-0299-44c7-99b1-60537273aa35 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Clapper: Compact Learning and Video Representation in VLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.766618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.766618Z digest=sha256:eb30477be49ad1911bed4fbcef4c595861458a00ff033ce46afa4aaa257d929d

Observation 6123c3b3-0e54-484e-b7dd-3352b9d48e95 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Clapper: Compact Learning and Video Representation in VLMs mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.848486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.848486Z digest=sha256:aa17c5ccc5bf3c19e71042b565d68515e82fef302d904f207de16508ad18f907

Observation b7036e20-0a50-4793-a074-fb882ae4e945 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.964092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.964092Z digest=sha256:22a1183c69242497c31236c69265e4c1a7a15a77eb888d1f8a804948a6edf613

Observation b080d67c-5c39-4dbb-8d51-116ecac20bc5 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.040187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.040187Z digest=sha256:c950dc6a0448b2239d6f37f229538b95d231c0bb6b606e370012e27a121570a2

Observation e0f45e9c-7100-445f-a5c3-ea4118e6ebbc · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.150431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.150431Z digest=sha256:04eca976394481d20035d3742f28a90f090573a49174d697330836aa96ae0955

Observation cdfe6a61-b8c6-4b66-8526-ef453138e929 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Clapper: Compact Learning and Video Representation in VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.289056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.289056Z digest=sha256:9ec9b62c88f09c34dad36fb77e06fb5cf5713822ee8eb719fd8521ae268381cd

Observation 0fec3900-d476-4a76-b440-735d2d295676 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Clapper: Compact Learning and Video Representation in VLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.439015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.439015Z digest=sha256:af51313b3d77d3204334d4e166e61cb52a62584cf8bb2957093cf757ba5f2cac

Observation 896f522a-1fd3-4c1f-9435-e1462eb735e9 · outbound

This paper cites Long Context Transfer from Language to Vision.

Clapper: Compact Learning and Video Representation in VLMs Long Context Transfer from Language to Vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.560323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.560323Z digest=sha256:bec6abbe1e46c5330bb67a4e7bf2aa2bb70084c2993556902835a20af855e8f7

Observation a2c8dda5-b1a7-4752-a01d-d624987f1540 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Clapper: Compact Learning and Video Representation in VLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.669699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.669699Z digest=sha256:4a488f9e95d2d2a545a0b8ec16d0a6a2d5c8d66f795393125543668ca730ebc7

Observation 61da8d38-d993-463d-b042-23d06d7753db · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:49.835526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T15:20:48.791387Z digest=sha256:63d45d2b9ee1794f7f164fe3c113964dcea266ef42d0d3a1f89bee83eb94a26e

Observation 8eacf7cc-13ab-4b30-9181-67af9725b7e1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.883282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.883282Z digest=sha256:9eed1f8de4d8ee5380fb5eaac5adcfda1f51e79e21a4aa859ef9fa90ceba9ad8

Observation 20c7a929-9e13-4928-860f-e2097167ac49 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Clapper: Compact Learning and Video Representation in VLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:49.012698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:49.012698Z digest=sha256:08d8e4880ec7f0364ab914cc04fd6c5bcfa670d0b286f10e7552a732b8dbe310

Pith citing papers

No inbound Pith citation observations are available.