Pith. sign in

Paper Citation Record · LEDGER

Clapper: Compact Learning and Video Representation in VLMs

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2505.15529.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15529 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:49.012698Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f111aa1-f365-4e08-a008-587d8d0b50db · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan.

Clapper: Compact Learning and Video Representation in VLMs Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:20:51.732946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.435871Z digest=sha256:b66f212c77207b1622deb98e3ec35c39a0c7b70d471260f8c85a17aace13c374

Observation 4757746a-bce9-4cbc-b8e7-4def18f07d21 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.581798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.511993Z digest=sha256:30a38b8074f85c3425af5e751d575d33681961785d7a2546a4c166ed0024cdbf

Observation 7e5aaeaf-d987-421e-b41d-72332f633ec6 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.407479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.586369Z digest=sha256:e38946547cf532521e9de629d2a52627cfddfccf19d99d3cd518a57d5c0875f8

Observation aeb95aec-7c5b-48a8-9143-5c9c25861d5b · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Clapper: Compact Learning and Video Representation in VLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.667515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.667515Z digest=sha256:c5a3e1eb7d8cb2b2b062cd25d2e1a600779d6ad372b18490f72079013ae49164

Observation 89c9b623-186a-4330-be78-cabcf333b42b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Clapper: Compact Learning and Video Representation in VLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.764951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.764951Z digest=sha256:7bbc5e4e8c907d7d6935a2fe88cb9088b61588e781713b5bd5d37be6148993b8

Observation 042da8ab-4e83-4132-bc45-8a49da6f3094 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Clapper: Compact Learning and Video Representation in VLMs VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:44.871031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:44.871031Z digest=sha256:ba4b8f4bee7f7403d0451202379eb86c2ff31e77604387a0e9d268f3fef96cd6

Observation 9df7712b-7efb-4210-95f3-c02909292127 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.190608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:44.997198Z digest=sha256:46855ed075393a3071ea2ae452fad0696ec8c56bfa7f5ef9a967f394ff499fc0

Observation 6537dfb7-4d62-42b2-8b9e-152561038d2e · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Clapper: Compact Learning and Video Representation in VLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.092112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.092112Z digest=sha256:00c51a3359a5b0c74c73b58132304364aaca5b193d056c4226130027995c655d

Observation 32001848-f00c-4b73-b278-eb58c1f57253 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.191691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.191691Z digest=sha256:6b06440e17193880f565162c50a6ab671622332c6210918a8662ea14df20d0e7

Observation fc5946e2-dd72-48a1-a1bc-98da3fea1e84 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.283481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.283481Z digest=sha256:8a3ac6040c1b63e85dcb9bc80b438b43c481443622f0bb420ba4cbaba4455655

Observation 568b2f86-13e7-4fb3-bedd-e721c409f31d · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:51.011985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:45.401124Z digest=sha256:2ad2a37c98cb92f4b9a21bb2bc2c9d98eac0be60849ac356870824405dad9b3c

Observation 6dcc8ae6-4df9-4b87-a376-5dbf66e45257 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.816531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:45.516975Z digest=sha256:02e8d6705896110230b5bd0c45c23afa81938443dfbd4fc72026a67904c32298

Observation 0833c6ca-2656-44aa-93b4-173aa01292d6 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Clapper: Compact Learning and Video Representation in VLMs Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.618643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.618643Z digest=sha256:f4cca0e49d5642051d91580bd2d96bca5b0c3660f1d1bb68d7ecf29772f07dc6

Observation 81ccf352-fc30-40ad-a176-ed4d766704f4 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Clapper: Compact Learning and Video Representation in VLMs Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.733144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.733144Z digest=sha256:5ea5dd29b0021c294aa7e7984539ce7ac690be4eef516fc4d3617a550824d87e

Observation 69f0435a-ef32-44d7-b5ee-a8290edf9e0f · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Clapper: Compact Learning and Video Representation in VLMs TempCompass: Do Video LLMs Really Understand Videos?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.814866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.814866Z digest=sha256:29f7db8361722e332af02378a62817c080314d643641b293264953add32912b7

Observation 68cd9ea3-10f6-419d-b479-3cc699219b56 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:45.926076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:45.926076Z digest=sha256:b90404ef14b3be50470207be057d0653fb145d2768e7fb402ab2e18dbbc643b9

Observation 5e32be1b-96f0-47c5-a829-3e799dc63adc · outbound

This paper cites OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation.

Clapper: Compact Learning and Video Representation in VLMs OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.025304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.025304Z digest=sha256:d2c6921d96ab2d7f993625c9078f64e3cf42b02974d98e9366d162fbb830f29e

Observation 2c42a8a7-e8b0-4b89-a580-674d9dd00723 · outbound

This paper cites GPT-4 Technical Report.

Clapper: Compact Learning and Video Representation in VLMs GPT-4 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.130255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.130255Z digest=sha256:1f1c9e6ff5855bbb2d6f0e02cfa0c9e49b39762132fdf0b3da3e248a274ac079

Observation 29148629-9ec8-41a5-8c18-a42e3e65cfc3 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.240845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.240845Z digest=sha256:f323f75f2e385d80b4a8990c0614384a7616e14aefbfd9c626cca33d22937705

Observation ccdf8af0-55ec-4984-87a4-2d1ccafd23d7 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.356809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.356809Z digest=sha256:a665c20a4a0dfa28d63bf0b1a6d63216ba88e943d2798c22b0475a4614b38e56

Observation 11bfe300-e9c2-47e6-91a2-4e1d3cdb879a · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.438533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.438533Z digest=sha256:d08c943ddb7212260752b1ef75a25cf6efdf17b71c52a8a30f95cac14a83611c

Observation 53be6b26-fba9-45f5-9341-c07df4c0678b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Clapper: Compact Learning and Video Representation in VLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.538312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.538312Z digest=sha256:c56a688b21b7b505aa14657e34b90b42ed4112397a5786467b4138e6106191b4

Observation df10b4bf-19b9-4cc1-9cda-ea7ba31abd8d · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

Clapper: Compact Learning and Video Representation in VLMs LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.656613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.656613Z digest=sha256:175aa3e2c951ea2d1904dd5eb5ffbd9531944a8a1d52314a7c64810619ee612a

Observation 8e97a8a8-ca31-4697-95fb-48e4305a115e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Clapper: Compact Learning and Video Representation in VLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.774925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.774925Z digest=sha256:33d9f26fa45212fca681dcbca14ccfdc2e7ce48dc7b4b84f1acebfda20cfe5a0

Observation 4e8687f3-97cb-4618-809d-b22cb0a964e5 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:46.882268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:46.882268Z digest=sha256:5cc2aa95750acbe5df001ab5160b89188f6953018fb2a3e0775841e9f52887ca

Observation 52373ded-3d67-44cb-b8ae-820b898f4e2f · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.489650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:46.992993Z digest=sha256:53bad50f060bd0ead098eee855e58ef3e5bee77de0c2ca05f723e239b88e6792

Observation e80577a3-53a0-4e6a-b701-074b0ce0ca31 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.314362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:47.097954Z digest=sha256:7ef0403eb072e2d8487ede4ace8974ba9d9154c69d15ac72cca040ce1ac049d5

Observation 6d1c5864-6482-4522-bde2-8406b87655e1 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:50.061845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:47.209607Z digest=sha256:103fe3d58d16455ba66689613f542f3f9641bd4c0d37c96ab551195a4ef031d6

Observation 65d26525-3b22-4b76-8c52-335ada6560d3 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.360957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.360957Z digest=sha256:935994faca5a4deda26f0b37ab51b1cacb8bd2543780ae0abf274e9327f87303

Observation 4e19bcb9-4faa-4af7-8a25-9f01f45e34d9 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

Clapper: Compact Learning and Video Representation in VLMs PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.442019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.442019Z digest=sha256:65c884a77a0434ec8abd2c65e5024c1e6a227b29f6499b93127d0cd0f147b038

Observation 0a170583-a703-4bcf-9e43-4138824d1160 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

Clapper: Compact Learning and Video Representation in VLMs SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.549466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.549466Z digest=sha256:e30330f605a7705f0a589f9c3d96cf6d4e6ea94cc4b4cf9239b4bfcc8fcb0ecb

Observation 504d6f4d-54f0-4a77-8703-264b2636ec34 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

Clapper: Compact Learning and Video Representation in VLMs LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.582468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.582468Z digest=sha256:30465c79b14cc0c920140b93d62ecca29ecfded92eebbab3e0aa232670538d75

Observation 3b7b5785-9727-433c-8acf-043e29bc57cb · outbound

This paper cites Qwen2 Technical Report.

Clapper: Compact Learning and Video Representation in VLMs Qwen2 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.666280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.666280Z digest=sha256:817d1a43e485b084261dad10da88f8fcd0215eb7efeed76b741b5922f52d92e4

Observation dcf3998a-0299-44c7-99b1-60537273aa35 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Clapper: Compact Learning and Video Representation in VLMs MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.766618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.766618Z digest=sha256:9dd530d11a53ab712fee7fd662caeccbf6e6e90ab0a552b2901e47d90b866e49

Observation 6123c3b3-0e54-484e-b7dd-3352b9d48e95 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

Clapper: Compact Learning and Video Representation in VLMs mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.848486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.848486Z digest=sha256:e0211a6e9264786777ccb1c9a6c93be9732c385e4fa3d031ef43ed0f7b2e20af

Observation b7036e20-0a50-4793-a074-fb882ae4e945 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:47.964092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:47.964092Z digest=sha256:12dc763f6e0f4a78e9df1070daa887ae46706432849bdee9e8df5c5dd48f6628

Observation b080d67c-5c39-4dbb-8d51-116ecac20bc5 · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.040187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.040187Z digest=sha256:c649a6bfa41e4fae6e6aa08d3190a97f611b90553accb590e4529ced9d3d482f

Observation e0f45e9c-7100-445f-a5c3-ea4118e6ebbc · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.150431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.150431Z digest=sha256:85d0ef01b5e86652f786204b11340b82225b55af94559ceb3f0a2d5bdeebfa8c

Observation cdfe6a61-b8c6-4b66-8526-ef453138e929 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Clapper: Compact Learning and Video Representation in VLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.289056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.289056Z digest=sha256:c32a91a9a902850613a0d8292d3079aedc81ee99b502a6e54869860fb8ad8119

Observation 0fec3900-d476-4a76-b440-735d2d295676 · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

Clapper: Compact Learning and Video Representation in VLMs InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.439015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.439015Z digest=sha256:2474d95515becf36f5ade174ea35077e7e0f7ebdcde0c4942ba75c7366574067

Observation 896f522a-1fd3-4c1f-9435-e1462eb735e9 · outbound

This paper cites Long Context Transfer from Language to Vision.

Clapper: Compact Learning and Video Representation in VLMs Long Context Transfer from Language to Vision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.560323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.560323Z digest=sha256:1171f2c3801f760be3990c18655e3731730519eb5ea95b3389d354ff85bcd7bd

Observation a2c8dda5-b1a7-4752-a01d-d624987f1540 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Clapper: Compact Learning and Video Representation in VLMs Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.669699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.669699Z digest=sha256:75f8fbd74040cbf0df50e442e6e80f86f01e261b865c7898eade9d24c19d34bb

Observation 61da8d38-d993-463d-b042-23d06d7753db · outbound

This paper cites an unresolved cited work.

Clapper: Compact Learning and Video Representation in VLMs Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:20:49.835526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T15:20:48.791387Z digest=sha256:384727d86dc29905a38dd5627be9071e40c57e8dcbc78047a4421886b98e6f8e

Observation 8eacf7cc-13ab-4b30-9181-67af9725b7e1 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Clapper: Compact Learning and Video Representation in VLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:48.883282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:48.883282Z digest=sha256:bef394315c5f990ea5e69baa6acac42cc5c106dd7d3d76098526ba08cd837144

Observation 20c7a929-9e13-4928-860f-e2097167ac49 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

Clapper: Compact Learning and Video Representation in VLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:49.012698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:20:49.012698Z digest=sha256:57192f408c4781560313580da09079380ab597c77f2088682909e9d14a5b1fdc

Pith citing papers

No inbound Pith citation observations are available.