Pith. sign in

Paper Citation Record · LEDGER

Valley: Video Assistant with Large Language model Enhanced abilitY

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2306.07207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.07207 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:37:42.248728Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.409899Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 897e4d1c-875a-441b-b2f8-26c72af5d579 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.650439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:585c0f81a904caf75d8d5efd824ee63e6699598c43173c2f5aad42c0603bec61

Observation c342be09-1bd7-449d-b81d-5e098411cd7c · inbound

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection cites this paper.

Video-LLaVA: Learning United Visual Representation by Alignment Before Projection Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:08:01.242531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T18:08:01.166072Z digest=sha256:44b07c32eab7444d787e269cb4a64339ebd7aee1a247a085f9620b3c5972c030

Observation f560f83b-f72e-4b21-9b3f-27f69a86c5a5 · inbound

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark cites this paper.

MVBench: A Comprehensive Multi-modal Video Understanding Benchmark Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:22:35.014045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:22:34.954228Z digest=sha256:bf4caf13ea5fd6dc99b54b2d3b43c84ed4a9292977e41fab53c3ffd62e57473c

Observation da2205a2-3c68-494f-bd58-63204b3bf502 · inbound

TempCompass: Do Video LLMs Really Understand Videos? cites this paper.

TempCompass: Do Video LLMs Really Understand Videos? Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:16.769669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T02:46:16.632743Z digest=sha256:5068edf3844d5e1af31bfedd2fdff0c34fed968f88bcf1de56f48882afb5053d

Observation d8272357-eb90-40ef-b937-06bdfc05800d · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.884272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:3bc490ed73b37f03a835fc05abcb132e159e6778ee9410a6d99c874ac2e3d1b2

Observation 634a0eee-6bd6-4860-99f0-19f0893c45f8 · inbound

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding cites this paper.

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T13:53:33.675561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T13:53:33.585035Z digest=sha256:871d03ae4ed9e6a3a9c95c5d9a96b8ea52b6b4adeea95e477fd78b1455d6df32

Observation dc8993eb-df65-4539-8b6b-4bd24028e0e2 · inbound

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance cites this paper.

PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:33:15.706077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T17:31:59.030963Z digest=sha256:43dd186e48783e1f15210dca620e952d6d37bc9b0b95f07a8490c086cb8678c1

Observation 8f5e0e6f-f569-4ead-809a-2a3b4e6d6d3e · inbound

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos cites this paper.

TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T08:12:43.870984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T08:08:01.889675Z digest=sha256:c73de1a3b90a5540e838c642e072c8c3ea1cf5d600eb83c6f0d8a01fe4434dec

Observation 08271f24-be62-4741-8711-197e91098711 · inbound

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO cites this paper.

FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:22:18.746828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T00:21:51.621582Z digest=sha256:2088aef0b3a46f6cf51ff57d69f80b7928c840eb17b1e10765a0811ddf816ac3

Observation c4229669-f0ae-4e90-a50f-ac08f48c9807 · inbound

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding cites this paper.

UniMind: Unleashing the Power of LLMs for Unified Multi-Task Brain Decoding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:47:10.369499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:43:26.736409Z digest=sha256:03ce207a5fe944fd75beeb4a52078786bcfb6ffe56b64ee6cc6ff4a977fd38c4

Observation 34b2886f-ccd6-447f-b1c7-55ef0e380930 · inbound

Video Understanding by Design: How Datasets Shape Video Models cites this paper.

Video Understanding by Design: How Datasets Shape Video Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 248

Resolution
unresolved
no resolver link, observed 2026-08-04T19:37:42.248728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:37:42.248728Z digest=sha256:29737cef34a142d528db1deebbcf7d4c440ba31ab53c0ef622af48019044daf0

Observation d5431206-c977-4ddb-9a17-869ded7abf2b · inbound

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM cites this paper.

SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T21:42:47.855446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:42:47.855446Z digest=sha256:c07582a653659f2701182788592ab10d4b1018e09bfaba7fc88f6bd98cc2de42

Observation 0536ee5a-5b72-4bb6-b5f3-d8f5ecdfcc87 · inbound

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning cites this paper.

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:30.906964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:57:30.906964Z digest=sha256:d9efd6e02a5858d2ee9d9939487ca8c091c9d01832642593009178e17e3210e3

Observation 3780dad6-ab2a-414b-8f9d-9e81fe75ccdc · inbound

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding cites this paper.

Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.556812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:54:53.789523Z digest=sha256:83433b58a501d716ba2b67531f4e7a5ad4ffb834c1b892238d21cc79b5a5b7fd

Observation 89f4997e-db72-495e-a46d-391c9c729afc · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.044127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:2764d8e011e6ae0fcf427d887584542b8484a7eb5022e9b6d53a8a1f15b08af8

Observation 053cc833-2807-415b-b9d3-2620cdc741e8 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.924293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:5466afb54eaef2c7c7970f52d4f42973ed61a5c136050655b40d85d52560dec1

Observation 6d2e3842-bbb3-4e92-84ad-800e932d1c89 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:9ee70a74efe68024e083c1c7adf99948053bb1e644db76488342f1ff29a88569

Observation 6ef3b4ba-347d-4592-8f41-d38510568d4b · inbound

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding cites this paper.

One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.566195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:28:58.920442Z digest=sha256:b2ed5c3142db7b0aa72d4a7daa46333670b61d2341a74fd445be37d66871d6a2

Observation 9a223249-f3da-42a3-af03-839e912aeace · inbound

ClimateVID -- Social Media Videos Analysis and Challenges Involved cites this paper.

ClimateVID -- Social Media Videos Analysis and Challenges Involved Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:30.448614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-07T05:25:05.760276Z digest=sha256:965162592df1c33b234a57516d0c97a952bcb3d63c91bf639902152764a34bf8

Observation 696d60da-b03f-4d11-a459-80b296cf68c3 · inbound

Dynamic Model Merging Made Slim cites this paper.

Dynamic Model Merging Made Slim Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:18:25.408077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T15:16:24.651868Z digest=sha256:b0e0ac53b3a75a9c1ac6be41b9edb8673391ab685e7f551e8683d4779eeb986e

Observation 54e5d202-bf23-46d1-98fb-d3bd45514b1f · inbound

Closed-Form Spectral Regularization for Multi-Task Model Merging cites this paper.

Closed-Form Spectral Regularization for Multi-Task Model Merging Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:27:09.379995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:40:00.510742Z digest=sha256:e63be4bf70d55b1a5085a3d9262c4683fa09ac2da3bc8de5eb9f08fb37937fb9

Observation 3b7dfae4-d4b9-41b7-8a07-ff1bc762f762 · inbound

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning cites this paper.

Audio-Visual Exchange-Aware Token Pruning for Efficient Audio-Visual Captioning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:36.879950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T13:53:53.520545Z digest=sha256:e02f7ea5d026303e17d3db9bee80783078ab4a8a444df341add63f0b51cb098b

Observation c2a7f144-34cd-4dad-9773-a3100e589982 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:48:02.979640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:b79a2427aeb596c70f3920f874c54739e7e7f773b180546ad9efc302fdae958e

Observation 0b1ceda6-c948-4963-b7a4-e89a65659738 · inbound

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning cites this paper.

On the Sparsity-Storage-Accuracy Tradeoff in Parsimoniously Activated Dictionary Learning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:39:42.125759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:13:49.266859Z digest=sha256:fbee2a656655845522c501cdc05ae2e5b2906359af03c7ceb0e13c4ba391e992

Observation 2bf24e0f-0851-454d-97e7-f1d6c87a3583 · inbound

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning cites this paper.

Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.411854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T21:31:38.450382Z digest=sha256:a38b532765774d50e166753efb10a6637d369a023b8755c0d47779c6e570d1a0

Observation c1801018-9997-4953-8a27-fcc10b024224 · inbound

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models cites this paper.

MoHallBench: A Benchmark for Motion Hallucination in Video Large Language Models Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:58.685914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T13:44:48.250838Z digest=sha256:7bf70a197b71ad0b2729bd848267bb929ac22db283c8326dbbc2aaa53a193af2