Pith. sign in

Paper Citation Record · LEDGER

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models

As of 18 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2601.22574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.22574 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:26.484082Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:31:40.463891Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:36:31.820599Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9eef630c-9d12-4349-937f-3ceb8eba3994 · outbound

This paper cites Grounding language with vision: A conditional mutual information calibrated decoding strat- egy for reducing hallucinations in lvlms.arXiv preprint arXiv:2505.19678,.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Grounding language with vision: A conditional mutual information calibrated decoding strat- egy for reducing hallucinations in lvlms.arXiv preprint arXiv:2505.19678,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.276084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.276084Z digest=sha256:57bd19576ef986605ce6ac79af7880c8b62364b24e2a3b814c1ba13e641dd8ab

Observation 47c1b3e6-57a4-4dab-94d9-c39b6b685622 · outbound

This paper cites Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.319744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.319744Z digest=sha256:a647f0b0ff6f18f31cff38b6f712d98aff1c6862bc15a3affefbdd37c44ee2b8

Observation f68a30ba-060e-45b2-8058-16df557cbff3 · outbound

This paper cites Mentalmac: Enhancing large language models for detect- ing mental manipulation via multi-task anti-curriculum distillation.arXiv preprint arXiv:2505.15255, 2025b.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Mentalmac: Enhancing large language models for detect- ing mental manipulation via multi-task anti-curriculum distillation.arXiv preprint arXiv:2505.15255, 2025b

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.378480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.378480Z digest=sha256:e5ba2255031539b02cd8a4f893263d0843ea829354d0e9994fce044a7c6e4e86

Observation 41f780d3-8f56-4fd1-b112-2ccb19bdf0e7 · outbound

This paper cites VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.474447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.474447Z digest=sha256:b6abc71a2321ff3d21a82b2a7b67ba22cac4f929a12eadc406b369a20870222f

Observation bd967670-d41d-46f1-b8f7-289c78e505c0 · outbound

This paper cites H., Jo, Y ., and Seo, M.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models H., Jo, Y ., and Seo, M

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.569043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.569043Z digest=sha256:6f7dc029db5dcbbd53bc07d0450ab3ccdb05b137dde2daca52a5d0344143b4be

Observation d5808543-5658-4a0d-a05f-415c4feb81f9 · outbound

This paper cites Decoupled Weight Decay Regularization.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Decoupled Weight Decay Regularization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.779451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.779451Z digest=sha256:e4b8ee88b04ac027cdbeb0167c49e56458cea359be12f00c1a0828e08fa08087

Observation da3e0954-924c-4902-80ea-91d18c307674 · outbound

This paper cites Countervid: Counterfactual video generation for mitigat- ing action and temporal hallucinations in video-language models.arXiv preprint arXiv:2601.04778,.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Countervid: Counterfactual video generation for mitigat- ing action and temporal hallucinations in video-language models.arXiv preprint arXiv:2601.04778,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.834881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.834881Z digest=sha256:b56a14aa85a5fcac564159262e46cfc608259ab2b80714b44977fd67fea9192f

Observation 4de1f7b8-f60c-4ef0-a021-b1bbfe3febcc · outbound

This paper cites Smart- sight: Mitigating hallucination in video-llms without com- promising video understanding via temporal attention collapse.arXiv preprint arXiv:2512.18671,.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Smart- sight: Mitigating hallucination in video-llms without com- promising video understanding via temporal attention collapse.arXiv preprint arXiv:2512.18671,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.911702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.911702Z digest=sha256:59e1720b7771da0182113f1bfb39ed38d3f6e5c3f97da621eecd03f4f2a02970

Observation d80fcd96-959f-4a72-9985-1e19fcb2a6e0 · outbound

This paper cites VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.055855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.055855Z digest=sha256:38bb95d9d60d2928f77214e57fd9b548aced6ec05e94c9801e001cab5cf38a15

Observation 4183cc5b-88aa-481a-af8f-b4baa078a506 · outbound

This paper cites an unresolved cited work.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.108846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.108846Z digest=sha256:cca22cabf4dee27def9d5db403946e7e9e7a25e397764c302fb80fd283731414

Observation 9498603f-c74d-4438-8ebf-1b6bc3972838 · outbound

This paper cites Kardia-r1: Unleashing llms to reason to- ward understanding and empathy for emotional support via rubric-as-judge reinforcement learning.arXiv preprint arXiv:2512.01282, 2025a.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Kardia-r1: Unleashing llms to reason to- ward understanding and empathy for emotional support via rubric-as-judge reinforcement learning.arXiv preprint arXiv:2512.01282, 2025a

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.211914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.211914Z digest=sha256:35990e8f4866e5f5b19b0c6e156c91059dcd5f12e32aa32d7a40c6de097c8d75

Observation d5c1c029-6c75-436f-b0b4-16d7c869f433 · outbound

This paper cites Video-llama: An instruction- tuned audio-visual language model for video understand- ing.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Video-llama: An instruction- tuned audio-visual language model for video understand- ing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.275069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.275069Z digest=sha256:0aeff452e2839cd62915d29e331af21f63fa455e214cd0c7ba65f3e7916709c1

Observation 96dd6031-5c4b-4cab-a7c0-83252aa71abb · outbound

This paper cites Eventhallusion: Diagnosing event hallucinations in video llms.arXiv preprint arXiv:2409.16597, 2024a.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Eventhallusion: Diagnosing event hallucinations in video llms.arXiv preprint arXiv:2409.16597, 2024a

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.315860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.315860Z digest=sha256:664247d37184880204ef8c462e1bf298a9efbeb91c78dd664317199e4e6e6f58

Observation 237834a3-6d1f-41ef-9563-5d0b483f3496 · outbound

This paper cites j., Gui, L., Fu, D., Feng, J., Liu, Z., and Li, C.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models j., Gui, L., Fu, D., Feng, J., Liu, Z., and Li, C

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.367685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.367685Z digest=sha256:eff4d5235bb69c1d70bc625666e826a646e2d4e535e5a6bd28a95b808f7d3871

Observation 659eaaf6-511a-473d-9f5c-9b76a50e4198 · outbound

This paper cites Can pruning improve reasoning? revisiting long-cot compression with capability in mind for better reasoning.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Can pruning improve reasoning? revisiting long-cot compression with capability in mind for better reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.425417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.425417Z digest=sha256:4dd4955cd74c6e59edc63507d01af4366cfe9b549effeb0ca8fd834dc2276759

Observation 5e077945-d7b4-47c6-84b1-e83c05dd89c0 · outbound

This paper cites Layernorm Linear GeLU Linear GeLU Linear Tanh Layernorm Linear GeLU Linear GeLU Linear Tanh Figure 6.Overview of the architecture of SSD.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Layernorm Linear GeLU Linear GeLU Linear Tanh Layernorm Linear GeLU Linear GeLU Linear Tanh Figure 6.Overview of the architecture of SSD

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.484082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.484082Z digest=sha256:daf3301446957999e881faa51c3a52617a28dcd30052d6de745257365d7ae203

Observation b7b21c44-bcda-46b7-982c-d67aab7717c5 · outbound

This paper cites Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.992400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.992400Z digest=sha256:a6666d3ed6c270a679caf8bab8d71cc0e0dc1ca4b2046020ad13c84bf34426ec

Observation 0a6763e4-2535-425b-9843-2594ba84dcf7 · outbound

This paper cites Helpd: Mitigating hallucination of lvlms by hierarchical feedback learning with vision-enhanced penalty decoding.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Helpd: Mitigating hallucination of lvlms by hierarchical feedback learning with vision-enhanced penalty decoding

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:26.165448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:26.165448Z digest=sha256:39e26c3338c77fc18bda692768cfc42f58ece929643028294c3f2b75b8fc8c1c

Observation 79c897ce-3b31-45da-9dc2-7ccd05ab297e · outbound

This paper cites an unresolved cited work.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Unresolved cited work

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.694297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.694297Z digest=sha256:f7b36c81fb234c7ceb1dbf7e29b288cab8a9740af6261c4f5c3ff5cb7ddd43b3

Observation ddf08541-6493-4209-b923-2632e50ca56d · outbound

This paper cites Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.124649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.124649Z digest=sha256:6e53bbf2703d624a5d5ef4c66dc504a756b3b57c4f27547705618e81631509b4

Observation 8f265e5e-95f3-4a29-a098-211c1d0d2a1c · outbound

This paper cites Decoupling contrastive decoding: Robust hallucination mitigation in multimodal large language models.arXiv preprint arXiv:2504.08809,.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models Decoupling contrastive decoding: Robust hallucination mitigation in multimodal large language models.arXiv preprint arXiv:2504.08809,

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.182617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.182617Z digest=sha256:4eb17cbb186b848a95a87762e323e4ded4b036cc016b9212f4610f7313bc38f1

Observation 3c5176a0-b158-44a3-b733-7f69cc794a5b · outbound

This paper cites PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning.

Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:25.228855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:25.228855Z digest=sha256:ee19c08f9492a2d7f4352ba3c4005e9bc63995d794842e595fd5f2215d88ca0c

Pith citing papers

Observation a41e6d9f-3019-400d-b960-51b86cea64b0 · inbound

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models cites this paper.

Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models Enhancing Video Representations with Spatiotemporal-Semantic Residual to Mitigate Hallucinations in Video Large Multimodal Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-08T02:03:50.869564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:31:40.463891Z digest=sha256:68d8c3732a9af0e7b7e18df6e388f0f2633d2016198d9f76771c33e5aa796a67