Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:25.588833Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2508.15903.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:25.588833Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
27 of 27 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 30546734-549b-421e-9c22-975274f512e2 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Human action recognition from various data modalities: A review,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d252faa4-4aab-465e-b185-618752afe075 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos 2d versus 3d convolutional spiking neural networks trained with unsupervised STDP for human ac- tion recognition,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e50b3ee1-68e1-40bf-8e19-38eda428f0c7 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Overview of the transformer-based models for NLP tasks,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75379562-ca53-4ae5-9323-5bcc340240f1 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 393296a3-aa62-407b-9c96-27bfb523ed99 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b93cb2b3-2c7e-46dc-a229-4e7af91113e1 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Visual in-context learning for large vision-language models,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f7e4b3b-4da5-4477-ab48-b442500aef50 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Weak to strong generalization for large language models with multi-capabilities,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2c5cd88-7f4b-4abf-a9a9-6be46a81d246 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation beb377fc-8139-40d5-8afe-405dedf492f2 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Adaptive prompt: Unlocking the power of visual prompt tuning,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bdffa8b-4257-4b3c-b07d-b184751a1123 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos NTU RGB+D: A large scale dataset for 3d human activity analysis,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e6dc61f-30be-4c19-9a88-f3b716458e0a · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Human action recognition and prediction: A survey,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2361002e-23bb-49df-8111-06a8d05a507e · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A comprehensive study of deep video action recognition,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79808f44-f637-4276-a7fe-b7fa2f194032 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Action recognition based on efficient deep feature learning in the spatio-temporal domain,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7ebe177-7957-4a59-a98a-6e0c5f04b272 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Cross-fiber spatial-temporal co-enhanced networks for video action recognition,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d38b2284-ef73-454c-8d50-24c3e7f36bde · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Mutually reinforced spatio-temporal convolutional tube for human action recognition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b2ad58e-545b-4a36-8881-17ec5cec0e07 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Multi-scale spatial- temporal integration convolutional tube for human action recognition,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a6e302f-a6d6-43c0-a56b-c4ee621fe105 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Long-term temporal convolutions for action recognition,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ec2b4758-7c2f-4a5c-9952-0cd638e80b82 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Finegym: A hierarchical video dataset for fine-grained action understanding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6cc55d62-1891-4051-8279-ad907b4c8366 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos End-to-end video-level representation learning for action recognition,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9ba7246f-d83e-48b7-9661-d24dd9c51378 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Stnet: Local and global spatial-temporal modeling for action recogni- tion,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f36d86e0-9f0b-4436-b838-d6922f39f59e · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A survey on efficient vision-language models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60923db0-71b2-4be5-a736-e99c2b677598 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Cheap and quick: Efficient vision-language instruction tuning for large language models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ea32719-ee71-46ea-b40b-b0638e01085c · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos PEFT A2Z: parameter-efficient fine-tuning survey for large language and vision models,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa2bd55c-8085-4894-a159-46f8e5e017c8 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Visualgpt: Data- efficient adaptation of pretrained language models for image captioning,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ed8435e0-1800-4210-b87d-c91cf362d305 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Multi-modal large language models are effective vision learners,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1280096c-bfe2-4c67-938d-9ca493b7f1bb · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e97a3f1-20f4-44d3-a72d-793938ec7988 · outbound
VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Aligngpt: Multi-modal large language models with adaptive alignment capability,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.