Pith. sign in

Paper Citation Record · LEDGER

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos

As of 7 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2508.15903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.15903 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:45:25.588833Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30546734-549b-421e-9c22-975274f512e2 · outbound

This paper cites Human action recognition from various data modalities: A review,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Human action recognition from various data modalities: A review,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:29.123621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:23.060592Z digest=sha256:cc737d481643539012438071d5af0db44deb3e8640f200cd958336bbb4001156

Observation d252faa4-4aab-465e-b185-618752afe075 · outbound

This paper cites 2d versus 3d convolutional spiking neural networks trained with unsupervised STDP for human ac- tion recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos 2d versus 3d convolutional spiking neural networks trained with unsupervised STDP for human ac- tion recognition,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.949894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:23.146665Z digest=sha256:3336b8b780f3dc1fd5e25d2fbc1c53d9c8793f3e4ab342e6975961bfb941f4fb

Observation e50b3ee1-68e1-40bf-8e19-38eda428f0c7 · outbound

This paper cites Overview of the transformer-based models for NLP tasks,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Overview of the transformer-based models for NLP tasks,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.789963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:23.228078Z digest=sha256:539a666c500dbfb45aebc42243a78b00f4cda3b6b0ccf85aa849efbae8a4a9db

Observation 75379562-ca53-4ae5-9323-5bcc340240f1 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:23.362858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:23.362858Z digest=sha256:ef726d98e5d2045e6322233be8defec6bad20acecfa31b2039a290ea7f0ccdea

Observation 393296a3-aa62-407b-9c96-27bfb523ed99 · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.625634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:23.471682Z digest=sha256:9d0521d3419f7fa665dcfca971c677a29a6a23097c25aa9f267972b0f6614b0e

Observation b93cb2b3-2c7e-46dc-a229-4e7af91113e1 · outbound

This paper cites Visual in-context learning for large vision-language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Visual in-context learning for large vision-language models,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:23.600178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:23.600178Z digest=sha256:61eadd3ef653236a062ceb0fc9760522345373d2a833e28bdee416fa253846e5

Observation 9f7e4b3b-4da5-4477-ab48-b442500aef50 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Weak to strong generalization for large language models with multi-capabilities,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:23.702577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:23.702577Z digest=sha256:32beba715668157100e83f7ce6a3286de109e9b7a94a562a64ccd2a51321c014

Observation c2c5cd88-7f4b-4abf-a9a9-6be46a81d246 · outbound

This paper cites Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Llava-more: A comparative study of llms and visual backbones for enhanced visual instruction tuning,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.462300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:23.829216Z digest=sha256:66865ba2842f7333500f4c8d2c3e1d33325c3fed5bca452b8cfe56bf0ad26968

Observation beb377fc-8139-40d5-8afe-405dedf492f2 · outbound

This paper cites Adaptive prompt: Unlocking the power of visual prompt tuning,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Adaptive prompt: Unlocking the power of visual prompt tuning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.243612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:23.958129Z digest=sha256:8b85e6458bc435fe3f62fca23b808a8a687fe527922dc286c48cf83f66de5937

Observation 4bdffa8b-4257-4b3c-b07d-b184751a1123 · outbound

This paper cites NTU RGB+D: A large scale dataset for 3d human activity analysis,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos NTU RGB+D: A large scale dataset for 3d human activity analysis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:28.046819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.084139Z digest=sha256:8b4d8009ca3e18c6ef59313b54a2cd449e060d8e8940f8287f3756da929b730d

Observation 4e6dc61f-30be-4c19-9a88-f3b716458e0a · outbound

This paper cites Human action recognition and prediction: A survey,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Human action recognition and prediction: A survey,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.168894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.168894Z digest=sha256:3ab1f859cba3daffb16f9cd0bd2d6f34259bd8221212f77fa817b01e47d37a8e

Observation 2361002e-23bb-49df-8111-06a8d05a507e · outbound

This paper cites A comprehensive study of deep video action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A comprehensive study of deep video action recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.884575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.257880Z digest=sha256:d9516a3d59c3428418742af8901dafa0ad702a7b65895b45aec4b55accdd61ef

Observation 79808f44-f637-4276-a7fe-b7fa2f194032 · outbound

This paper cites Action recognition based on efficient deep feature learning in the spatio-temporal domain,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Action recognition based on efficient deep feature learning in the spatio-temporal domain,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.720535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.356694Z digest=sha256:25f0ff1fbf364de6ed449ac3e087a496b5aa16c050f29b3bec44b70e22c17ee0

Observation e7ebe177-7957-4a59-a98a-6e0c5f04b272 · outbound

This paper cites Cross-fiber spatial-temporal co-enhanced networks for video action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Cross-fiber spatial-temporal co-enhanced networks for video action recognition,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.437028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.437028Z digest=sha256:7cfd428653973a3cbfb5826f1cb66c136701aaf73ec09810c1571c6a376919d3

Observation d38b2284-ef73-454c-8d50-24c3e7f36bde · outbound

This paper cites Mutually reinforced spatio-temporal convolutional tube for human action recognition.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Mutually reinforced spatio-temporal convolutional tube for human action recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.563094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.563094Z digest=sha256:ff6223a40368ee6083571dd226bb5fb8f6e431ecf660cb03fe770d7044f723b3

Observation 1b2ad58e-545b-4a36-8881-17ec5cec0e07 · outbound

This paper cites Multi-scale spatial- temporal integration convolutional tube for human action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Multi-scale spatial- temporal integration convolutional tube for human action recognition,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:24.625203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:24.625203Z digest=sha256:64d843ef3e7f22e6877f8af6f095b0414ddacb6214e7acc897c8cc0d7f1f9883

Observation 5a6e302f-a6d6-43c0-a56b-c4ee621fe105 · outbound

This paper cites Long-term temporal convolutions for action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Long-term temporal convolutions for action recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.583506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.707195Z digest=sha256:6a0dabbc6be9497d92290ad978950f3595570f419b32633e9aeea398c5efe97c

Observation ec2b4758-7c2f-4a5c-9952-0cd638e80b82 · outbound

This paper cites Finegym: A hierarchical video dataset for fine-grained action understanding,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Finegym: A hierarchical video dataset for fine-grained action understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.375395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.810896Z digest=sha256:aeeb9ee25f0356adb5489913b2bb695e2c16e5a37f261d8ec91d9e6ba79acce3

Observation 6cc55d62-1891-4051-8279-ad907b4c8366 · outbound

This paper cites End-to-end video-level representation learning for action recognition,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos End-to-end video-level representation learning for action recognition,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:27.153608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.875343Z digest=sha256:3f5a09d130deb422a2403c331e2145a046a4e35e41e195dd6ef65483de6ba5fd

Observation 9ba7246f-d83e-48b7-9661-d24dd9c51378 · outbound

This paper cites Stnet: Local and global spatial-temporal modeling for action recogni- tion,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Stnet: Local and global spatial-temporal modeling for action recogni- tion,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.955176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:24.943847Z digest=sha256:e00fdcc069c30536d61cb756c3d80247bd35e17604c54d17a255aac8e6bb4ff0

Observation f36d86e0-9f0b-4436-b838-d6922f39f59e · outbound

This paper cites A survey on efficient vision-language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos A survey on efficient vision-language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.739960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:25.030702Z digest=sha256:9e014ecdcd914ba52db4ebeb92538ec9536b0e0689d0b120d38dd2ae7d971f4c

Observation 60923db0-71b2-4be5-a736-e99c2b677598 · outbound

This paper cites Cheap and quick: Efficient vision-language instruction tuning for large language models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Cheap and quick: Efficient vision-language instruction tuning for large language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.578959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:25.141410Z digest=sha256:1a5ebe07616bffdffdd0ec8df20de752d3cb2f48c25ad1b71a6309d0c9926bc0

Observation 5ea32719-ee71-46ea-b40b-b0638e01085c · outbound

This paper cites PEFT A2Z: parameter-efficient fine-tuning survey for large language and vision models,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos PEFT A2Z: parameter-efficient fine-tuning survey for large language and vision models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.404939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:25.221123Z digest=sha256:12612fa8d1b99255f7d6920a7dfe66fdc6698f4e40e2a300c4f4f07d0379807f

Observation fa2bd55c-8085-4894-a159-46f8e5e017c8 · outbound

This paper cites Visualgpt: Data- efficient adaptation of pretrained language models for image captioning,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Visualgpt: Data- efficient adaptation of pretrained language models for image captioning,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:26.204200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:25.317266Z digest=sha256:e2a0e5acd174269704f5c0c7044f5dbc99e98df5d06546003a70ca6e1c68cb58

Observation ed8435e0-1800-4210-b87d-c91cf362d305 · outbound

This paper cites Multi-modal large language models are effective vision learners,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Multi-modal large language models are effective vision learners,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:25.997782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:25.381519Z digest=sha256:13db573abedc3466b5ad909b3796cab08477018408f29dfd04612fe60e6f2064

Observation 1280096c-bfe2-4c67-938d-9ca493b7f1bb · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:25.477587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:25.477587Z digest=sha256:41b7271227e48e1e9a580a8c640f9c7b32209ae62aee7126087d8a369b8dcd2e

Observation 0e97a3f1-20f4-44d3-a72d-793938ec7988 · outbound

This paper cites Aligngpt: Multi-modal large language models with adaptive alignment capability,.

VT-LVLM-AR: A Video-Temporal Large Vision-Language Model Adapter for Fine-Grained Action Recognition in Long-Term Videos Aligngpt: Multi-modal large language models with adaptive alignment capability,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T17:45:25.836109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T17:45:25.588833Z digest=sha256:5a652b7d3ecb41db5bffcd9507581b668019acc17687923743e4e80027e44fc2

Pith citing papers

No inbound Pith citation observations are available.