Pith. sign in

Paper Citation Record · LEDGER

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition

As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2508.17442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.17442 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:55:04.858493Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3f1d62b3-e4ff-44e3-88cc-74fe63ac909f · outbound

This paper cites Human action recognition and prediction: A survey,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Human action recognition and prediction: A survey,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.377306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:02.710999Z digest=sha256:3d14ecc34aae6d34b36944da09c223a1b773b5ef47fb8b6e49d10652cb2f2ef4

Observation ca917620-26f2-4f5a-a9ef-15b7c3fb264e · outbound

This paper cites How would autonomous vehicles behave in real-world crash scenarios?.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition How would autonomous vehicles behave in real-world crash scenarios?

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.365851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:02.768294Z digest=sha256:62648e4f52c16774bf615879678e5fd21fb129978651903b8b51f72b7541affb

Observation f23be0e3-7c48-4726-b026-9ec4087861db · outbound

This paper cites Crash-based safety testing of autonomous vehicles: Insights from generating safety- critical scenarios based on in-depth crash data,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Crash-based safety testing of autonomous vehicles: Insights from generating safety- critical scenarios based on in-depth crash data,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.355542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:02.849541Z digest=sha256:54a4d4e57d3464ae0a62c45dd3c2260773dd8bf6f69ca9e377786e4335a8f357

Observation 3c4d22bc-9847-4c6b-be31-85d70a0292ae · outbound

This paper cites Diffcrash: Leveraging denoising diffusion probabilistic models to expand high- risk testing scenarios using in-depth crash data,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Diffcrash: Leveraging denoising diffusion probabilistic models to expand high- risk testing scenarios using in-depth crash data,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:09.181785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:02.930306Z digest=sha256:92a7114e20014a0a5d6b857d3623ff3218719961356fdd2de3af1d5d3c8edea7

Observation 6bf11373-c719-4f82-8435-1b6698d8172d · outbound

This paper cites Search-to-crash: Generating safety-critical scenarios from in-depth crash data for testing autonomous vehicles,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Search-to-crash: Generating safety-critical scenarios from in-depth crash data for testing autonomous vehicles,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.901226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.044456Z digest=sha256:c9ded3fa6f342fa3a7e9cebe6f8e5a92796c7ab0d010527f30e3499b1cbf47c9

Observation b35bfb34-040c-4576-8ba5-94b3c409ed6b · outbound

This paper cites Visual features of interme- diate complexity and their use in classification,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Visual features of interme- diate complexity and their use in classification,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.705485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.140644Z digest=sha256:7de8110ba6ea1b7f66887265479abfec3f01725414f65fe6d37cd90baf140c7d

Observation 0bf434c0-2c9b-4dc3-bd96-29e13993419b · outbound

This paper cites Visual in-context learning for large vision-language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Visual in-context learning for large vision-language models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:03.247944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:03.247944Z digest=sha256:2b3f2411b77cb29d2f41b81a73868e85ab234fdb18bf7bc0d90fea0e0afcf119

Observation 9211e064-77c1-4010-bfdf-c90ba7e07438 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Weak to strong generalization for large language models with multi-capabilities,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.527606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.335362Z digest=sha256:a6c35946c7566b99ded971707fca733e6871f98ebac838e41b8b03897e352dad

Observation 224aaed7-67a0-4e16-a210-5ce0c1e3997f · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:03.423729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:03.423729Z digest=sha256:6eb71476e88e9c6493f8d09e76ed903bcdc1fbe5797b4bac342f5797b99e14f3

Observation d4a7da62-e621-486a-9fa3-16dc0b86d2cd · outbound

This paper cites Cbr-net: Cascade boundary refinement network for action detection: Submission to activitynet challenge 2020 (task 1),.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cbr-net: Cascade boundary refinement network for action detection: Submission to activitynet challenge 2020 (task 1),

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.342940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.530428Z digest=sha256:b6cb0b73731ce0b863ab74c5c21aa7445f8d6f146fa5beec4c552e8cd4ad389a

Observation 03899956-4933-490a-a4c7-629d929cd2a9 · outbound

This paper cites Fineaction: A fine- grained video dataset for temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Fineaction: A fine- grained video dataset for temporal action localization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:08.146218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.602521Z digest=sha256:d5295e93dd3af19bca477f4479d2bdf04def8daaef73ea34f8e1dada918b9c12

Observation ef41a983-1680-4f33-b553-8d6fb8d8fe5b · outbound

This paper cites The THUMOS challenge on action recognition for videos.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition The THUMOS challenge on action recognition for videos

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.940537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.718104Z digest=sha256:fe260e5c07afecd9b190c09967057837fd8f4cd08f3516233bd6df31f4bd08af

Observation a7b9a5fe-129e-4185-9f99-899b1726cfc1 · outbound

This paper cites A survey on temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A survey on temporal action localization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.725629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.778064Z digest=sha256:11f343a1440b7ddf3feae811d779663c13a3be81f1c7ca753ed87c0635b87b12

Observation 9382a4f2-34cd-4395-8309-5206f9e61901 · outbound

This paper cites Tallformer: Temporal action localization with a long-memory transformer,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Tallformer: Temporal action localization with a long-memory transformer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.524109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.844680Z digest=sha256:04b7aab414286fa8ac98a5ecc99b98f93893ebb4dbb900a251114ef8e3e9c1da

Observation f3f02e0a-06f0-4cdf-b677-c39fa59aa05a · outbound

This paper cites Cross-fiber spatial-temporal co-enhanced networks for video action recognition,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cross-fiber spatial-temporal co-enhanced networks for video action recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.294764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.899195Z digest=sha256:49ca3e63e0d0d10471bae30d9d851581a9345cfbf954dd0f645a498cd4aeed33

Observation ac0b2a75-226b-45d7-b09d-38dcca55f0db · outbound

This paper cites Mutually reinforced spatio-temporal convolutional tube for human action recognition.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Mutually reinforced spatio-temporal convolutional tube for human action recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.192111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:03.980769Z digest=sha256:0fd9f5c01513b9d098b9294cf47e59dc327c7c50c61c7c1674a2b0c2f73e0efe

Observation 02d844f6-21ee-476a-8b3b-42c23c94aa4b · outbound

This paper cites Multi-scale spatial- temporal integration convolutional tube for human action recognition,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Multi-scale spatial- temporal integration convolutional tube for human action recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:07.050514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.066995Z digest=sha256:e3d8af0836bac27e791f6c64b9a04a657f139be06fa7398289c37996d638d3bf

Observation 567dc6b1-2132-407c-9605-a08dc09a600b · outbound

This paper cites Cross-video contextual knowledge exploration and exploitation for ambiguity reduction in weakly supervised temporal action localization,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Cross-video contextual knowledge exploration and exploitation for ambiguity reduction in weakly supervised temporal action localization,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.901413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.094400Z digest=sha256:9a89032ef2594a1364339624c7ecd1960b1788c7df4d1abaa3e8c975ed890fde

Observation 0f799125-56a3-4d42-bd85-2a1197b630a1 · outbound

This paper cites Trigger is not sufficient: Exploiting frame-aware knowledge for implicit event argument extraction,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Trigger is not sufficient: Exploiting frame-aware knowledge for implicit event argument extraction,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.793739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.157516Z digest=sha256:27accd99c8e56448db19c6fe4a0ef44e808bbca7b6afe2d16470719a102aa102

Observation d4beb4ae-e0ac-4bd4-80e5-35d32144f7e8 · outbound

This paper cites Guide the many-to-one assignment: Open informa- tion extraction via iou-aware optimal transport,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Guide the many-to-one assignment: Open informa- tion extraction via iou-aware optimal transport,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.618967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.215955Z digest=sha256:902728980d48c2455efce80632a05abf7119d476d6fff9d39009f749c2112635

Observation 53ef3b59-ce1f-4abe-b2b6-42c2904fb1cb · outbound

This paper cites Video activity localisation with uncertainties in temporal boundary,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Video activity localisation with uncertainties in temporal boundary,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.438970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.236450Z digest=sha256:38c76950a3a2578fc160a909d817ca73a8e8c16c932fab0fcb3fad54b309bd6b

Observation ea17fdfa-a670-4471-85df-f46710cca323 · outbound

This paper cites Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Exploring the reasoning abilities of multimodal large language models (mllms): A comprehensive survey on emerging trends in multimodal reasoning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.310617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.290223Z digest=sha256:2267744a07028811bb45be464270c437c61038af2e622a933b72cfe63e5fb4de

Observation 551a893b-fd60-4b3e-a6f2-1f062e4f7392 · outbound

This paper cites How vision-language tasks benefit from large pre-trained models: A survey,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition How vision-language tasks benefit from large pre-trained models: A survey,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:06.167646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.314652Z digest=sha256:83d2e4cb7bcce5ff6eb0df863d1c0b16487d57ea410a27c6791267ed78cdbf7b

Observation ad1ced01-72c5-4aaf-8cad-4b5b40f9fa1e · outbound

This paper cites From linguistic giants to sensory maestros: A survey on cross-modal reasoning with large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition From linguistic giants to sensory maestros: A survey on cross-modal reasoning with large language models,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.993465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.373024Z digest=sha256:66f4b5d83c68a4a48091dc9704c7d73e51cd5ef12c996ebe0b2409c583ac3b30

Observation f94019dc-ab73-426b-a526-9fd2d4d125db · outbound

This paper cites A systematic survey of prompt engineering on vision-language foundation models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition A systematic survey of prompt engineering on vision-language foundation models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.845189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.457379Z digest=sha256:0563d6fd1971d0f86e1515cf67e0e016f1de59a21fa5113f988223dde9048b29

Observation 950d90b7-9a6d-4b03-923f-f2ae9d0f4cc6 · outbound

This paper cites Evaluating multimodal vision- language model prompting strategies for visual question answering in road scene understanding,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Evaluating multimodal vision- language model prompting strategies for visual question answering in road scene understanding,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.669528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.548139Z digest=sha256:c506c5f8275492a482bb17e8e738579e3aab4b5ffe7ed867d2103e62b63b4af2

Observation 4b5364e7-7739-4df5-9ba8-15d560a207f0 · outbound

This paper cites Towards grounded visual spatial reasoning in multi-modal vision language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Towards grounded visual spatial reasoning in multi-modal vision language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.509382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.608059Z digest=sha256:4c3bfad950d3d55a9e3c1b1d15ad4a2d0576793517dce89fc1ea0d0d46f5ab5c

Observation 51bb8a7c-3808-432c-b3e3-95078126c0b1 · outbound

This paper cites Enhancing advanced visual reasoning ability of large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Enhancing advanced visual reasoning ability of large language models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.335787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.641263Z digest=sha256:2dac98d533d0329cc548ecc447c712d1b588fdd455bcc1d3cee74b1d3a22b340

Observation 3357097d-1759-4009-8866-21e72b58ae99 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:55:04.691360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:55:04.691360Z digest=sha256:efe0c647163b6f50162994c8d436b97feec9da5799ffaf748e316b26598d000d

Observation ce6576ed-c6cb-49b3-8ca9-b10ef9db3a17 · outbound

This paper cites Chain-of-specificity: Enhancing task-specific constraint adherence in large language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Chain-of-specificity: Enhancing task-specific constraint adherence in large language models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.221014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.745657Z digest=sha256:98c44db317793e89f30923afc28555ee7e3fdb64c5a3c679a403d2b280185bcb

Observation f4f25b48-e5c4-4bb2-bae9-b8d309399ca6 · outbound

This paper cites Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,.

Multi-Level LVLM Guidance for Untrimmed Video Action Recognition Lvlm-ehub: A comprehensive evaluation bench- mark for large vision-language models,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:55:05.054392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T16:55:04.858493Z digest=sha256:d246b47b6ecad108b298f297c3a34fd08f0bacc88b5b8c0d0014807a58a49344

Pith citing papers

No inbound Pith citation observations are available.