Pith. sign in

Paper Citation Record · LEDGER

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation

As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2505.01713.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01713 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:17:00.065856Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:55:44.834170Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T17:55:45.159103Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b23a21c9-dacb-4034-ad02-8fe9e7aa55fd · outbound

This paper cites GPT-4 Technical Report.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:16:59.916631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:16:59.916631Z digest=sha256:6a060a070e4d661772010bb35f779b719259a5371a01c533c7f5266be6f17838

Observation 45cdc215-036c-47dd-8de3-c187714b5a6e · outbound

This paper cites The epic-kitchens dataset: Collection, challenges and baselines.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation The epic-kitchens dataset: Collection, challenges and baselines

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.489441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.934262Z digest=sha256:d3a45cb9a0563206bfe4b1e1c01de41248b1c81aa88002a0e883154d5be6224c

Observation c903edfa-5d06-45ad-b073-80e193ab1f0a · outbound

This paper cites A technique for computer detection and correction of spelling errors.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation A technique for computer detection and correction of spelling errors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.478671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.937843Z digest=sha256:30b67e44a70dc9b00843708e0ee7401551fde46146a98fc161bd8ce4df02aad0

Observation bf4fbe5e-102b-4948-b6cf-40adccfb6ceb · outbound

This paper cites Rolling-unrolling lstms for action anticipation from first-person video.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Rolling-unrolling lstms for action anticipation from first-person video

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.468213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.945334Z digest=sha256:2f3d20bead3706cce7ce7a8df2dee50c9c9f532610c7113bc752d368aa337fdd

Observation 966d78ff-85f6-41c5-b0a4-f39ae2fa16f0 · outbound

This paper cites Future transformer for long-term action anticipation.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Future transformer for long-term action anticipation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.457225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.948823Z digest=sha256:c671ac500e9a570a1d3f4596d46b79b2d97bca7267a6e53ebb4ca16ff5f19019

Observation 134fc880-21de-436e-a5c3-d78ba8ae891a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation LoRA: Low-Rank Adaptation of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T04:16:59.955210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:16:59.955210Z digest=sha256:a5bc6465f16184c4f2a401f8b093c3031c58dfc1a946ecf26b81068094370755

Observation f1b701bd-79f4-4c23-a6b0-d0bc66b07881 · outbound

This paper cites Technical report for ego4d long term action anticipation challenge.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Technical report for ego4d long term action anticipation challenge

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.425374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.966646Z digest=sha256:7e05f723413cc5b999513afcf6c1fb17fd0fa2cb8cc8e145f3465d77745d1591

Observation b838ab20-a63b-4e46-81f7-e4fa068c7359 · outbound

This paper cites Technical Report for Ego4D Long Term Action Anticipation Challenge 2023.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Technical Report for Ego4D Long Term Action Anticipation Challenge 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T04:16:59.970868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:16:59.970868Z digest=sha256:03d11ba3422be1ecfea39e08d3e22a5c72f9da4e8e12d86198a2da101b032457

Observation 346cb39d-1f4a-47f4-856e-9d8bfb4f1ee6 · outbound

This paper cites Anticipating the start of user interaction for service robot in the wild.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Anticipating the start of user interaction for service robot in the wild

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.414572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.974858Z digest=sha256:3e1470ed1e74ac695d63b523726f59e8ae3771e1072cc5f369a0a9f96b210b30

Observation 493db213-0aa1-4eb4-867e-9e081f79fb03 · outbound

This paper cites Palm: Predicting actions through language models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Palm: Predicting actions through language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.403915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.978393Z digest=sha256:669f30d512e6c15e4bbe8a36ce2cb71acf99287dce4897ea8a6cca5b0ca53f3f

Observation c29b4694-c060-4178-a363-3202275621d1 · outbound

This paper cites Anticipating human activities using object affor- dances for reactive robotic response.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Anticipating human activities using object affor- dances for reactive robotic response

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.393245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.981857Z digest=sha256:98bef9a42b2041c7516a83193f67fb8141340fcba3b700ff496bafb3a25ac365

Observation e76c131d-b624-4bc3-9f5a-612d6ab8fbf1 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Llama-vid: An image is worth 2 tokens in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.358287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.992047Z digest=sha256:9a57e3b677667e4c84c747f84fdfda36bb8552e358fe7c30a7c25863d19dc307

Observation 54a044e8-3d9d-46ed-9577-4eee2b2bc33d · outbound

This paper cites Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video-language models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.334653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.998989Z digest=sha256:f43c942014bf092d025445dc46b40ae2c465251fef891e0f181c857425571e82

Observation ac04f3ad-eaf4-4502-96d5-190c0a1511db · outbound

This paper cites Ego-topo: Environment affordances from egocentric video.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Ego-topo: Environment affordances from egocentric video

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.323854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.002447Z digest=sha256:eb96d310a388e3dc0d1572cc1a46e368099ef67a2c8a36460b41d87d6c83faf6

Observation 4e59dfc2-f5c3-40e6-a55a-a4b39842175d · outbound

This paper cites Rethinking learning approaches for long- term action anticipation.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Rethinking learning approaches for long- term action anticipation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.312862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.006378Z digest=sha256:7d3d63e6f3e1a00335350d13185943aa9ba8fa4f81577bc96b5f0a87bcd898bd

Observation 2cc08456-6da6-47d8-82a4-11d7a56e6d06 · outbound

This paper cites Summarize the past to predict the future: Natural lan- guage descriptions of context boost multimodal object interaction anticipation.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Summarize the past to predict the future: Natural lan- guage descriptions of context boost multimodal object interaction anticipation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.301486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.010065Z digest=sha256:641b0954e07291f54353b71eeb36c68ca35b075e93a18de7eb64dd44e80ee86a

Observation 089ac1e6-e416-4e2d-ba63-2d0d867454b6 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:00.013657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:00.013657Z digest=sha256:3ffdbbc5154630476e30b2db03dd0c0a261c15016ba67032ff59ae74b7359e14

Observation 4702ef4c-8203-408b-9923-3c6b4558219d · outbound

This paper cites Learning transferable visual models from nat- ural language supervision.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Learning transferable visual models from nat- ural language supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:00.018207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:00.018207Z digest=sha256:cf7a96751fc7cbfbc98d1291452749ce076420814177572f4d817c34a4d9b21e

Observation 09e33535-a780-44d2-b0b7-5cc90c7db377 · outbound

This paper cites Pre- dicting the future from first person (egocentric) vision: A survey.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Pre- dicting the future from first person (egocentric) vision: A survey

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.282967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.022425Z digest=sha256:75900136225f033bef2070db23142330d88711be24df6e71f63f3d126bc881a1

Observation d3bb4032-a725-4d2c-8655-b13187b891e1 · outbound

This paper cites En- couraging lstms to anticipate actions very early.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation En- couraging lstms to anticipate actions very early

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.272008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.026392Z digest=sha256:afd2c3d5b8c6ff6c637b2526c430ee51a9ebc1722aee4c1702dfdeeb9949ef5d

Observation 16c508a7-dc07-4503-8466-14239cd0594a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation LLaMA: Open and Efficient Foundation Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:00.033906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:00.033906Z digest=sha256:cb4c1489e19025a5897cb34d29f4108386237dd2329cccc976129f5ef8aa8629

Observation 4cbdc48c-decb-4d83-b840-7caa03f667ff · outbound

This paper cites Attention is all you need.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Attention is all you need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:00.038188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:00.038188Z digest=sha256:8aa5e3b60c5e3a65c70f1ee865481c7c75eca8ee5371c3c6438d88dec0609025

Observation fb7b1e05-d3c1-452c-b982-e326b95fba8c · outbound

This paper cites Memory-and-anticipation transformer for online action understanding.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Memory-and-anticipation transformer for online action understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.241348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.042922Z digest=sha256:d8c410bf5e23d64871f94ca1bb7d2631cdd3d7e4e3a268b0ff25253259da6ca8

Observation bd421d6a-6661-470d-a005-533be7b31b9e · outbound

This paper cites Visionllm: Large language model is also an open-ended decoder for vision- centric tasks.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Visionllm: Large language model is also an open-ended decoder for vision- centric tasks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.229477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.046487Z digest=sha256:23f9b037d077bd163acb5f3920057184e92cb24a3bdf246981cbb0991dcc6861

Observation a4cae2e8-dbf4-46f2-8399-6c957c1fc5fb · outbound

This paper cites Black-box prompt tuning for vision-language model as a service.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Black-box prompt tuning for vision-language model as a service

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.217089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.050599Z digest=sha256:ab79b9cccbfe0c7c6f87b06a4f52d8bbf5e49f146d2cb0701d30a6d10c734679

Observation 0aacd786-8ec1-4f77-87aa-06406453995e · outbound

This paper cites AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:00.054405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:00.054405Z digest=sha256:986ad93c09c521e07c1e92091636c770689fd0fe007b4c3fb214cf702e1d8164

Observation 9b5082f4-f0f6-471b-a946-3db1f8ecd69b · outbound

This paper cites Anticipative feature fusion transformer for multi-modal action anticipation.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Anticipative feature fusion transformer for multi-modal action anticipation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.205522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.058814Z digest=sha256:b8cbfb987486f9153f5dfdc11560455461070b18370bc79fc876b18d1063fea9

Observation 380a38c4-06a5-40fc-9484-30e59dd93977 · outbound

This paper cites an unresolved cited work.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:17:00.190975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.062338Z digest=sha256:88551be1924bbf0a480b9e22627f3f0c895ce8f9a51b4ec2c1fe44e89ddbc7a7

Observation 9f27aa03-8246-463c-ba20-1948ea37e648 · outbound

This paper cites QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T04:17:00.065856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:17:00.065856Z digest=sha256:2912ea45bdc39f67454b0e670a5ec88ee53313b6707b3160536b8c48784626b5

Observation 47abcfae-6f50-438f-9b37-0320acd41712 · outbound

This paper cites The Llama 3 Herd of Models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation The Llama 3 Herd of Models

Reference 1964

Resolution
unresolved
no resolver link, observed 2026-08-16T04:16:59.941847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:16:59.941847Z digest=sha256:98370cddb312bd44ace84e641108eeba36352a7f6fbf9c1776e0bf7df63f6f7a

Observation 949ff927-2b61-4868-908f-17012bb77f5d · outbound

This paper cites In the eye of beholder: Joint learning of gaze and actions in first person video.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation In the eye of beholder: Joint learning of gaze and actions in first person video

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.381374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.985408Z digest=sha256:2de04545d3d6764c9f2048dc9d45c2157f2a5954a9b2ebaf6ca2be36366be118

Observation aacadb1e-17b2-429d-98c5-224a8eeea937 · outbound

This paper cites Temporal aggregate representations for long- range video understanding.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Temporal aggregate representations for long- range video understanding

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.260535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:17:00.030334Z digest=sha256:418c3e585b95e24b8191c2a7eacbc9c3845a45e3a1437ad6efc271c48aebdc1f

Observation ac976ea0-fd66-4dfc-a8d4-9caf30491c62 · outbound

This paper cites Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Blip-2: Bootstrapping language-image pre- training with frozen image encoders and large language models

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.370295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.988587Z digest=sha256:695c1b0d2e14b95dc0560a55b7ba9b743359d4e86f283453c60d32302c00f042

Observation c9658f80-6d31-4319-83dd-b9d99b0cb6ed · outbound

This paper cites VideoGraph: Recognizing Minutes-Long Human Activities in Videos.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T04:16:59.962650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:16:59.962650Z digest=sha256:9f7bc4b7f08fea29a5377d0e240f5b35739088b41aa7f5cd11e38c7cea6ffca7

Observation 4bec848c-3850-4325-bb76-ae7419bda8a5 · outbound

This paper cites Sgdcl: Semantic-guided dynamic correlation learning for explain- able autonomous driving.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Sgdcl: Semantic-guided dynamic correlation learning for explain- able autonomous driving

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.512190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.926640Z digest=sha256:82dccfa6a99ad836651907223d7a10454e3f5976fea3c3f4da6d34579ba81075

Observation 29be6859-98f4-457e-bdce-8813f3155c1e · outbound

This paper cites Using gaze patterns to pre- dict task intent in collaboration.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Using gaze patterns to pre- dict task intent in collaboration

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.436667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.959028Z digest=sha256:b0a9871ae8d63b13b4db74f8a2fd0ce89e4830f834a35c716a5e0b4f6a854148

Observation 6ccbc618-a512-4251-bba8-0a34906fb4d4 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Ego4d: Around the world in 3,000 hours of egocentric video

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.446800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.951867Z digest=sha256:3b7f9a7114080f1e97f95bcbc78bb9d820e6802d3db5a2e2b9a272bf12efd21d

Observation 84f25e48-fe3f-435f-b837-06d6e38768f7 · outbound

This paper cites Language models are few-shot learners.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Language models are few-shot learners

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T04:16:59.921983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:16:59.921983Z digest=sha256:8b55b96b1a8de184d917da1292a70e6e1f67324f8c1777fab669fbb8e84c4cfd

Observation 79d26287-c79d-4db5-9add-acee4bed298e · outbound

This paper cites A survey on multi- modal large language models for autonomous driving.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation A survey on multi- modal large language models for autonomous driving

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.500801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.930383Z digest=sha256:1989cb112328d095e0a7cfec49a2b8988a88bfc8b36966c44bca52c8a659398a

Observation 0b1b8fc5-8d06-4414-b6b2-03dd8cd766a8 · outbound

This paper cites Ego- centric video-language pretraining.

Vision and Intention Boost Large Language Model in Long-Term Action Anticipation Ego- centric video-language pretraining

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:17:00.346589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T04:16:59.995283Z digest=sha256:139d96ddf6319644b5026d2f0fa0cac674cb44718b2adfddc849f0d1a5b8c44a

Pith citing papers

Observation c57a096f-e4b6-494f-b811-0b43be3e0627 · inbound

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition cites this paper.

Compositional Benchmark Synthesis for Hierarchical Human Action Recognition Vision and Intention Boost Large Language Model in Long-Term Action Anticipation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-12T17:55:45.166090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-12T17:55:44.834170Z digest=sha256:b483f7e6c3888884fbbd46b820a375d2b01a9b34f23577067b630fe52bce944b