Pith. sign in

Paper Citation Record · LEDGER

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.15130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15130 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:44:49.592008Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T11:39:15.308355Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T11:40:03.350024Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact6
  • verified fuzzy34
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1ef1572-a27c-4efd-b6ee-8a89deccadc8 · outbound

This paper cites When will you do what?-anticipating temporal occurrences of activities.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction When will you do what?-anticipating temporal occurrences of activities

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.173046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:43.981738Z digest=sha256:4441da6e2245cd9022043b69e9879ac748b19403954a1097484b9b9b9d669554

Observation cb381446-e785-4243-9a8f-b90d5f812668 · outbound

This paper cites GPT-4 Technical Report.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.065986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.065986Z digest=sha256:cd011ba43cd16c21678a3a23948aa50128dc05b75f3874062ba6e6ce3916a127

Observation 72238229-4323-4029-bb15-8fa4759981fb · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.155587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.155587Z digest=sha256:aa824c045400638b438a9fcbf8bd29add62d7c7c0311876604cc8dca9467c6f0

Observation fff40a00-3055-483b-99c0-73bf8ca694e5 · outbound

This paper cites Hiervl: Learning hierarchical video- language embeddings.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Hiervl: Learning hierarchical video- language embeddings

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.161138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:44.281924Z digest=sha256:60e3d2f9bdf3b65b918bf14020f769b64ff39ce74abc2bea54da99d6e971d7d7

Observation f9f18253-afc6-42b3-9605-c383a5d72b26 · outbound

This paper cites Procedure planning in instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Procedure planning in instructional videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.153750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:44.356600Z digest=sha256:6a6cc14ee97d13e97318f369df1d603ff4415fe7bed011f80db426f1f56ed2c7

Observation fbf579ff-4efe-43d4-889e-b9e8957a91b8 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.146331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:44.497438Z digest=sha256:e72c02a23fbff65033c1cf23d7be87942739ab8f73194dc20cd6ecc3cdd8c025

Observation b1b39347-9368-4e03-9849-e3285bb02838 · outbound

This paper cites EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.617672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.617672Z digest=sha256:78a7ea7862b3ae99555d2fbcc730682bb02dba07bcf4a8a11c0796c746adc29d

Observation 92f198ec-ee75-48ac-8f2a-66d84f03eb99 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.712467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.712467Z digest=sha256:e6f6d644fb7640ab14e4670416be924cc488094bb7006ad611f54ece767b2596

Observation 6c23d4da-96e1-43a4-b120-7c4db34acdaa · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Better & Faster Large Language Models via Multi-token Prediction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:44.806734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:44.806734Z digest=sha256:1e79c05e02471638aa15b2b1e448cd5653134203f26018771c286b0c4f6757e6

Observation 9b8f4eb0-eead-471e-a8b0-4397c62e683a · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego4d: Around the world in 3,000 hours of egocentric video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.138897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:44.914759Z digest=sha256:2b792d6f6a51f4def686573470b89153718e2bf2af255c9da64d0cc57d183bf3

Observation 1e7b96ed-19db-489a-b249-63f02ce724b2 · outbound

This paper cites Reasoning with Language Model is Planning with World Model.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Reasoning with Language Model is Planning with World Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.036118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.036118Z digest=sha256:5dca513e0103b624bb1f10f5aeaaf1959cc0a9e7346370ed4c59d845afcd6ce6

Observation 66bf25b6-aa1e-4d23-82b3-1a50ffecafe2 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LoRA: Low-Rank Adaptation of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.143895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.143895Z digest=sha256:6abf75134cafdb2ec925b1aaba44d440c9903bd2ad0309cc42999020c4435302

Observation 3231d02b-bfad-4a3e-9862-4454351e6988 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vtimellm: Empower llm to grasp video moments

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.259004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.259004Z digest=sha256:9962c612e121f6f8494af4d26986ded25b606edadb834e4e7011ce5dcfd7ec9f

Observation b452a7b1-8d21-4b5b-a131-554a48270480 · outbound

This paper cites Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Language models as zero-shot planners: Extract- ing actionable knowledge for embodied agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.066608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:45.361973Z digest=sha256:42c084dcc12f25e0b970f2ddd15e34adc480a5f964301791b729c021f3f4449d

Observation 5ebe8f6a-baad-455f-a467-80ef575e3610 · outbound

This paper cites Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.792034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:45.482849Z digest=sha256:fed85029ae674da7f176dda332c9fad1ace41732b2f2a018fd8e297ba6b23ebe

Observation 0f03c10e-7a52-4b22-a0c9-7d99c64ba0c9 · outbound

This paper cites Palm: Predicting actions through language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Palm: Predicting actions through language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.059681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:45.617947Z digest=sha256:466a0a1226d9f66b2e2244c3ed073b7effdeb6d8ee68f9258cf1545e6cc252d4

Observation df76deea-f9a7-4c2b-9cea-6dc95f0b37bf · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.736431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.736431Z digest=sha256:56a0c5acbf5c938de4d0a08a26cce4eb3befa79c3ca9f56eb5b813ec1f8317e0

Observation bbd8abc5-45e6-4a6a-9c4d-1e2ad8ac49ec · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.052622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:45.818880Z digest=sha256:13b77da5547ea295206e07b00346f3ad66d6508cacd8343cc686ddd30b7baf4b

Observation 010fbb76-6cf4-4479-82d7-62955d2fc731 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:45.907158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:45.907158Z digest=sha256:1a104e1a2610357a7fe930e199526112ee2bc92b8d91908bcdef5acb38cd7462

Observation 8e2c3f67-7f3d-463a-ae90-eab374619d4d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.045748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:45.981792Z digest=sha256:67de50c430068427edd65ee73ebd1d183fd4f3febf3e592bbe54c4e294945519

Observation 83e61c2f-2559-4abc-bed2-2b054d6e6cc0 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama-vid: An image is worth 2 tokens in large language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.038452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.075955Z digest=sha256:3831c423abf3c014638dd619512ad8ac96ef6a351b89fadb267c85aa4f00ebb4

Observation e32811d9-557a-45dc-8a30-0dcf25808e09 · outbound

This paper cites Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Skip-plan: Procedure plan- ning in instructional videos via condensed action space learn- ing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.031304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.185631Z digest=sha256:549ba94d4a8a47f1d6a2543bf5a7199dab35c111bba2bd3c5fc13cf2de901033

Observation 5d031845-fb8f-4d37-83cd-620bed5edc2f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:46.272623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:46.272623Z digest=sha256:ee88f4594666fc8b2b08ae5c0fbec5ec9d7b1a363c03c958c206fbd0cb9a2e2b

Observation 1826b2e1-b5d0-4166-95ad-9199fcbaf75d · outbound

This paper cites LLM+P: Empowering Large Language Models with Optimal Planning Proficiency.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction LLM+P: Empowering Large Language Models with Optimal Planning Proficiency

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:46.355848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:46.355848Z digest=sha256:0fd9d07fd63467cc5782c625ba161f968943ef0ff6924f7db2e7ae80309e20ef

Observation 0d423aea-4a84-411c-8400-7a297d87cd58 · outbound

This paper cites Visual instruction tuning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Visual instruction tuning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.024103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.434456Z digest=sha256:0735a687016dcd2c971b9db9958b0a93f67f4cc8160ae021127a7ca4b4d545f0

Observation 19b722b5-fe3a-422c-807c-cf83d362d10d · outbound

This paper cites A language-first approach for procedure planning.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction A language-first approach for procedure planning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.017054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.528710Z digest=sha256:83bc60828ba9a8a4631cee50639f38c0f2b6d55cb8d9da430fb1424c8ce7810b

Observation ec54eccf-757e-4470-a178-8c20ac2ca022 · outbound

This paper cites Intention-conditioned long-term human egocentric action anticipation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Intention-conditioned long-term human egocentric action anticipation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.009663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.610609Z digest=sha256:35d0166cc34e7adcd6def8b35d25784d21efb8de5c4722a69e627d92f5a3acf1

Observation e70b9c84-0c3a-4466-b36c-262583ebbfa4 · outbound

This paper cites Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Can’t make an omelette without breaking some eggs: Plausible action anticipation using large video- language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:50.002240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.680605Z digest=sha256:223623c10bcfd86cb61ac4f70addb0d637ebc256d9df24e048f5ddddf5223b73

Observation 7e57bf75-9560-4cb3-a881-8c99d592fa92 · outbound

This paper cites Any- mal: An efficient and scalable any-modality augmented lan- guage model.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Any- mal: An efficient and scalable any-modality augmented lan- guage model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.994645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.734541Z digest=sha256:c01cdd3c1806abd78391dd013963b519b5fab5ec72a119edbe7bd0647dd2b8bf

Observation acc44834-e809-4f34-9a37-ff9452287226 · outbound

This paper cites Embodiedgpt: Vision-language pre-training via embodied chain of thought.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Embodiedgpt: Vision-language pre-training via embodied chain of thought

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.986211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.817312Z digest=sha256:f91dd36cbe504937dad42a2c46e3d6bb7397e2734eceb191f1849d94c221c711

Observation c0cd7f5f-bff3-49d8-b818-96adb1df2653 · outbound

This paper cites Ego-topo: Environment affordances from egocentric video.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Ego-topo: Environment affordances from egocentric video

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.979003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.889917Z digest=sha256:9d55b46010e6d10965ce2651020669effaf6ffbabee2f0ef3f821acf9c5b761b

Observation b8de1b13-6e6a-4cc0-b53e-59a503d393c0 · outbound

This paper cites Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Why not use your text- book? knowledge-enhanced procedure planning of instruc- tional videos

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.971464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:46.976969Z digest=sha256:726450d1a354dfddf9afe6947e543253bfb4c7ddabb53dc3bd5d9fac5a055b57

Observation 4d55b9dd-512e-4f59-9f33-0bc19cd72ffc · outbound

This paper cites Re- thinking learning approaches for long-term action anticipa- tion.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Re- thinking learning approaches for long-term action anticipa- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.964005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:47.046680Z digest=sha256:9cc48c288220940e5329e67ddbbe935d9124e724cc14e22ca0a41de6bbe0afd8

Observation 003f8245-b33f-43fc-8cb8-f28d80703fbb · outbound

This paper cites Do Pre-trained Vision-Language Models Encode Object States?.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Do Pre-trained Vision-Language Models Encode Object States?

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.751802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:47.143277Z digest=sha256:71ca26854b4f2ac2cab694b5a63479527f6a089fa57fa69fbc118da96011f23f

Observation 23e267b8-7b9e-44fa-ba12-6ecdce73340d · outbound

This paper cites SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction SCHEMA: State CHangEs MAtter for Procedure Planning in Instructional Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.232091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.232091Z digest=sha256:42cf04ed345171274a9c1d6ed7c0afeef241c47cd9ebaf244ae36ea4c2fa63ce

Observation eb7d0d66-827a-41ca-975d-84c6678d8883 · outbound

This paper cites Pretrained language models as visual planners for human assistance.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pretrained language models as visual planners for human assistance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.956596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:47.334176Z digest=sha256:7b5ddc3a572480ab10c33a73eaecb83d689452c62743f1f7c84fa3de6cd60ae6

Observation ad121c6a-9f47-4e65-820a-c8f927d7c9d2 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.422726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.422726Z digest=sha256:6dae611f5cdf666da073f336f7b7022d322746ede73d8ad28d4f1545b1f558f1

Observation 212abde4-2bc1-42ed-82a1-c02e251bec32 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.520419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.520419Z digest=sha256:fb5625235ebbdf9ee82cda6abee54b636b7327f24a5d4da458a15e59b01ced68

Observation 8d5e895f-1c46-43d7-9700-6300edf4e55b · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.619011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.619011Z digest=sha256:4b09a2351b2c4ca101dd1cee942dcff3f327067eff8ae4d1ca6997f6a72dbdeb

Observation 9f2e9711-22f4-4fe6-98c5-e186b4b5e058 · outbound

This paper cites Llm-planner: Few-shot grounded planning for embodied agents with large language models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llm-planner: Few-shot grounded planning for embodied agents with large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.944499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:47.715754Z digest=sha256:732212a82bf2852c51e010391958b81125369e61d5fd36265b0d9e2cd40aafee

Observation 0332c79f-4bd2-4097-a37d-4d215e106510 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Moviechat: From dense token to sparse memory for long video understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:47.800874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:47.800874Z digest=sha256:fdf8b9b3cfe4a607d043cfc17692e5fe0c141bd057ec87887a362c815c8b93c0

Observation 5d54eda3-0659-45a6-9c88-e1e0f13aaf97 · outbound

This paper cites Coin: A large-scale dataset for comprehensive instructional video analysis.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Coin: A large-scale dataset for comprehensive instructional video analysis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.932602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:47.933683Z digest=sha256:0eb3d246630611b01c39b66b0b3de1968f4fb3d53bac399230176e47c4b5cc7e

Observation f219ca49-d55a-4d09-96f2-08470039b654 · outbound

This paper cites Learning Multiple Object States from Actions via Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learning Multiple Object States from Actions via Large Language Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.717370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.030903Z digest=sha256:68108dfdd18d819bb6d59e9c9934c0767a5be29dc6da8e04dcfc68b1c0491d7a

Observation 5f6f2c71-ab4a-44a8-a5de-3d2bf337fcc5 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:48.175250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:48.175250Z digest=sha256:2427926381aa4a35fc6ee376cb05d16b647a9d2103f51d738aa275fc6999976f

Observation ddc243e0-b54d-4218-ba0d-974bfd281f57 · outbound

This paper cites User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.697709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.352122Z digest=sha256:078796a1a06d7fe9d18d31d4a65a4d3a6c588934afd2354de02c8e0eacba62f5

Observation 3b647031-8ff2-4b39-9107-29dbc407fde3 · outbound

This paper cites Event-guided procedure planning from in- structional videos with text supervision.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Event-guided procedure planning from in- structional videos with text supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.925468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.507610Z digest=sha256:cf0a5b6a8bc2b1da418d38fbbb8289f2f6fb0985d50463a90519a87d55c8ea6e

Observation 7bd79037-3e2b-4fd2-afcb-c52500d9e309 · outbound

This paper cites Pdpp: Projected diffusion for procedure planning in instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Pdpp: Projected diffusion for procedure planning in instructional videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.917737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.650999Z digest=sha256:bdf019dae175305aa8efaf12fa467c1172e807e28611c76e2f1a609bb05d4d22

Observation 9cb4e834-73d3-49bc-aaee-44abd6ea5728 · outbound

This paper cites Vamos: Versatile Action Models for Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Vamos: Versatile Action Models for Video Understanding

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.687256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.688715Z digest=sha256:02a2ca8d64ceed6cd7b01a67de6bd26237b9d129bfc197eb91390d733f3eaaff

Observation a088458d-7634-4f1c-8ca8-51e26ea5be01 · outbound

This paper cites Learn- ing object state changes in videos: An open-world perspec- tive.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Learn- ing object state changes in videos: An open-world perspec- tive

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.909513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.748675Z digest=sha256:32af529409201e4d5040d6b1c2ba8a2201df893c143e955851c6b526f625deea

Observation 956f8e27-856e-4355-936d-0666a9808099 · outbound

This paper cites Octopus: Embodied vision- language programmer from environmental feedback, 2023.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Octopus: Embodied vision- language programmer from environmental feedback, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.901873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.801693Z digest=sha256:140ef39900df3205aa6e5b537b18d6e0f014527f4c35d98251960d752fc23bad

Observation 1d5423c0-02f8-49f9-9e76-ff78bfff4087 · outbound

This paper cites RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:44:49.675996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.842813Z digest=sha256:22b52be23954b76306f164a1dd970a06f5f269c57063eceb00871a6f8b8b0553

Observation a03f4ecb-e404-4648-bf90-07472dcd2fc3 · outbound

This paper cites Object-centric video representation for long-term action anticipation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Object-centric video representation for long-term action anticipation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.894247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:48.910963Z digest=sha256:6e7a8049e32d3ba81afe7a3a9482527944edc0784ded7a51460b5c5bd3416116

Observation 23a278c6-bd3f-43e7-ab9b-259c65b1eb56 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:48.976855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:48.976855Z digest=sha256:b0f073b9f64861f9de4ad4271f01b3e652e1f53131295d264b291b1912560da7

Observation b78c2393-32b0-4215-884f-cf024c8db353 · outbound

This paper cites P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction P3iv: Prob- abilistic procedure planning from instructional videos with weak supervision

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.886095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:49.021366Z digest=sha256:15994f8c89b5ae8b9781327207901d83b1dd381151706468d51b0e50d44796ee

Observation 70571645-9ad2-4347-bea0-571aa42e7e8b · outbound

This paper cites AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction AntGPT: Can Large Language Models Help Long-term Action Anticipation from Videos?

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.029086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.029086Z digest=sha256:88718456ab34f394c84a2d3ba2a515488527ed6df72ba3f13a1330f3ff8a97bd

Observation 653059be-6758-4f8a-ad93-a510d079ded0 · outbound

This paper cites Towards learning a generalist model for embod- ied navigation.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Towards learning a generalist model for embod- ied navigation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.878331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:49.034663Z digest=sha256:644a2355a9eb469973a8b0ecc40d26b69532d6c7fcc59ef304e735ac8a855c0d

Observation c8f553c1-b4d0-4bc1-9634-e562077195be · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.120307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.120307Z digest=sha256:173f0ec7571f5d34a01b3de3a32ed6d94fbdc03cb7988fd3d91508ef817d7ce4

Observation ed1535af-4470-4e8e-bb0b-5efafbe389f7 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:44:49.190266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:44:49.190266Z digest=sha256:dc5036a0cd208eabed6f6d57942e052b3bcade0e184521dab67de67e7132597d

Observation 58977a31-b981-4e08-ae42-3f13da79417e · outbound

This paper cites Cross- task weakly supervised learning from instructional videos.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Cross- task weakly supervised learning from instructional videos

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.866298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:49.288340Z digest=sha256:61a14e778fec9e70812df9e05b3eb8c58d431ae260d7fa2e75551caa666ba6f0

Observation aa874697-635c-4d25-8905-c864c4106a8e · outbound

This paper cites before” with “after.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction before” with “after

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.858942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:49.417662Z digest=sha256:a5f1d280a3a109bb9ad470674f8bce6c7a83e387922b6fab653669f106bf5cb9

Observation df875fdf-dfaf-490e-b2ae-bd12d17b3fe7 · outbound

This paper cites Training We train our model for 1 epoch with a batch size of 1024.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction Training We train our model for 1 epoch with a batch size of 1024

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.851659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:49.495928Z digest=sha256:99001252adc064a5c7a4548c974979e62127c58c60153d8d5f611c90d8dbfbd2

Observation 889eb96f-8bc6-4f6a-b0ec-b5d88ce41d6b · outbound

This paper cites dough”, “container.

Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction dough”, “container

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:44:49.843986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T15:44:49.592008Z digest=sha256:53be88f5e13613c875cd13577b984f21ce401b32ca3a45bdb13fe9fb12309249

Pith citing papers

Observation 115e9f3b-1d83-4ccd-9d0b-7f6cbd293e61 · inbound

GeoWorld: Geometric World Models cites this paper.

GeoWorld: Geometric World Models Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:40:03.352082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T11:39:15.308355Z digest=sha256:a02d4653dfc43118dfba96f631a835278ca1b7c62e27e903a48d5f80c65fb031