Pith. sign in

Paper Citation Record · LEDGER

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 9 inbound Pith citation observations for arXiv:2506.09987.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09987 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:42:37.838863Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:56:48.420938Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T10:34:36.459072Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e1d51c16-d587-425b-8a09-5d466e971a4a · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:38.021779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.791497Z digest=sha256:0066e1b98bbaa85c5f6c327612fc8bfc6aaa02e195afcf711b47b393084aaa00

Observation b41858ef-cdc5-44d6-99b9-31afd1f77cdb · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:38.013541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.794495Z digest=sha256:c16d976bcd77de5506c2657dfd9d334b1257afcbabdb0644544e5c2dd725e3c7

Observation 66c2b1e8-7371-4b74-8cc5-df5c8aeeaa2a · outbound

This paper cites Both the other options.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Both the other options

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:38.004229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.797183Z digest=sha256:dcc66976cd22ce652dab35db68dcef4d0d06ddbf49e513d326e0fd000c882bdf

Observation bd009d75-00f7-4b7a-b8c9-29b293f3d716 · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.895789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.833710Z digest=sha256:9a594ceeba3e6aa3fa33f0b2b1a22cc720dbfa7bd24553fd61b958b647fe0d70

Observation c6261c81-d004-4762-a5f9-8001f49cfa31 · outbound

This paper cites red triangle.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs red triangle

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.995692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.800448Z digest=sha256:717e851a5ad79638584cd60a528e1a1f0cc390b129c8bf39b8e6077bc9f09644

Observation f99e1750-b010-420e-b9d3-d8970378befc · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.987394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.803395Z digest=sha256:5e367f12912caa8b71d1152b0b012e9d0420a36277d1b7edc918860dfcadf4af

Observation 69ea94b9-b263-472e-9366-353ef21944d8 · outbound

This paper cites Move yel- low triangle to blue heart.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Move yel- low triangle to blue heart

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.979713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.806541Z digest=sha256:739dd7d5558205f0ae5c22c85a71625d71b2fe33d03395f98b8ad7fb3550324e

Observation 005c38f2-f694-44fe-9342-861277fdb40b · outbound

This paper cites spinning something so it continues spinning.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs spinning something so it continues spinning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.971538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.809225Z digest=sha256:9edc4054bf939222cdd0e48ba12e561f2208cbaa6766be9be51f92ed167d58d4

Observation b72bd6c6-65e7-4901-9a8f-65bf87f22f33 · outbound

This paper cites If no pairs fulfill this strict criterion, we relax it such that only one object must overlap: P ′ ={(x i, xj)|obj(v i)∩obj(v j)̸=∅}.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs If no pairs fulfill this strict criterion, we relax it such that only one object must overlap: P ′ ={(x i, xj)|obj(v i)∩obj(v j)̸=∅}

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.963363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.812436Z digest=sha256:1d1706f5257d38639be9c698a1f379cb716e6f5ba4db6f3124d3935fdd3071e1

Observation 3f9c877e-4d02-463e-a496-bc3020564916 · outbound

This paper cites ,(xik , xjk ) sim(vim, vjm)≥sim(v im+1 , vjm+1 ).

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs ,(xik , xjk ) sim(vim, vjm)≥sim(v im+1 , vjm+1 )

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.955040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.815170Z digest=sha256:2e0a4654798fd96bc5f1a8d50a99effa576f173f76d7fcf703fe33a2bc97da6c

Observation 6e3998be-cef0-4861-8f75-4d4d4d98171a · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.947008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.817732Z digest=sha256:ebc52932a8da74c0e86a0648adfbf438c71c16d8c382a911f8c5029ae6d4fee0

Observation 8e81b4fe-efa6-42bf-8555-74848e2652dd · outbound

This paper cites How many objects are moving when the video ends? A) 2 B) 3.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs How many objects are moving when the video ends? A) 2 B) 3

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.939002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.820670Z digest=sha256:4136771e976440723b7d3351bbb7b713bf057ea155e22e4a879f2406d337b997

Observation e75a8ed1-44d9-420d-bf2c-b7bf1387efa0 · outbound

This paper cites fuzzy subset.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs fuzzy subset

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.930499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.823555Z digest=sha256:5b6d00b57fabf8d43d01ae7889413fb5db7a0a5476960b8f7ddb13d2834c9dac

Observation 702d5902-e00f-43fa-9fe8-30291a4df7e4 · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.921742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.825928Z digest=sha256:c3249dc7f26077774def9bbebfb34403078cf157822c2677d3e4091b5454705d

Observation babc19e3-831e-4baf-bcdc-c15f13d6c74a · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.912490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.828749Z digest=sha256:54a5ee07b8b440ba5e6c621391f8af46d1b6154abf1fbc2a74f0124556c98fa3

Observation 6818c665-9438-4ae7-a293-f15165473c8f · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.903908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.831292Z digest=sha256:3eed0cc0885950b72e70d07de1b5f1e3c22ba8b77e4fa16a203bb4116863a75f

Observation 0134955b-604f-4347-8bc9-54b88e603cb6 · outbound

This paper cites In order to push the field further we are now asking the models more and more nuanced questions, and the answer may lie only in a short span of a less than second.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs In order to push the field further we are now asking the models more and more nuanced questions, and the answer may lie only in a short span of a less than second

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:42:37.886812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.836507Z digest=sha256:f7eb701aacc7fba9ca737a4bbccd2c6b694c857ab30b18c231028cd9e7972fd3

Observation 3dc7ba2b-2020-4a9b-9dca-c64f364d549c · outbound

This paper cites an unresolved cited work.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:42:37.876885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:42:37.838863Z digest=sha256:d090989b760a9bdeab89f176a35a820f594198addba668da668740290a7cb39e

Observation 19cd38e0-24ab-4abf-a6f4-3b2267239606 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs LLaVA-OneVision: Easy Visual Task Transfer

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:42:37.786681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:42:37.786681Z digest=sha256:0f25f70c24c81ac16a816ca9e6e4c258de28abe729993717b8bb6247ecc5444b

Pith citing papers

Observation e6b2e4ab-88ac-49fb-8b66-3b16879d33fa · inbound

Embodied AI Agents: Modeling the World cites this paper.

Embodied AI Agents: Modeling the World A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:28.130255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:28.130255Z digest=sha256:f694d771e27a0232c3fbf848e1ffa963bb823268f88ea9172c3c4b4894890f01

Observation c04fef8c-e3c7-438a-8f02-d83742c4efab · inbound

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction cites this paper.

VisPhyWorld: Probing Physical Reasoning via Code-Driven Video Reconstruction A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T11:11:27.528157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T11:06:38.102348Z digest=sha256:bb41caa0348bebfb918b5093e2d748b9d7d5e825226129c862456b2a172ba908

Observation 1d155d78-ff93-41bd-a425-c7d5f8c5fd6f · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:31:26.671992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T04:13:21.487431Z digest=sha256:d2e6440ba3b416bfd4ffa588c7b4ee7ff21e1293f1daa13f0485e08f4e4c0fd1

Observation 96a3f3a5-4152-47f6-a1d6-735db32b4ae3 · inbound

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models cites this paper.

TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:57:28.342496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T06:53:42.726350Z digest=sha256:d6c71068b3815da407bb6290bd49171dbb98f89d128a8fed934f162f6ecd854c

Observation 7a01867c-a937-408a-8b7d-9184f6135dba · inbound

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning cites this paper.

Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T07:46:14.800731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T07:45:56.473188Z digest=sha256:ea9986b20be792b744e0665fa63c29b81e174a30a6629de0f1f4c9bdf749b2d4

Observation 3f6b82da-50f9-4164-84f1-77cc48a726ec · inbound

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis cites this paper.

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T07:14:42.783296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T07:12:02.612292Z digest=sha256:071758f1b19fac9b60c43079b6b618f22e7b2d566f6c1257d428966d7d6b0c43

Observation 2a6531c2-e3a3-43da-8c2f-ded6c2f038b7 · inbound

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position cites this paper.

How Should World Models Be Evaluated for Embodied Decision-Making? A Decision-Making-Centric Position A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T10:34:36.460645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T10:31:37.108792Z digest=sha256:55600ac19f086f05b568fd291d20434e4691407cb8380abddc8fb55ea458e27f

Observation 4e1e5bf9-0685-4249-9df0-f4c04de7cc11 · inbound

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning cites this paper.

AdaThinkV: Adaptive Thinking for Token-Efficient Video Reasoning A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T17:24:16.066057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:24:16.066057Z digest=sha256:bd864fa29e55dd8dfeb0f57e8729ec144a3232e4b842776ee7a30874cefcd2ad

Observation 74fed687-0104-4443-a078-658533090feb · inbound

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding cites this paper.

FADE: From Passive Verification to Active Discovery in Counterfactual Video Understanding A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T17:56:48.420938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:56:48.420938Z digest=sha256:87d50ea0e7909ebee9d169b37b5e12e767175a06def5ec8ffe7b38b649bfe45c