Pith. sign in

Paper Citation Record · LEDGER

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models

As of 22 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2509.08270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08270 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:58:34.564196Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy21
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a919164d-d87d-44bd-845f-800b766b4de9 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:40.015509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.228497Z digest=sha256:1fac0a98dc085c422b5ef268843ec43ee4a4fb8f985a0ab9b06a91c8a360dc87

Observation b2020f27-e0fc-40fa-ae50-86b8a54b45f6 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:39.678327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.318689Z digest=sha256:e95bc1197ecb5e6a26e3f04feff30b7313445a2835fecd85c26529e7772925f9

Observation aa6152b0-6786-4d9f-88fe-d14e16c690f0 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:39.369301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.432981Z digest=sha256:4edcf223f51eab3162ee1fdd5639a40a53ee4b21b9bdfb1ef38de4e423bb1f9b

Observation 1cbb1d13-365d-4068-be35-5226e324960e · outbound

This paper cites Phyre: A new benchmark for physical reasoning.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Phyre: A new benchmark for physical reasoning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:38.992192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.506855Z digest=sha256:2a259c875e3e9b8e03036c3637bbe45292a4963287be818f495045884610f1e9

Observation b4246a39-eb68-4527-a67d-0a6b64c8ba6c · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:38.670472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.598914Z digest=sha256:87e4625510a89c988fa46a70a8cb82d58a4f5652cfb7cafac07100df9e348a7f

Observation b41b0920-12b6-4376-a9b9-9d9e616e445a · outbound

This paper cites Patil, Peter Clark, and Wen tau Yih.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Patil, Peter Clark, and Wen tau Yih

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:38.452080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.696044Z digest=sha256:a5763d2e5f421cfb1da749d3e9bec7cfb48a3c60b72bcc0e8f50c82c3099f252

Observation 8d5db976-eb74-4c7a-be1b-5f0bd70cf17d · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:32.782839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:32.782839Z digest=sha256:9ec32bde03179f1b1a10d386d731fb0a3d7848f0c88f1759b7ea7f1f46e0d043

Observation b9b26cd4-26f3-4ea1-916f-b3d17118d50f · outbound

This paper cites Training verifiers to solve math word problems.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Training verifiers to solve math word problems

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:38.206663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.895078Z digest=sha256:62191bfd94e57896f28ba9d34a0a0ac7208a8b7bd052ab9711bd8b176f8c7066

Observation 005e84d0-f5e9-4bc4-a55b-a8f97d7ee731 · outbound

This paper cites Construction of a Surrogate Model: Multivariate Time Series Prediction with a Hybrid Model.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Construction of a Surrogate Model: Multivariate Time Series Prediction with a Hybrid Model

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T20:58:35.119599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:32.956056Z digest=sha256:bba4486f9cbd77d7239be129752dc9111d1fb823942f82f431901f6b7417bb68

Observation 4abad966-84ef-437d-ae28-0dc8b896779e · outbound

This paper cites an unresolved cited work.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:38.006689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.013610Z digest=sha256:0beef0340c6c4d3c85e3a032d92a2e741f0043d80a0a348fcfd9f15ef63fdd21

Observation d52e6d40-6bfc-41c2-8d0a-50565de44a89 · outbound

This paper cites Measuring coding challenge competence of language models.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Measuring coding challenge competence of language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:37.816210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.091972Z digest=sha256:e486653de97c73891ae49dd504f8f07abb7d878eb5fcf2ee71a71638a4a43c4a

Observation bb8ec7f3-661b-44af-8291-489dbbd997ad · outbound

This paper cites Scaling Laws for Neural Language Models.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Scaling Laws for Neural Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:33.194108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:33.194108Z digest=sha256:a7b4c4bbf131d12697f35f464b35396018345679d1438d9f9568e01a3a323927

Observation ab935133-f3d8-476c-8c1f-487e04f704eb · outbound

This paper cites Procedural generation of physics problems for visual reasoning.IEEE Transactions on Visualization and Computer Graphics, 26(1):234–244, 2020.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Procedural generation of physics problems for visual reasoning.IEEE Transactions on Visualization and Computer Graphics, 26(1):234–244, 2020

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:37.604500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.286245Z digest=sha256:05aef325c1acd59be0f574472a803b194dccf3959cd06ba1aad33b862bf77f6c

Observation 46b7beb5-6ee1-4a6c-995c-e592ec989bef · outbound

This paper cites BLIP-2: Bootstrapped language–image pre-training with frozen image encoders and large language models.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models BLIP-2: Bootstrapped language–image pre-training with frozen image encoders and large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:37.412272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.376322Z digest=sha256:2fe810ececf5d9f5cd9bbd54004c697ee2380bab7cde50352732d0883f7dccd4

Observation 918a2682-236b-4246-8176-ca67e772d88f · outbound

This paper cites Lee, and Tamara L.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Lee, and Tamara L

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:37.174142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.430010Z digest=sha256:f922f764d1715dd71968907195faa1c513784da22b99e57adc8141099c5f0bfc

Observation f6721ae9-46e8-41a4-bf8e-3ff8b0908260 · outbound

This paper cites Few-shot prompting for physics word problems with large language models.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Few-shot prompting for physics word problems with large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:36.946313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.502444Z digest=sha256:8a201fd472c09ea19a44cbe79917583d40cd7cd10877925742b8f4e22a8c42da

Observation b941ff2d-c4d6-4716-9ebc-81f9d806f00d · outbound

This paper cites Scienceqa: A large-scale multimodal dataset for scientific question answering.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Scienceqa: A large-scale multimodal dataset for scientific question answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:36.728305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.591312Z digest=sha256:19329844773b53d4463b247fbb4e8fe8c18cf86f2e5b5ba6888425569784ea41

Observation 3cf08a2f-9bce-4a8c-991e-1cb335bca5a3 · outbound

This paper cites Mathvista: A benchmark for visualizing and reasoning about mathematical expressions.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Mathvista: A benchmark for visualizing and reasoning about mathematical expressions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:36.556529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.641615Z digest=sha256:4907b924231969f7525aaf2ff311bf05bce8c30f71717f48f700ac26341cf07b

Observation 22cc50fd-5182-4f77-a47c-6c72db6475ba · outbound

This paper cites Analyzing common failure modes in large language models.Transactions of the Association for Computational Lin- guistics, 11:612–633, 2023.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Analyzing common failure modes in large language models.Transactions of the Association for Computational Lin- guistics, 11:612–633, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:36.335221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.724164Z digest=sha256:704e064f1f172ff4fdd5ee39a74d43fdd774bb930e6ee9b3b4d3f2c49791a197

Observation 241a2288-0a65-4b01-bab3-7a82d8422cce · outbound

This paper cites Adaptive numerical integration for physics-based simu- lations.Journal of Computational Physics, 435:110260, 2021.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Adaptive numerical integration for physics-based simu- lations.Journal of Computational Physics, 435:110260, 2021

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:36.146695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.769419Z digest=sha256:e96f228ac5a6a8b2323140163d902b3b38dd555deae7addd59e6ee30ebf17e77

Observation a9b27762-7779-4c56-be2a-214b9e56eda9 · outbound

This paper cites Physion: Learning to predict physical interactions through video simulation.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Physion: Learning to predict physical interactions through video simulation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:36.039538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.825014Z digest=sha256:7d7f6bde3c435a4b9db0489ba1a21fe86bc389dc7858798ab8e8eaa86ec7baa8

Observation 19f93100-543a-4394-9548-cb7f91a50322 · outbound

This paper cites Physical interaction question answering (piqa): A test of physical commonsense modeling.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Physical interaction question answering (piqa): A test of physical commonsense modeling

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:35.896695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.917450Z digest=sha256:11e2206e800addb6e4c849b53ba61287f3bf6f5e0e12a8b9c56743fe12ed9f9d

Observation 94df43ec-7a9c-45f6-978b-cad41efa83e1 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large lan- guage models.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Chain-of-thought prompting elicits reasoning in large lan- guage models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:35.790847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:33.993416Z digest=sha256:59c1a959495f4bd10d4610ffbfcb353ea0f550afc6cfa88b24226767186ddf13

Observation 5ee4405e-c1fa-4fce-889a-2ce71fb29c0f · outbound

This paper cites Emergent Abilities of Large Language Models.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Emergent Abilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:34.051204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:34.051204Z digest=sha256:12c2fd5f3a3433fb3db18d0958964f6379bde37e37ed131afc6e10831835e060

Observation 6cc90d26-9645-4021-9b2c-4c277e271d6a · outbound

This paper cites A Comprehensive Survey on Multi-hop Machine Reading Comprehension Datasets and Metrics.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models A Comprehensive Survey on Multi-hop Machine Reading Comprehension Datasets and Metrics

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-04T20:58:34.844659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:34.173167Z digest=sha256:fb6202bbdd541559d62ffc93aec1da28fe7d48bca6004a747021aac50d484a07

Observation 494d9954-6943-4481-91f7-e18006b01fdb · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:34.283178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:34.283178Z digest=sha256:c3e3f05f5e2f744908d33b21418c6a1d50076552ac006210cb49d2a3c0770262

Observation 963377d4-efb1-4d99-b1e3-44c35332ead9 · outbound

This paper cites Llama-adapter: Efficient fine-tuning of language models with zero-init attention.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models Llama-adapter: Efficient fine-tuning of language models with zero-init attention

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:35.668974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:34.415057Z digest=sha256:4b33b339e0d0784e462a79d60f107c9feedd8f35a38aa24ae5bc93af2da694ef

Observation 46da5d6d-8ee8-4bf8-8fcc-00cb4b906218 · outbound

This paper cites The cater dataset: A diagnostic dataset for compositional actions and temporal reasoning.

Interpretable Physics Reasoning and Performance Taxonomy in Vision-Language Models The cater dataset: A diagnostic dataset for compositional actions and temporal reasoning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:35.418730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-04T20:58:34.564196Z digest=sha256:c2dc65c45300f7761518d4a545c142552dd1616eaff0b0ae8135d526aec12138

Pith citing papers

No inbound Pith citation observations are available.