Pith. sign in

Paper Citation Record · LEDGER

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

As of 14 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.07584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07584 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:22.245545Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4476f292-0ae6-45be-aa36-0a19615192ed · outbound

This paper cites 2022 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2022 , publisher =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:23.035059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.069655Z digest=sha256:8f1ee5e3c1d6e0831ff01a4c0616d61e9277e5d521b8b139083c65b139e8a4cd

Observation 64d35d65-af2f-4476-9bc7-8600396687eb · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.078210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.078210Z digest=sha256:34f0d61021dcf5a34c93a548dff2a74de4e2215a8a18904fada121061a092b83

Observation cd3324ef-fff3-4469-8f7b-967e10675454 · outbound

This paper cites Lawrence and Girshick, Ross , booktitle =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Lawrence and Girshick, Ross , booktitle =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.990804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.088566Z digest=sha256:e8359f3a07eff7741e5547f31226564fb2d13abed64a84d1999ca1ebd62ee2c1

Observation 51e95402-573c-4d2e-8e2a-c49209618063 · outbound

This paper cites 2024 , doi =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2024 , doi =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.969167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.096685Z digest=sha256:09672b4daa44db4d07d44cd6eda51638f8f6b90d9504df8293d4624b8f3c4b82

Observation c5f6c67b-10d0-45bf-a23a-ddc118cb8283 · outbound

This paper cites 2025 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , publisher =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.946307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.103389Z digest=sha256:d8be6493b48d7e4c46933bd180ed140ceac893f6f3e15e5505054edbf6ac10c3

Observation b4265b22-c5f0-4ae5-bf4f-0d0147d6d303 · outbound

This paper cites Journal of Data-centric Machine Learning Research , volume =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Journal of Data-centric Machine Learning Research , volume =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.928124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.108889Z digest=sha256:ca7be7972dd124758a7702d87c5b66324b614c106b314fcf06fc99e18c02b6ae

Observation 387f56a9-c392-48cb-9e1a-109170f7b229 · outbound

This paper cites 2019 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2019 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.114861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.114861Z digest=sha256:228001ceb7ec9fae33014658d6743826e141fc87369a82753dc7ab5c748f211c

Observation 16cf6ca3-e2f6-4e97-8a37-0cf641434036 · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.121965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.121965Z digest=sha256:e5ccf0a8802d07be98d03104288c1dca4b4253fad1df11b529a7cc7a5f99f160

Observation b6432445-2633-4b02-b7c0-d69144babb8c · outbound

This paper cites MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.128677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.128677Z digest=sha256:e37a581f13325452835c176452eac0abcb90c9122b3317c8ee232b0199145b00

Observation 59df4741-b015-4241-ad15-04b1e257e67f · outbound

This paper cites 2024 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2024 , publisher =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.889531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.134634Z digest=sha256:b615d2fe745cd58188ffc9a1160ddd2360e4b2d28785a070eda1bea8fd6e2a97

Observation 2ead2ad2-c935-42e2-a944-dc8ad8a058b9 · outbound

This paper cites 2025 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , eprint =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.861618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.142290Z digest=sha256:9ac4612555960cbc24de9b74de1282c3a3c6573c3293c36a78e9631b5a440115

Observation 6964239f-fadf-4c63-bfa3-2830d73f31de · outbound

This paper cites PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.148366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.148366Z digest=sha256:83525a999fca59d5fb22474b89d80fb7f3a3d3fc2149107c67af8a691e2e6b9a

Observation 6fdbe709-ec3d-4eae-a840-88dabe50aeec · outbound

This paper cites 2026 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2026 , eprint =

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.842026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.156416Z digest=sha256:11d37a8654caf76d1730775923c50d5b6a09b088e5b471a00d3cb57c37cec07c

Observation 7c5160d0-5f51-4c7f-994d-825831750f7c · outbound

This paper cites NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.162543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.162543Z digest=sha256:63edd122a78a56e0bfb4c96a69ddd32542336bfbce14fb16f4f2476f9f462ffb

Observation 0ae47376-8cfc-49ae-b668-44c7468c4ec9 · outbound

This paper cites 2025 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , publisher =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.820733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.170552Z digest=sha256:d1b4869a03566f432881d126ba586ec11b5162ac971c0712ed73f555c15a5ac9

Observation 801da51e-d348-486c-acc6-532317dd6cee · outbound

This paper cites MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:36:22.517208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.176532Z digest=sha256:580b7c4548cf1bdfdd52db03d7c2c86e20b194bdc29d3cd98db43de8add8ddd4

Observation eae8b963-b47e-41b7-a638-46c48763c3fb · outbound

This paper cites 2602.08367 , archivePrefix =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2602.08367 , archivePrefix =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.184820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.184820Z digest=sha256:b45c8dcbf95bbf28f819b80fd512aadf6695c1995b6e098132d311ba27c4199e

Observation bb30cc29-e875-43cd-ab43-ff33ec22c586 · outbound

This paper cites 2026 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2026 , eprint =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.801065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.191335Z digest=sha256:0a065793fbe287f27e3d45b1650c071f53788141c93481a77aa0c27e96b12420

Observation 9f7642af-c173-4bb1-86d9-bf0125a3b057 · outbound

This paper cites ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:36:22.351158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.197364Z digest=sha256:50ef5cda6c0d221c509a12e374782a62acdb3c4673e08f51d97b5714f7497e63

Observation 8cb0bcc6-4c1a-4a03-878e-d14bcfffb816 · outbound

This paper cites 2025 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , eprint =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.781866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.204494Z digest=sha256:9e19ef349afb6db50195c254e44f862e7b1ad0775ec6406670e9a752f382fa80

Observation 34c41c24-873b-457e-9a62-02262ed09492 · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.211047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.211047Z digest=sha256:25bb45c74d50490be362ad628c5cd5ef87a77efa836854191bdf74ccd0b7f782

Observation d87097a0-a98e-4758-9d20-9779e2db31aa · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.216429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.216429Z digest=sha256:70a0f8189c5c8f48b937cf825a99b889f7a752979a23762985c07f462261541f

Observation 2577fa03-b8e9-4dcb-83bf-5749fee3e8a5 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.742909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.222214Z digest=sha256:32d9ca4c64ca49840cf2bedb03217f0763bc5a01b2fe67298841d01011818469

Observation f19be4b3-7942-493d-84b7-0c85d85eb92d · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Advances in Neural Information Processing Systems , year =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.226802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.226802Z digest=sha256:60a5706ea9ef62d52004a474bdef2c1a94e934e6a38163d7b496222f9340e596

Observation 9c4f1d09-8883-45f5-8ae2-e9bb2c002866 · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.232109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.232109Z digest=sha256:b5ee0f7ecc419c1b609cb951b887e956f3cbc133b856f5fe8c0dd78a6970de23

Observation 84ad65ce-7676-410b-8f3d-1f48132d4a42 · outbound

This paper cites Complexity of Computer Computations , editor =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Complexity of Computer Computations , editor =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.686421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.237385Z digest=sha256:2c397078279f44b15d6f0f726ff202aa59196be85c3385669778a908a58f63eb

Observation d99f49fb-09d8-4954-bc74-e5ae756b4aba · outbound

This paper cites 1979 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 1979 , publisher =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.664591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.245545Z digest=sha256:630413a10421ffc64fa8cbefe3e53cd1a281f7a014da0e5a85c0f3d96582d2c7

Pith citing papers

No inbound Pith citation observations are available.