Pith. sign in

Paper Citation Record · LEDGER

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

As of 13 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2608.07584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.07584 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:22.245545Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4476f292-0ae6-45be-aa36-0a19615192ed · outbound

This paper cites 2022 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2022 , publisher =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:23.035059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.069655Z digest=sha256:d0f05a73f4de21a46f3b6502cd7c02b34180e60fd2a3bf92f630d5a6a3a51b8a

Observation 64d35d65-af2f-4476-9bc7-8600396687eb · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.078210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.078210Z digest=sha256:4fb2398783358fec1bd41e80f32a21f5f44a8e16016191c0af76fac15d4e50d2

Observation cd3324ef-fff3-4469-8f7b-967e10675454 · outbound

This paper cites Lawrence and Girshick, Ross , booktitle =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Lawrence and Girshick, Ross , booktitle =

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.990804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.088566Z digest=sha256:47505a25805bc5ba1401d819f764d0567b16cb3231726196b3039f3b9dc41903

Observation 51e95402-573c-4d2e-8e2a-c49209618063 · outbound

This paper cites 2024 , doi =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2024 , doi =

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.969167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.096685Z digest=sha256:fd328972ee1eae4c800590e1abbdaab46a8affc3edf4b57235e14621ac0d0b89

Observation c5f6c67b-10d0-45bf-a23a-ddc118cb8283 · outbound

This paper cites 2025 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , publisher =

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.946307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.103389Z digest=sha256:2679505c419d82e1048aadab783bb50fd3e50c3edb1462d4c23804b7c1c2be85

Observation b4265b22-c5f0-4ae5-bf4f-0d0147d6d303 · outbound

This paper cites Journal of Data-centric Machine Learning Research , volume =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Journal of Data-centric Machine Learning Research , volume =

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.928124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.108889Z digest=sha256:9355c344aff238c3c89ac551231770c7b333f96f1ee013ff8422f8043b884062

Observation 387f56a9-c392-48cb-9e1a-109170f7b229 · outbound

This paper cites 2019 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2019 , eprint =

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.114861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.114861Z digest=sha256:ababb4714fc36c9313c779e7e3b51028feeed5ae90656d5e752784e281d71e11

Observation 16cf6ca3-e2f6-4e97-8a37-0cf641434036 · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making BabyVision: Visual Reasoning Beyond Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.121965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.121965Z digest=sha256:16cbd1de2da5c63eb1119c60c4abc0ffb99d589c2fa3fd4043e17c743f9d5648

Observation b6432445-2633-4b02-b7c0-d69144babb8c · outbound

This paper cites MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.128677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.128677Z digest=sha256:c039fc29114fea8bf3ea9687d122beb3a56eae60fbb0ddb7bda7ee757c331f66

Observation 59df4741-b015-4241-ad15-04b1e257e67f · outbound

This paper cites 2024 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2024 , publisher =

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.889531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.134634Z digest=sha256:89bff9a40e898301a3b41a048448f8d2064528fa1fc45457f250f1025743f20c

Observation 2ead2ad2-c935-42e2-a944-dc8ad8a058b9 · outbound

This paper cites 2025 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , eprint =

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.861618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.142290Z digest=sha256:eee89e9b683ffa872615aca82b5f6a5429912588c9d2cdcb2907549996ef6d13

Observation 6964239f-fadf-4c63-bfa3-2830d73f31de · outbound

This paper cites PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.148366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.148366Z digest=sha256:9f88461c6a9608761a6541c04f3e00ececebd1763f44da15ec202c17e96cc024

Observation 6fdbe709-ec3d-4eae-a840-88dabe50aeec · outbound

This paper cites 2026 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2026 , eprint =

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.842026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.156416Z digest=sha256:ed5c8578faa5dfbb0dec10da443448d4a3f4b774ed70a27e60d49957212cb85b

Observation 7c5160d0-5f51-4c7f-994d-825831750f7c · outbound

This paper cites NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.162543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.162543Z digest=sha256:edbb308972d6ae0020f747ce3c5ea27d7bea6e609cd80891200df9cfadf220e0

Observation 0ae47376-8cfc-49ae-b668-44c7468c4ec9 · outbound

This paper cites 2025 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , publisher =

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.820733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.170552Z digest=sha256:120652ef94342ecfdc40cf29fe7a9f6ef9ce7c01e008845b363bb9b81509da31

Observation 801da51e-d348-486c-acc6-532317dd6cee · outbound

This paper cites MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MPCC: A Novel Benchmark for Multimodal Planning with Complex Constraints in Multimodal Large Language Models

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:36:22.517208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.176532Z digest=sha256:4650793cfbbc6f05725b080f2f4a5af1be8e9e5212988db89dc726978e77e933

Observation eae8b963-b47e-41b7-a638-46c48763c3fb · outbound

This paper cites 2602.08367 , archivePrefix =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2602.08367 , archivePrefix =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.184820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.184820Z digest=sha256:68e63dac600c52d6b2506ea63c7bd8ad9c8a19c2ea36ff195da8ef3c7662f630

Observation bb30cc29-e875-43cd-ab43-ff33ec22c586 · outbound

This paper cites 2026 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2026 , eprint =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.801065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.191335Z digest=sha256:66bd43ad1f7893736d5a766cabc89401ae8e7c27c56b7704d3eabee5667abd98

Observation 9f7642af-c173-4bb1-86d9-bf0125a3b057 · outbound

This paper cites ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making ORAgentBench: Can LLM Agents Solve Challenging Operations Research Tasks End to End?

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T00:36:22.351158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.197364Z digest=sha256:867a1c393485eb02098276a36077ad018bfc8641d0e0e2a43df85a6a3e5ee4e5

Observation 8cb0bcc6-4c1a-4a03-878e-d14bcfffb816 · outbound

This paper cites 2025 , eprint =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 2025 , eprint =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.781866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.204494Z digest=sha256:6c9bbae3f0fc14250ad961a21d56053cbd2f7e239fd35d149c4ee3ac65abf09e

Observation 34c41c24-873b-457e-9a62-02262ed09492 · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.211047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.211047Z digest=sha256:57c3e51c3105c6b1a4eda011a1b66a9ed4eda14465e6db6d94d540e1fa52bb0b

Observation d87097a0-a98e-4758-9d20-9779e2db31aa · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.216429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.216429Z digest=sha256:084ddcdf63b34b5340d1790ffedee375e9feec4740a17bc03be4e54fa4b049d7

Observation 2577fa03-b8e9-4dcb-83bf-5749fee3e8a5 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Proceedings of the IEEE/CVF International Conference on Computer Vision , pages =

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.742909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.222214Z digest=sha256:480abb2819f78fdd8c82e7384bd919d808ea04ee7e53d8d98f8d01c89f4be281

Observation f19be4b3-7942-493d-84b7-0c85d85eb92d · outbound

This paper cites Advances in Neural Information Processing Systems , year =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Advances in Neural Information Processing Systems , year =

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.226802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.226802Z digest=sha256:1daf752f758bd2b8b3ad45a82d10db644d3815e728cbf4297031b153e7cad566

Observation 9c4f1d09-8883-45f5-8ae2-e9bb2c002866 · outbound

This paper cites an unresolved cited work.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:22.232109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:22.232109Z digest=sha256:f86e455d8113cc38165149be0078b2acc103db38e73ee98d82d836af898ff81b

Observation 84ad65ce-7676-410b-8f3d-1f48132d4a42 · outbound

This paper cites Complexity of Computer Computations , editor =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making Complexity of Computer Computations , editor =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.686421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.237385Z digest=sha256:8ea1330e6b7dd381ce2aeead229acfdebfc5305cf9d800586be63721016b31a5

Observation d99f49fb-09d8-4954-bc74-e5ae756b4aba · outbound

This paper cites 1979 , publisher =.

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making 1979 , publisher =

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T00:36:22.664591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T00:36:22.245545Z digest=sha256:a68593d34de5d387b718221ec9fe13f3d3584440d0fedbcbdd457e3a6f0929b8

Pith citing papers

No inbound Pith citation observations are available.