Pith. sign in

Paper Citation Record · LEDGER

Benchmark Everything Everywhere All at Once

As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2606.06462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06462 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:03:52.964870Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:44:30.868733Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact31
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 245b9b83-5266-42bf-ad68-400d35a68d55 · outbound

This paper cites Synthetic dialogue dataset generation using llm agents.

Benchmark Everything Everywhere All at Once Synthetic dialogue dataset generation using llm agents

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:46ff0e75241bf4d62e762bc182e4d7a4fdbd2253e15f91c836ea15d57f3ddb4b

Observation b35417bc-f33f-49ef-a6d3-0349718ddb53 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

Benchmark Everything Everywhere All at Once LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.041205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:a456aea85d9ed1795ba62fb17029d217138dd81dba1370005d82af2148770eb0

Observation 0d72cafd-15e9-42b9-9536-f401aba277c1 · outbound

This paper cites System card: Claude opus 4 and claude sonnet 4.

Benchmark Everything Everywhere All at Once System card: Claude opus 4 and claude sonnet 4

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:f78782c9e8ba7270fa45dd196276aa83bbe750eb9994f41028b9179cf9e4bdf1

Observation ca964959-c795-4c7a-8c9a-378261b5d8b9 · outbound

This paper cites Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks.

Benchmark Everything Everywhere All at Once Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.100580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:cb3cbd06e2aef43ea67f269822156d94709dc5ccd55de82a9d9ee0669808b695

Observation 619d7574-9f7c-4165-8416-b1564b4e6669 · outbound

This paper cites Qwen Technical Report.

Benchmark Everything Everywhere All at Once Qwen Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.105788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:1efcbb0445b9944ad135195f38a6d95342a4e797333f923c9753d2a93d0a8a99

Observation 9968c45b-c9ef-4ee3-9a70-3dcf01497f2f · outbound

This paper cites Qwen2.5-VL Technical Report.

Benchmark Everything Everywhere All at Once Qwen2.5-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.046343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:a217c63ff928b6152fe5ff4744ccb0a85274810e27b90d7fb924e737c1fb9331

Observation fd87ddf5-a1fe-4f20-9b9b-0fc505c747b9 · outbound

This paper cites Benchagents: Automated benchmark creation with agent interaction.

Benchmark Everything Everywhere All at Once Benchagents: Automated benchmark creation with agent interaction

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.037157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:0e696c1fb96545c0884db24a034260202ca232a11a653c74b58ef640088dbac7

Observation 4ad58176-15f9-4c22-806d-0f3e7d0d395f · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Benchmark Everything Everywhere All at Once ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.063664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:bc1791416be83c741a4a89b75b8d3d3c14ddaea2ba13bad81bd1bc8cad2e89c4

Observation f1489bf8-79d0-401b-b41a-8b7177db5e19 · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as-a- judge with vision-language benchmark.

Benchmark Everything Everywhere All at Once Mllm-as-a-judge: Assessing multimodal llm-as-a- judge with vision-language benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c0d5650f2f48f3acc8e59a1e74ff4863efd6b1211c7700705383e06699c3aff2

Observation 768d7b4f-9bec-4cf9-b8b0-a431fe83a748 · outbound

This paper cites Are we on the right way for evaluating large vision-language models?NeurIPS, 2024.

Benchmark Everything Everywhere All at Once Are we on the right way for evaluating large vision-language models?NeurIPS, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:ca323c366bf4201052463994923a79b0fab839d4488066d809d319afe6825710

Observation 5db47bd7-7218-4732-b59d-ef9a38d09c68 · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

Benchmark Everything Everywhere All at Once Can Large Language Models Be an Alternative to Human Evaluations?

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.078136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:a230980d5eec4aedce2c97de38a88c2d178ccafec37e81e4440adcfb8e332b66

Observation 388e95ee-e877-481f-8495-7b254550c651 · outbound

This paper cites Cl-bench: A benchmark for context learning.

Benchmark Everything Everywhere All at Once Cl-bench: A benchmark for context learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.019997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:3f66a8bf9691ebcbf748626e0fc100bb949315d133330c9468d0de23eb8de45b

Observation 1fd43dfd-dcc5-4ffa-8658-3d91369f039b · outbound

This paper cites On path to multimodal generalist: General-level and general-bench.

Benchmark Everything Everywhere All at Once On path to multimodal generalist: General-level and general-bench

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:07e5d53e1a0d3c4e34d3c92180411b7b8116ea2698e48bebe866d3c581e3dd90

Observation 6e701b4a-7098-42e5-ba26-00e5195a9927 · outbound

This paper cites Gptscore: Evaluate as you desire.

Benchmark Everything Everywhere All at Once Gptscore: Evaluate as you desire

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:e00d14e52ed94595902a81fe5ba50f554cf6105f5a7463e484910dc8d3f69bad

Observation 135ee187-2bb9-4a4b-8dd6-6f2edba2bb89 · outbound

This paper cites Gemini 3 pro model card.

Benchmark Everything Everywhere All at Once Gemini 3 pro model card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:48aea1046b2394c739fca0d5d4cf81877b82b995a20a751751352dedc46ad226

Observation c7563fcb-991c-420c-82b9-1ccdefeede39 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Benchmark Everything Everywhere All at Once Measuring Massive Multitask Language Understanding

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.043591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:b9248d15f1ab803e7a368b05cd6d3474b51e76e25d0d9fca0df5150eb5a92dcd

Observation 633ea2af-0f31-417c-a0b4-4da42887e5f3 · outbound

This paper cites Cerebellar output shapes cortical preparatory activity during motor adaptation.Nature Communications, 2025.

Benchmark Everything Everywhere All at Once Cerebellar output shapes cortical preparatory activity during motor adaptation.Nature Communications, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:73b471c13cd877c7c87d4adf02cd109f02544668a66a87d8f59a49d422806ec6

Observation dbde44d1-873b-4a2a-a320-a89eb2c38689 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Benchmark Everything Everywhere All at Once Kimi K2: Open Agentic Intelligence

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.064881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c158fb94cd12df773a5a860d6634824793a953eb0c423ccd46bd35ba61adf9d5

Observation 37e4bbe7-3d81-4191-ae13-f8c311d235a8 · outbound

This paper cites Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai.

Benchmark Everything Everywhere All at Once Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.094802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c566a303b23d77883c6867c301e63615fd618f50a38b0e28b00e61aa63fa6cf2

Observation 669c9553-7c22-4295-b5dd-648e1f15482f · outbound

This paper cites Act as human: Multimodal large language model data annotation with critical thinking.arXiv preprint arXiv:2511.09833, 2025.

Benchmark Everything Everywhere All at Once Act as human: Multimodal large language model data annotation with critical thinking.arXiv preprint arXiv:2511.09833, 2025

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.097773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:58a3c2a2a1ec3f5f5540f73327d6c8bd616b350291cf2ef5e897878f666466d5

Observation e30cf7db-2964-49d8-8f2b-d15bb965016c · outbound

This paper cites Agentbench: Evaluating llms as agents.

Benchmark Everything Everywhere All at Once Agentbench: Evaluating llms as agents

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:2a7468a990d47fc2fab6c8d59d673e9b26f346ccde1a84a2b07e9d7e9f233fd9

Observation c665b82b-7963-4cfa-a425-243e42b5182d · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV, 2024.

Benchmark Everything Everywhere All at Once Mmbench: Is your multi-modal model an all-around player? InECCV, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:9aeb42b3d33276b6fbb558444b26d642dfd5cefd05ad89b080a165538b51859f

Observation 0b02423a-e7ea-4845-aee9-98e2f7cd37c4 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Benchmark Everything Everywhere All at Once MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.084664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:05bd47022021a5278e19a87ec07611f3ade476716e43991f0df9d616b923c277

Observation 0d5f8588-81d8-460b-a234-dde1141d3f2f · outbound

This paper cites Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena.

Benchmark Everything Everywhere All at Once Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.086653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c957a70c8254ffec235521624d8425434416a2725ab6a61b4d35a2654c29a90a

Observation dbc013b3-4552-496a-a0ed-306f1f7f7d98 · outbound

This paper cites UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance.

Benchmark Everything Everywhere All at Once UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.040295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:77cad56f32b8d64eec855a0d6e67769cb9d995e779a8b80c1c2942ebc1ac68ff

Observation ddbe4b47-6817-45c7-bee9-db926870d4ef · outbound

This paper cites MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling.

Benchmark Everything Everywhere All at Once MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.075462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:162c0dc64ef9f6561e8aea6c2bd5b784b40f8e6ceb448c3122ce995c78ec147c

Observation 5d280fe4-9844-4fe7-a8f2-a4d22c845091 · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

Benchmark Everything Everywhere All at Once Autonomous Evaluation and Refinement of Digital Agents

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.083651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:2f84051b033ab00b932dac9348938819600261749d0e3115d437f99e90619b17

Observation c5408104-8e3a-4565-bad1-90e0e568c3b3 · outbound

This paper cites Benchmarkˆ 2: Systematic evaluation of llm benchmarks.arXiv preprint arXiv:2601.03986, 2026.

Benchmark Everything Everywhere All at Once Benchmarkˆ 2: Systematic evaluation of llm benchmarks.arXiv preprint arXiv:2601.03986, 2026

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.066628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:0e6f54e7fe5d98272bb97d4f47eff10dc4328038aa791f6c6ecfb7bea184075c

Observation c1fc5154-382b-42d9-8cf0-c8e5b7ad21f9 · outbound

This paper cites Autobench: Automatic testbench generation and evaluation using llms for hdl design.

Benchmark Everything Everywhere All at Once Autobench: Automatic testbench generation and evaluation using llms for hdl design

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:fe30c5082a8a9aa90c061c46ee42646356972d6b62d312a5a16c12df9bebe367

Observation 9a8ce739-6907-4129-aafb-40fbf6c90da8 · outbound

This paper cites Qwen2 Technical Report.

Benchmark Everything Everywhere All at Once Qwen2 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.087084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:1964eb09dd80f03f5eb546c33a3da9e3ce81d06614ec78f21cf359e9a34d1f02

Observation 2ecb2511-f8e8-4118-95b7-535678e94aa7 · outbound

This paper cites Qwen3.6-Plus: Towards real world agents, April 2026.

Benchmark Everything Everywhere All at Once Qwen3.6-Plus: Towards real world agents, April 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:ae8b3bbbc3b6431a3e1c3e44cc5e40e52fb94d70f279b1b66bbdfaecc4ceb5fd

Observation c61b8177-133d-4102-bd6b-e94a801eb13a · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Benchmark Everything Everywhere All at Once GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.058019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:fe95c8185bb5d5425ec559e0e3efa9561471d09c98f2d6174f25e2f79337e395

Observation 457f59f2-cb04-48de-b19a-59c3bc8710f2 · outbound

This paper cites Neuronal dynamics of cerebellum and medial prefrontal cortex in adaptive motor timing.Nature Commu- nications, 2025.

Benchmark Everything Everywhere All at Once Neuronal dynamics of cerebellum and medial prefrontal cortex in adaptive motor timing.Nature Commu- nications, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:28cc38378cc6c02b59ff0c61a4aa09fce532d467915653e66dfbef30bda4908d

Observation e9217cc7-ee50-47bc-82a2-f9083a7fe733 · outbound

This paper cites TAGAL: Tabular Data Generation using Agentic LLM Methods.

Benchmark Everything Everywhere All at Once TAGAL: Tabular Data Generation using Agentic LLM Methods

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.082147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c389ce93af6a8773d512347f5bc2199fb33b4b45ae721063691bb3f85fe1ed9f

Observation e9961485-1c81-4be7-86b5-6715f7991c79 · outbound

This paper cites One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026.

Benchmark Everything Everywhere All at Once One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.072630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:bc93137d92a8dbe3e4e085cf7cd81e3a70578e808fa23376ba07b1f50c0aa3aa

Observation af2d7b06-9ba6-4c0c-9441-0e8fef07efd9 · outbound

This paper cites OpenAI GPT-5 System Card.

Benchmark Everything Everywhere All at Once OpenAI GPT-5 System Card

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.019426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:ae34a296fbe6e346c8ec9c8f6c265c69a2a6556b7fb73014d7cfccd18052ea83

Observation 658fff00-4afe-4123-8d5b-e3a75383f82a · outbound

This paper cites Towards vqa models that can read.

Benchmark Everything Everywhere All at Once Towards vqa models that can read

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:fdfea722813ed4f68876cccc6d9c0d0a656e83847c1c9a2222b15168946e27cc

Observation 731e3fc0-afdf-4757-9bd1-4ec1b37d7e34 · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

Benchmark Everything Everywhere All at Once Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.077010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:e6a3159dd38dd975000bfe2204d65523b6f9c93c398c4882ff09993a02cd4f52

Observation 9131cd49-9410-49ce-a764-9bd7f6103cf2 · outbound

This paper cites Spacevista: All-scale visual spatial reasoning from mm to km.ICML, 2025.

Benchmark Everything Everywhere All at Once Spacevista: All-scale visual spatial reasoning from mm to km.ICML, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:143ce2ad78c9e257c80a05c5cc22ae35f83bf203145a305b7a5b245a86be7eee

Observation a90dccf7-b1b2-42ea-95ad-6cb33a07ac0e · outbound

This paper cites RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration.

Benchmark Everything Everywhere All at Once RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.100601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:84a09473bbd58d441bf85f7ba5c0e46a82e99ba665aa1de2821c73d60518a8eb

Observation 3f29bb34-5214-4f29-9e2e-4fe9b0dd9799 · outbound

This paper cites AI-Researcher: Autonomous Scientific Innovation.

Benchmark Everything Everywhere All at Once AI-Researcher: Autonomous Scientific Innovation

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:46:59.103155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:109242a9fb96f516ae72d27287aa4be8c5516492b3f82bc17182b31759dd522c

Observation 40a447e6-993b-4698-9229-e4f57f43f647 · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February.

Benchmark Everything Everywhere All at Once Qwen3.5: Accelerating productivity with native multimodal agents, February

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:416d1eb54551965138cfe929469e2081cdb1c024a86746495cff5127764bfaa9

Observation a39ce569-e168-4b2a-a9d3-a2f1f5617958 · outbound

This paper cites an unresolved cited work.

Benchmark Everything Everywhere All at Once Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:5af3d5709cb56ababc80e01303de06202684a49e8ec42bee0fd48339da77077a

Observation ec5913dc-a94b-44d0-b720-99c0414fd1d9 · outbound

This paper cites Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration.NeurIPS, 2024.

Benchmark Everything Everywhere All at Once Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration.NeurIPS, 2024

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:7f0f7d2d608430b1044676b04e1cc0a24ae8c5a4c0727bbaf5b7929d640796a2

Observation 6fdf7826-c3b7-4a84-9710-e209c26340e6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Benchmark Everything Everywhere All at Once Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.092119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:4cce4d69bcf8a36d8025424b181ac261fd0c42dd4d4f07da2246a67a04f3efa3

Observation 6443c399-756a-496d-9d80-1cc4671ee353 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Benchmark Everything Everywhere All at Once InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.079441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:5db6e167719b0b728878395390be6d4590dad8e897a3a7808a0a716f2e754aa2

Observation 3d55a4d1-7b9f-40b4-9326-73aecf0dd419 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.NeurIPS, 2024.

Benchmark Everything Everywhere All at Once Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.NeurIPS, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:9314f18571aa1dc58b46710275f6e49086072123225577d43d144358316581d4

Observation 36e5be77-918a-42c5-8775-12e586c47bf1 · outbound

This paper cites Finevision: Open data is all you need,.

Benchmark Everything Everywhere All at Once Finevision: Open data is all you need,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:365af988b0faa1229d46e136bdc763330373fbac9b56fdba445f217245c7bb63

Observation 0c9e453a-0d14-49e1-87ac-baa3d9589898 · outbound

This paper cites FineVision: Open Data Is All You Need.

Benchmark Everything Everywhere All at Once FineVision: Open Data Is All You Need

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.105452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:3c846bcb77a7cc2d2e28bbb8528c29d024c9e9982672d899b8d9978df3e5eae5

Observation 0254e79d-595d-4b03-ad8b-2fe13bcbb4ad · outbound

This paper cites Language prompt for autonomous driving.

Benchmark Everything Everywhere All at Once Language prompt for autonomous driving

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:3aa1cc746709df2aac9fee36e9ee654925dc3c212f89a9105c3c0ca7be619fd6

Observation a568e52c-c097-470b-b9d6-29a1d16ce8af · outbound

This paper cites Qwen3 Technical Report.

Benchmark Everything Everywhere All at Once Qwen3 Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.092402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:cd1cd128dc166b57338024591d5eaa320af81efe5829d90e8fddcf0dfa9b49e6

Observation d6ef1379-6292-4bd3-b045-8d04ff4a8af5 · outbound

This paper cites From Web to Pixels: Bringing Agentic Search into Visual Perception.

Benchmark Everything Everywhere All at Once From Web to Pixels: Bringing Agentic Search into Visual Perception

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-02T13:46:59.103114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c5edcf27aab7d4b2aab1cb341dfc51c208b1307bf798c9b76143c121c5b80e9b

Observation ce0cb930-e299-4436-970d-e04dd1ba1e54 · outbound

This paper cites Swe-agent: Agent-computer interfaces enable automated software engineering.NeurIPS, 2024.

Benchmark Everything Everywhere All at Once Swe-agent: Agent-computer interfaces enable automated software engineering.NeurIPS, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:c7b4445a3c1b1f90a7fbd480c5302f42ab6e759fb99002b2fc19ca2c1a58909a

Observation 0fc4f0f7-16ad-462b-917e-e9fbe6a1b066 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Benchmark Everything Everywhere All at Once React: Synergizing reasoning and acting in language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:194384be55c74d25f9e5e85cc5d0deea63def7bdb35653aefdc6c5e19927e152

Observation 9edae0f2-f3d7-4a2a-96d1-3fffd218b7ef · outbound

This paper cites Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark.

Benchmark Everything Everywhere All at Once Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:f237404c7a403589a54747743f0703f605dcd7ea371163adc93933f2a046d4cc

Observation 8966b782-672b-42b9-8d31-46d2ecd65cd6 · outbound

This paper cites Evaluation agent: Efficient and promptable evaluation framework for visual generative models.

Benchmark Everything Everywhere All at Once Evaluation agent: Efficient and promptable evaluation framework for visual generative models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:9ee8acfacdae0b8946ae63fe920bd6e8308efe40483838197aeb21be7f602d4b

Observation a5d74065-ab02-45f3-8954-d77a1791aeef · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InECCV, 2024.

Benchmark Everything Everywhere All at Once Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InECCV, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:6fa250af18054bd92a613deec327d21e5f1538b5112f90519cdf7237c180a86b

Observation 649636fe-f6c3-4c70-84f5-13dd6d08f569 · outbound

This paper cites Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?ICLR, 2025.

Benchmark Everything Everywhere All at Once Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?ICLR, 2025

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:617133763b21575613fff908d9c0fb86b7e46a66667b6b2338252f4ee8077f05

Observation e4110844-b819-4354-a6e6-7c3ffe1e643c · outbound

This paper cites Dual and plasticity-dependent regulation of cerebello-zona incerta circuits on anxiety-like behaviors.Nature communications, 2025.

Benchmark Everything Everywhere All at Once Dual and plasticity-dependent regulation of cerebello-zona incerta circuits on anxiety-like behaviors.Nature communications, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:0385b1132a520d25f8fd2abe5f60d86f0ba6d1b44942299403fafbfeaa7fedae

Observation 6bb49017-3926-4ed8-9eb3-4492b1769d0e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023.

Benchmark Everything Everywhere All at Once Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:6ebeeebf3470f6977c46e709d7d423d456505222ea751603d3cc449d9fdc1c90

Observation dcce68a2-5c0b-41d5-ad47-c87a8d197f85 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Benchmark Everything Everywhere All at Once Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:df99f3a48beb7cbf1660126c10c4ccaf7a80eed9713cc1942210dbd7dcee9480

Observation f81f8836-6be6-4578-9110-59818dba5789 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

Benchmark Everything Everywhere All at Once JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.089343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:728d761e917131d70a930ee17b201447657a7be9f6f4bc67ad538ebae28f1f3e

Observation 28547017-5184-4f09-8b7d-5e50d4347339 · outbound

This paper cites Q., and Shou, M.

Benchmark Everything Everywhere All at Once Q., and Shou, M

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T13:46:59.080772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:34cd7185e62cb8dc1644275732bcc3ab3f1f94271fe5e7a1ead78f70143415c5

Observation 5489fc31-c597-4ff0-9975-edd479e72c4c · outbound

This paper cites at the same time.

Benchmark Everything Everywhere All at Once at the same time

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-28T01:03:52.964870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:03:52.964870Z digest=sha256:9a2802b413929d6d713d0779f0b8aec7150881e537ae4030f175260c0b8c2b36

Pith citing papers

Observation e0f61985-60e2-42d7-a880-9f575693589f · inbound

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs cites this paper.

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs Benchmark Everything Everywhere All at Once

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T03:19:12.528300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:19:12.528300Z digest=sha256:708081c3a0f7ae72b3ce2d7de8761a87ee2418114b1f333d04bbdcee4eae02c3

Observation 6b4e567e-dc86-4d7c-b7da-0460d0f2c561 · inbound

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs cites this paper.

NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs Benchmark Everything Everywhere All at Once

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T01:44:30.868733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:44:30.868733Z digest=sha256:69aabcd87a7a688b38507fa25d57e325ad7ad74c86b1da41fcf43cf3ad60ec88