Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:03:52.964870Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2606.06462.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T01:03:52.964870Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T01:44:30.868733Z
A source-named dated measurement, never combined with another source.
Source: cited_works
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 245b9b83-5266-42bf-ad68-400d35a68d55 · outbound
Benchmark Everything Everywhere All at Once Synthetic dialogue dataset generation using llm agents
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b35417bc-f33f-49ef-a6d3-0349718ddb53 · outbound
Benchmark Everything Everywhere All at Once LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0d72cafd-15e9-42b9-9536-f401aba277c1 · outbound
Benchmark Everything Everywhere All at Once System card: Claude opus 4 and claude sonnet 4
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca964959-c795-4c7a-8c9a-378261b5d8b9 · outbound
Benchmark Everything Everywhere All at Once Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 619d7574-9f7c-4165-8416-b1564b4e6669 · outbound
Benchmark Everything Everywhere All at Once Qwen Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9968c45b-c9ef-4ee3-9a70-3dcf01497f2f · outbound
Benchmark Everything Everywhere All at Once Qwen2.5-VL Technical Report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd87ddf5-a1fe-4f20-9b9b-0fc505c747b9 · outbound
Benchmark Everything Everywhere All at Once Benchagents: Automated benchmark creation with agent interaction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4ad58176-15f9-4c22-806d-0f3e7d0d395f · outbound
Benchmark Everything Everywhere All at Once ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f1489bf8-79d0-401b-b41a-8b7177db5e19 · outbound
Benchmark Everything Everywhere All at Once Mllm-as-a-judge: Assessing multimodal llm-as-a- judge with vision-language benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 768d7b4f-9bec-4cf9-b8b0-a431fe83a748 · outbound
Benchmark Everything Everywhere All at Once Are we on the right way for evaluating large vision-language models?NeurIPS, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5db47bd7-7218-4732-b59d-ef9a38d09c68 · outbound
Benchmark Everything Everywhere All at Once Can Large Language Models Be an Alternative to Human Evaluations?
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 388e95ee-e877-481f-8495-7b254550c651 · outbound
Benchmark Everything Everywhere All at Once Cl-bench: A benchmark for context learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1fd43dfd-dcc5-4ffa-8658-3d91369f039b · outbound
Benchmark Everything Everywhere All at Once On path to multimodal generalist: General-level and general-bench
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e701b4a-7098-42e5-ba26-00e5195a9927 · outbound
Benchmark Everything Everywhere All at Once Gptscore: Evaluate as you desire
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 135ee187-2bb9-4a4b-8dd6-6f2edba2bb89 · outbound
Benchmark Everything Everywhere All at Once Gemini 3 pro model card
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7563fcb-991c-420c-82b9-1ccdefeede39 · outbound
Benchmark Everything Everywhere All at Once Measuring Massive Multitask Language Understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 633ea2af-0f31-417c-a0b4-4da42887e5f3 · outbound
Benchmark Everything Everywhere All at Once Cerebellar output shapes cortical preparatory activity during motor adaptation.Nature Communications, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbde44d1-873b-4a2a-a320-a89eb2c38689 · outbound
Benchmark Everything Everywhere All at Once Kimi K2: Open Agentic Intelligence
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 37e4bbe7-3d81-4191-ae13-f8c311d235a8 · outbound
Benchmark Everything Everywhere All at Once Dataflow: An llm-driven framework for unified data preparation and workflow automation in the era of data-centric ai
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 669c9553-7c22-4295-b5dd-648e1f15482f · outbound
Benchmark Everything Everywhere All at Once Act as human: Multimodal large language model data annotation with critical thinking.arXiv preprint arXiv:2511.09833, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e30cf7db-2964-49d8-8f2b-d15bb965016c · outbound
Benchmark Everything Everywhere All at Once Agentbench: Evaluating llms as agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c665b82b-7963-4cfa-a425-243e42b5182d · outbound
Benchmark Everything Everywhere All at Once Mmbench: Is your multi-modal model an all-around player? InECCV, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b02423a-e7ea-4845-aee9-98e2f7cd37c4 · outbound
Benchmark Everything Everywhere All at Once MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0d5f8588-81d8-460b-a234-dde1141d3f2f · outbound
Benchmark Everything Everywhere All at Once Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dbc013b3-4552-496a-a0ed-306f1f7f7d98 · outbound
Benchmark Everything Everywhere All at Once UnitCoder: Scalable Iterative Code Synthesis with Unit Test Guidance
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ddbe4b47-6817-45c7-bee9-db926870d4ef · outbound
Benchmark Everything Everywhere All at Once MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5d280fe4-9844-4fe7-a8f2-a4d22c845091 · outbound
Benchmark Everything Everywhere All at Once Autonomous Evaluation and Refinement of Digital Agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c5408104-8e3a-4565-bad1-90e0e568c3b3 · outbound
Benchmark Everything Everywhere All at Once Benchmarkˆ 2: Systematic evaluation of llm benchmarks.arXiv preprint arXiv:2601.03986, 2026
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c1fc5154-382b-42d9-8cf0-c8e5b7ad21f9 · outbound
Benchmark Everything Everywhere All at Once Autobench: Automatic testbench generation and evaluation using llms for hdl design
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a8ce739-6907-4129-aafb-40fbf6c90da8 · outbound
Benchmark Everything Everywhere All at Once Qwen2 Technical Report
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2ecb2511-f8e8-4118-95b7-535678e94aa7 · outbound
Benchmark Everything Everywhere All at Once Qwen3.6-Plus: Towards real world agents, April 2026
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61b8177-133d-4102-bd6b-e94a801eb13a · outbound
Benchmark Everything Everywhere All at Once GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 457f59f2-cb04-48de-b19a-59c3bc8710f2 · outbound
Benchmark Everything Everywhere All at Once Neuronal dynamics of cerebellum and medial prefrontal cortex in adaptive motor timing.Nature Commu- nications, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9217cc7-ee50-47bc-82a2-f9083a7fe733 · outbound
Benchmark Everything Everywhere All at Once TAGAL: Tabular Data Generation using Agentic LLM Methods
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e9961485-1c81-4be7-86b5-6715f7991c79 · outbound
Benchmark Everything Everywhere All at Once One-eval: An agentic system for automated and traceable llm evaluation.arXiv preprint arXiv:2603.09821, 2026
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation af2d7b06-9ba6-4c0c-9441-0e8fef07efd9 · outbound
Benchmark Everything Everywhere All at Once OpenAI GPT-5 System Card
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 658fff00-4afe-4123-8d5b-e3a75383f82a · outbound
Benchmark Everything Everywhere All at Once Towards vqa models that can read
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 731e3fc0-afdf-4757-9bd1-4ec1b37d7e34 · outbound
Benchmark Everything Everywhere All at Once Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9131cd49-9410-49ce-a764-9bd7f6103cf2 · outbound
Benchmark Everything Everywhere All at Once Spacevista: All-scale visual spatial reasoning from mm to km.ICML, 2025
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90dccf7-b1b2-42ea-95ad-6cb33a07ac0e · outbound
Benchmark Everything Everywhere All at Once RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3f29bb34-5214-4f29-9e2e-4fe9b0dd9799 · outbound
Benchmark Everything Everywhere All at Once AI-Researcher: Autonomous Scientific Innovation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40a447e6-993b-4698-9229-e4f57f43f647 · outbound
Benchmark Everything Everywhere All at Once Qwen3.5: Accelerating productivity with native multimodal agents, February
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a39ce569-e168-4b2a-a9d3-a2f1f5617958 · outbound
Benchmark Everything Everywhere All at Once Unresolved cited work
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5913dc-a94b-44d0-b720-99c0414fd1d9 · outbound
Benchmark Everything Everywhere All at Once Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration.NeurIPS, 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fdf7826-c3b7-4a84-9710-e209c26340e6 · outbound
Benchmark Everything Everywhere All at Once Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6443c399-756a-496d-9d80-1cc4671ee353 · outbound
Benchmark Everything Everywhere All at Once InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3d55a4d1-7b9f-40b4-9326-73aecf0dd419 · outbound
Benchmark Everything Everywhere All at Once Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.NeurIPS, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36e5be77-918a-42c5-8775-12e586c47bf1 · outbound
Benchmark Everything Everywhere All at Once Finevision: Open data is all you need,
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9e453a-0d14-49e1-87ac-baa3d9589898 · outbound
Benchmark Everything Everywhere All at Once FineVision: Open Data Is All You Need
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0254e79d-595d-4b03-ad8b-2fe13bcbb4ad · outbound
Benchmark Everything Everywhere All at Once Language prompt for autonomous driving
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a568e52c-c097-470b-b9d6-29a1d16ce8af · outbound
Benchmark Everything Everywhere All at Once Qwen3 Technical Report
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d6ef1379-6292-4bd3-b045-8d04ff4a8af5 · outbound
Benchmark Everything Everywhere All at Once From Web to Pixels: Bringing Agentic Search into Visual Perception
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ce0cb930-e299-4436-970d-e04dd1ba1e54 · outbound
Benchmark Everything Everywhere All at Once Swe-agent: Agent-computer interfaces enable automated software engineering.NeurIPS, 2024
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fc4f0f7-16ad-462b-917e-e9fbe6a1b066 · outbound
Benchmark Everything Everywhere All at Once React: Synergizing reasoning and acting in language models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9edae0f2-f3d7-4a2a-96d1-3fffd218b7ef · outbound
Benchmark Everything Everywhere All at Once Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8966b782-672b-42b9-8d31-46d2ecd65cd6 · outbound
Benchmark Everything Everywhere All at Once Evaluation agent: Efficient and promptable evaluation framework for visual generative models
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5d74065-ab02-45f3-8954-d77a1791aeef · outbound
Benchmark Everything Everywhere All at Once Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InECCV, 2024
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 649636fe-f6c3-4c70-84f5-13dd6d08f569 · outbound
Benchmark Everything Everywhere All at Once Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?ICLR, 2025
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4110844-b819-4354-a6e6-7c3ffe1e643c · outbound
Benchmark Everything Everywhere All at Once Dual and plasticity-dependent regulation of cerebello-zona incerta circuits on anxiety-like behaviors.Nature communications, 2025
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb49017-3926-4ed8-9eb3-4492b1769d0e · outbound
Benchmark Everything Everywhere All at Once Judging llm-as-a-judge with mt-bench and chatbot arena.NeurIPS, 2023
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcce68a2-5c0b-41d5-ad47-c87a8d197f85 · outbound
Benchmark Everything Everywhere All at Once Dyval: Dynamic evaluation of large language models for reasoning tasks
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f81f8836-6be6-4578-9110-59818dba5789 · outbound
Benchmark Everything Everywhere All at Once JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 28547017-5184-4f09-8b7d-5e50d4347339 · outbound
Benchmark Everything Everywhere All at Once Q., and Shou, M
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5489fc31-c597-4ff0-9975-edd479e72c4c · outbound
Benchmark Everything Everywhere All at Once at the same time
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0f61985-60e2-42d7-a880-9f575693589f · inbound
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs Benchmark Everything Everywhere All at Once
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b4e567e-dc86-4d7c-b7da-0460d0f2c561 · inbound
NexForge: Scaling Agent Capabilities through Requirement-Driven Task Synthesis for LLMs Benchmark Everything Everywhere All at Once
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.