Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:04.057694Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18421.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:04.057694Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-30T22:18:45.189576Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T14:05:46.699261Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4b3a2a05-58ce-4e41-97c9-b4b49af9becd · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338be5f1-e943-4a6f-8e59-060ada6d38a0 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a0899b-0996-452e-a24a-0004dc95ce6a · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models R., et al
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4744b0a9-e1cc-423d-be1d-4ab347163d09 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37bf6e34-0550-4b47-9a1c-d55322eeacbc · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models RLHF Workflow: From Reward Modeling to Online RLHF
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da391928-f427-42e1-b027-8a07c6e7404f · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27e5dc92-2636-4092-8a11-e867b2856365 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models A Survey on LLM-as-a-Judge
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8735e25e-e6f3-4e3b-9d82-d8a27040ef97 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28ca5b0e-bb76-4375-bb69-a5af0144a6fe · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4799446-eba0-46b9-abc3-7b18fee89f36 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Qwen2.5-Coder Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80cd3bb1-90be-4ef7-bb44-8c12aa87e8be · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d00a86d-a6ac-4f38-ac47-b5ce33ac0601 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Ait-qa: Question answering dataset over complex tables in the airline industry
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 12d33cae-f6e2-4ee8-8fb0-29985d374f4c · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d897cdce-52af-46b2-ac59-5c62e1a1d68f · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableQAKit: A Comprehensive and Practical Toolkit for Table-based Question Answering
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa4b1cff-1354-4f94-a1b6-c2ffd55a2a6d · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86b4b131-021c-4617-80dd-67208fb4602b · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf2d175-1d2b-4802-868a-b97e70998156 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableGPT2: A Large Multimodal Model with Tabular Data Integration
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f6c6654-9fed-443f-8b78-24f80ffffc80 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c8885d-9fdf-40f1-b468-994f82b39700 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82b98e5-bc45-4474-9226-1faf9aee2f94 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ee7cea-3bc6-4461-85a4-b7c960ad444e · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Qwen2 Technical Report
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d44e477d-a053-445e-b0d9-c49b8f7aa476 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Yi: Open Foundation Models by 01.AI
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5424cf12-9e9c-4ac5-996c-feb3e76872a5 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Spider: A large-scale human-labeled dataset for complex and cross- domain semantic parsing and text-to-sql task
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 409a6220-26df-4c16-b46e-06bbdf72dbe6 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85074de9-9cc1-4f2f-a4cd-dd4e197a6f2a · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Totto: A controlled table- to-text generation dataset
Reference 2002
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 193dd8cc-9d80-4f51-8211-c7db35f3a691 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 111542bc-8bb5-4529-a60e-5f500ef79af2 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TQA-Bench: Evaluating LLMs for Multi-Table Question Answering
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7b13eb3-f7c8-4ef6-aa61-1394f2fbae1f · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Mistral 7B
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f7daa07-7220-4925-a7c1-5faf4bcceee0 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models McEval: Massively Multilingual Code Evaluation
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf0866df-e83f-417e-ba0d-bdbccb941bb4 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d2a58e-7701-44fd-a6bd-4891431660d6 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Table Foundation Models: on knowledge pre-training for tabular learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d4d8fcb-ca5b-4c36-b977-d8630df5f8ee · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d11dd7-80f8-4fce-929c-67c01f7d8945 · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c8c641df-fc1f-43b5-ae47-b30e5d7f0afa · outbound
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Tables as texts or images: Evaluating the table reasoning ability of llms and mllms
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 014cfda9-bca4-4b8a-98b3-9f9f3905be6e · inbound
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5dad171a-f708-44a9-9e65-1a3ab37af9d1 · inbound
ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.