Pith. sign in

Paper Citation Record · LEDGER

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

As of 7 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18421 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:04.057694Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T22:18:45.189576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:05:46.699261Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b3a2a05-58ce-4e41-97c9-b4b49af9becd · outbound

This paper cites GPT-4 Technical Report.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:00.709832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:00.709832Z digest=sha256:d5634f750d5a39c699a37fdd98b0f341e96feeefc4b57d75566d6db6b42c8d49

Observation 338be5f1-e943-4a6f-8e59-060ada6d38a0 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.001340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.001340Z digest=sha256:cd6d7db797fd44d7045bd3f4e3118eb7d632fa869b419444ac34534615d27519

Observation f5a0899b-0996-452e-a24a-0004dc95ce6a · outbound

This paper cites R., et al.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models R., et al

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:05.462141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:01.051627Z digest=sha256:130a75b8aaa09d3f83f4d052c391f738004aaf06f912ab0d2e473a2783e673a1

Observation 4744b0a9-e1cc-423d-be1d-4ab347163d09 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.103215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.103215Z digest=sha256:1f3ca6f842ea7b8d4f5965238b880f248ad155f261317f0a2f819a8053a504e4

Observation 37bf6e34-0550-4b47-9a1c-d55322eeacbc · outbound

This paper cites RLHF Workflow: From Reward Modeling to Online RLHF.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models RLHF Workflow: From Reward Modeling to Online RLHF

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.350485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.350485Z digest=sha256:d13a2d1185d4d1fcb8f919321ebbe5aa33f9ddf17c6373715cb62828ed45c536

Observation da391928-f427-42e1-b027-8a07c6e7404f · outbound

This paper cites The Llama 3 Herd of Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.408209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.408209Z digest=sha256:dadcd82589446407c634f2fdbf02113a4d3bb265684ba1de74d0803ce2f2450a

Observation 27e5dc92-2636-4092-8a11-e867b2856365 · outbound

This paper cites A Survey on LLM-as-a-Judge.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models A Survey on LLM-as-a-Judge

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.489710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.489710Z digest=sha256:577b1c84941d1291d7bfbf717ee3e570fbfaede611b0788d49edb31b7497c537

Observation 8735e25e-e6f3-4e3b-9d82-d8a27040ef97 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.630884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.630884Z digest=sha256:96d6e981b981e711d97b2704b4b48032595c5e4ea1c809fa2c70a43a9db54ac0

Observation 28ca5b0e-bb76-4375-bb69-a5af0144a6fe · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.729656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.729656Z digest=sha256:922b8894d2d8ac976ab4744e6575771b5d00957c3f1931da0bcfe415233b352f

Observation d4799446-eba0-46b9-abc3-7b18fee89f36 · outbound

This paper cites Qwen2.5-Coder Technical Report.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Qwen2.5-Coder Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.826117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.826117Z digest=sha256:9c9722a3265971298789f0aa8d2ea377f1a0def9a153e130e48c2d3285d166e9

Observation 80cd3bb1-90be-4ef7-bb44-8c12aa87e8be · outbound

This paper cites FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.097569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.097569Z digest=sha256:6c213724a9885085447eab6dc2ab50dfc62bc98fbc8d36743ac69777146e7131

Observation 8d00a86d-a6ac-4f38-ac47-b5ce33ac0601 · outbound

This paper cites Ait-qa: Question answering dataset over complex tables in the airline industry.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Ait-qa: Question answering dataset over complex tables in the airline industry

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:05.113736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:02.229374Z digest=sha256:4bec6c72290699082decdba72e43a1deba7f5940eee3ae72b82d8e0a6967cdd6

Observation 12d33cae-f6e2-4ee8-8fb0-29985d374f4c · outbound

This paper cites TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.470277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.470277Z digest=sha256:34c8b5826587d4a34c463ca7bc8fb10bf99d76ca434d68ea50c54626cf666bc1

Observation d897cdce-52af-46b2-ac59-5c62e1a1d68f · outbound

This paper cites TableQAKit: A Comprehensive and Practical Toolkit for Table-based Question Answering.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableQAKit: A Comprehensive and Practical Toolkit for Table-based Question Answering

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.598781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.598781Z digest=sha256:738f2135fe00fdbff218490ca50f981c08b51e4862e5a562441aed7b1825a36f

Observation fa4b1cff-1354-4f94-a1b6-c2ffd55a2a6d · outbound

This paper cites UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models UHGEval: Benchmarking the Hallucination of Chinese Large Language Models via Unconstrained Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:20:04.474776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:02.713992Z digest=sha256:f9ae8a81947d992826f09ac99eae913efb0177cb296a8160bc61db6771af5f36

Observation 86b4b131-021c-4617-80dd-67208fb4602b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.205192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.205192Z digest=sha256:1ff5ba633791f72654c2c3707e9143f0cb2b2664dc5c22e62c543e2edaff6984

Observation bdf2d175-1d2b-4802-868a-b97e70998156 · outbound

This paper cites TableGPT2: A Large Multimodal Model with Tabular Data Integration.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TableGPT2: A Large Multimodal Model with Tabular Data Integration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.333517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.333517Z digest=sha256:223b7acc881fb3a918cd274a4c62b92d14a59d9e1bcc3f84fd50034202f5283e

Observation 6f6c6654-9fed-443f-8b78-24f80ffffc80 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.461424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.461424Z digest=sha256:de582bac77085772ec9738d3f7ec215d93e167e5de73d7721010688711d81520

Observation 30c8885d-9fdf-40f1-b468-994f82b39700 · outbound

This paper cites MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MAC-SQL: A Multi-Agent Collaborative Framework for Text-to-SQL

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.588865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.588865Z digest=sha256:b5bc1dea4ad24a214d44afc4931d4a22dc9a32c9beddfa479c463d0b8b7b3030

Observation c82b98e5-bc45-4474-9226-1faf9aee2f94 · outbound

This paper cites Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.663026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.663026Z digest=sha256:1070619c64ed20649b3e786418764903032bee13b171efb5abe90ca818f6c23b

Observation b8ee7cea-3bc6-4461-85a4-b7c960ad444e · outbound

This paper cites Qwen2 Technical Report.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Qwen2 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.756297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.756297Z digest=sha256:3070e3831456ba4d54cee8236acebc79a24368c614bca7ceac6d70b3ebd65f62

Observation d44e477d-a053-445e-b0d9-c49b8f7aa476 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Yi: Open Foundation Models by 01.AI

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.872378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.872378Z digest=sha256:033e834bad4a58fc488a939749e68a6ffdb447070593568c950392114b702993

Observation 5424cf12-9e9c-4ac5-996c-feb3e76872a5 · outbound

This paper cites Spider: A large-scale human-labeled dataset for complex and cross- domain semantic parsing and text-to-sql task.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Spider: A large-scale human-labeled dataset for complex and cross- domain semantic parsing and text-to-sql task

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:04.802096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:03.958511Z digest=sha256:209ed3a2f96a32ff438af154f4fe42931135e540e097ecbced43df2a345c69bc

Observation 409a6220-26df-4c16-b46e-06bbdf72dbe6 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:04.057694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:04.057694Z digest=sha256:301a50feeba190bf47cf98367ec7c4e54ebc1b3f5f0d4def198368384269a252

Observation 85074de9-9cc1-4f2f-a4cd-dd4e197a6f2a · outbound

This paper cites Totto: A controlled table- to-text generation dataset.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Totto: A controlled table- to-text generation dataset

Reference 2002

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:04.964808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:02.940097Z digest=sha256:67f506985e19eb60874e4d056e97856bbf76f3955360daab6047128dba37aa2c

Observation 193dd8cc-9d80-4f51-8211-c7db35f3a691 · outbound

This paper cites MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.838439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.838439Z digest=sha256:f89bde15c3f3b87eb91f12150518665f9b682e0c478be011263820c0ff6cd927

Observation 111542bc-8bb5-4529-a60e-5f500ef79af2 · outbound

This paper cites TQA-Bench: Evaluating LLMs for Multi-Table Question Answering.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models TQA-Bench: Evaluating LLMs for Multi-Table Question Answering

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:03.071359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:03.071359Z digest=sha256:c46ef4c9e23a0ec42a46d96273432a23b9169935d726efe9f925b51feb23e628

Observation f7b13eb3-f7c8-4ef6-aa61-1394f2fbae1f · outbound

This paper cites Mistral 7B.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Mistral 7B

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.965467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.965467Z digest=sha256:7fd08172a8d683e18dd73082d559189210ca1c6232740e32135cd9b79a3d7ad2

Observation 0f7daa07-7220-4925-a7c1-5faf4bcceee0 · outbound

This paper cites McEval: Massively Multilingual Code Evaluation.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models McEval: Massively Multilingual Code Evaluation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:00.861149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:00.861149Z digest=sha256:941688b2b906b82db95817bbaabc19db40f7aca52b2318dbe4460483df8ec040

Observation bf0866df-e83f-417e-ba0d-bdbccb941bb4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:01.190008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:01.190008Z digest=sha256:eaef66123e7728e1024f0a926cf5208d83a22aa2010de38620e50c575b1583cc

Observation 48d2a58e-7701-44fd-a6bd-4891431660d6 · outbound

This paper cites Table Foundation Models: on knowledge pre-training for tabular learning.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Table Foundation Models: on knowledge pre-training for tabular learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.370073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.370073Z digest=sha256:995ad0c79c17e50ea559fd7c0f5cafec308b08f0ef7908443cb7b3598e95c6bf

Observation 1d4d8fcb-ca5b-4c36-b977-d8630df5f8ee · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:00.771769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:00.771769Z digest=sha256:cea4c6531bdf9cfd4a1d884b2512aae592fa751f8c992a74708a3e26bb7888a9

Observation 84d11dd7-80f8-4fce-929c-67c01f7d8945 · outbound

This paper cites an unresolved cited work.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:05.669312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:00.940491Z digest=sha256:4f6bc06ecd18593fa8224515b6e45d162d234578056087613a38d5d8644fa506

Observation c8c641df-fc1f-43b5-ae47-b30e5d7f0afa · outbound

This paper cites Tables as texts or images: Evaluating the table reasoning ability of llms and mllms.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models Tables as texts or images: Evaluating the table reasoning ability of llms and mllms

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:05.289155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:20:01.257151Z digest=sha256:0a7427c501526453c8f4c0bbc99619a142d56a7f5be799ecca97227c9f797ef2

Pith citing papers

Observation 014cfda9-bca4-4b8a-98b3-9f9f3905be6e · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-22T00:22:12.404705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:56:36.312877Z digest=sha256:f0d265ec1481c1429b837a88e687ab2b21cb445231d8f6e7ba5c4ae8b160332a

Observation 5dad171a-f708-44a9-9e65-1a3ab37af9d1 · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-22T00:22:12.404705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:18:45.189576Z digest=sha256:19ca993a9a422c6a7043171397b5e91e5c32ddb3dcee75bab46e5420fd324f00