Pith. sign in

REVIEW 32 references

Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering

T0 review · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A controlled study shows that for small models, the best table representation depends on table size and question complexity, and the proposed FRES selection rule improves accuracy by about 10 points on average.

arxiv 2505.14131 v1 pith:LDTPMLSD submitted 2025-05-20 cs.CL

classification cs.CL
keywords tablemodelsrepresentationsimagesquestionanalysisansweringcontrolled
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Table question answering (TQA) asks a model to answer a question given a table. The table can be fed to the model as serialized text or as an image, and the model can be a text-only large language model (LLM) or a multimodal one (MLLM). Prior work reported that images and text perform similarly, but the comparisons were not controlled for properties of the question or the table.

This paper builds a benchmark of 1600 instances from six existing TQA datasets, split into four controlled settings based on question complexity (retrieval versus reasoning) and table size (small versus big). It evaluates six small model pairs (4B to 12B parameters) and one large 72B pair. For the large model, images consistently beat text. For the small models, the best input changes: big tables do best with text passed to an MLLM; small tables with reasoning questions do best when text and image are both passed; small tables with retrieval questions only need text.

The authors then define FRES, a simple rule that reproduces these choices using table size and a question-type classifier. On four test sets, FRES beats the 'pass everything' baseline by an average of about 10 exact-match points, and it reduces the number of input tokens. The gains are mostly driven by a fine-tuned model, TableLlaVA, where feeding both text and image on large tables hurt accuracy substantially.

Extended reading notes

Core claim

The paper's central claim is that the optimal table representation for small TQA models varies by table size and question complexity: 'with big tables, using text representations with MLLMs leads to the best performance... with small tables, if the question type is reasoning, providing MLLMs with both table representations results in optimal performance.' This pattern is then packaged into FRES, which the paper says gives 'an average of 10% exact match gain compared to baseline approaches.'

Load-bearing premise

The load-bearing premise is that the question-complexity classifier, taken from the authors' own prior work (Zhou et al., 2024) and reused here with a Qwen-2-72B backbone, reliably separates retrieval from reasoning questions (reported accuracy 93% on 200 instances). This classifier labels both the benchmark settings in Section 2.3 and the FRES decisions in Section 4. If it is biased on the test sets, the observed conditional patterns and the FRES gains could be partly an artifact of mislabeling, not a property of inputs and models.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

No new physical or structural entities are introduced; FRES is a decision rule, not a postulated entity. The free parameters are thresholds and a rule chosen by the authors from their own data and prior work.

free parameters (3)
  • Table size thresholds (pixels, tokens) = 2e6 pixels, 288 tokens
    Averages over MMTab; used to split small vs big tables in the benchmark and in FRES. These are hand-chosen thresholds that condition all results.
  • Contamination cutoff = 20% accuracy under masked inputs
    Models with masked-input accuracy above 20% are excluded, and instances answerable by any model under masking are dropped. This threshold shapes the evaluation set.
  • FRES decision rule = Text for big; both for small+reasoning; text for small+retrieval
    The rule is derived from the authors' empirical observations on their benchmark rather than from a separate training set. It is a manually specified mapping from settings to representations.
assumptions (5)
  • domain assumption Retrieval vs reasoning is a sufficient dichotomy for question complexity.
    Section 2.2 defines these two categories and uses them to split the benchmark and to design FRES.
  • domain assumption The question classifier from Zhou et al. (2024) reliably labels retrieval/reasoning at 93% accuracy.
    Section 2.3 and A.4 reuse the authors' prior classifier, re-backed by Qwen-2-72B, with accuracy checked on 200 instances.
  • domain assumption The six small model pairs are representative of small open-weight MLLMs/LLMs.
    Section 3.1 selects six small pairs satisfying the contamination filter; conclusions about 'small models' generalize from this sample.
  • domain assumption Masking out questions or tables adequately detects pre-training contamination.
    Section 3.1 uses masked-input accuracy (<=20%) as a contamination filter.
  • domain assumption MMTab average size statistics generalize to the TQA datasets.
    Section A.5 computes thresholds from MMTab and applies them to WTQ, TabFact, HiTab, WikiSQL.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering." pith.science (2026). https://pith.science/paper/LDTPMLSD

@misc{pith2026250514131,
  author       = {Pith},
  title        = {Pith review of: Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/LDTPMLSD}},
  note         = {Machine review of arXiv:2505.14131}
}
read the original abstract

In table question answering (TQA), tables are encoded as either texts or images. Prior work suggests that passing images of tables to multi-modal large language models (MLLMs) performs comparably to or even better than using textual input with large language models (LLMs). However, the lack of controlled setups limits fine-grained distinctions between these approaches. In this paper, we conduct the first controlled study on the effectiveness of several combinations of table representations and models from two perspectives: question complexity and table size. We build a new benchmark based on existing TQA datasets. In a systematic analysis of seven pairs of MLLMs and LLMs, we find that the best combination of table representation and model varies across setups. We propose FRES, a method selecting table representations dynamically, and observe a 10% average performance improvement compared to using both representations indiscriminately.

Figures

Figures reproduced from arXiv: 2505.14131 by the authors.

Figure 1
Figure 1. Varying exact match (EM) for models and table representations under different settings (i and t stand for image and text representations of tables). We categorize our investigation into four settings based on table size (small or big) and question complexity (retrieval or reasoning). to table size, hinders a deeper understanding of their strengths and weaknesses. Moreover, existing investigations focus exclusively o… view at source ↗
Figure 2
Figure 2. Evaluation of table size robustness. The bar plot shows the number of instances sampled for each bin, and the line plots show the performance of different approaches against varying table sizes. ing both table representations better triggers the reasoning abilities of MMLMs, thus leading to su￾perior performance on reasoning questions. How￾ever, in terms of information retrieval, the input is best represented as tex… view at source ↗
Figure 3
Figure 3. Different table templates. Dataset Licenses We build our evaluation dataset based on a subset of MMTab (Zheng et al., 2024), CRT (Zhang et al., 2023) and TempTabTQA (Gupta et al., 2023). The three datasets are pub￾licly available under the licenses of APACHE-2.06 , MIT7 and CC-BY-4.08 , respectively. In terms of the test datasets: WTQ (Pasupat and Liang, 2015), TabFact (Chen et al., 2020), HiTab (Cheng et al., 2022)… view at source ↗
Figures from the paper (3 more)
Figure 5
Figure 5. Figure 5: Evaluation of table size robustness. The bar [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 4
Figure 4. Figure 4: Resolution distribution of MMTab. not in a table, a question is classified as a reasoning question. If a question contains comparative terms (we detect it using NLTK), the question is classified as a reasoning question. Next, an LLM takes in a question and a table and …
Figure 6
Figure 6. Figure 6: Prompts and table formats. Dataset #cell #table-text #table-img %small table %reasoning question %small_reasoning WTQ 165 429 2.4e6 57.7 71.4 41.6 TabFact 94 258 1.8e6 69.3 51.2 35.8 HiTab 180 360 3.5e6 54.5 46.0 26.0 WikiSQL 95 227 1.7e6 71.3 29.8 22.0 [PITH_FULL_IMA…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

32 extracted references · 6 canonical work pages

  1. [1]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Hewett, Jamie Huynh, Mojan Javaheripi, Xin Jin, Piero Kauffmann, Nikos Karampatziakis, Dongwoo Kim, Young Jin Kim, Mahoud Khademi, Lev Kurilenko, James R

    Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Hassan Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Singh Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, S \'e bastien Bubeck, Martin Cai, Caio C'esar Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allison Del Giorno, Gustavo de Rosa, Matthe...

  4. [4]

    Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774

  5. [5]

    Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Devendra Singh Chaplot, Jessica Chudnovsky, Saurabh Garg, Th \'e ophile Gervet, Soham Ghosh, Am'elie H'eliou, Paul Jacob, Albert Q. Jiang, Timoth \'e e Lacroix, Guillaume Lample, Diego de Las Casas, Thibaut Lavril, Teven Le Scao, Andy Lo, William Marshall, Louis Martin, Arthur Mensch, Pavankumar Reddy Mudd...

  6. [6]

    I \ n igo Alonso, Eneko Agirre, and Mirella Lapata. 2024. https://doi.org/10.18653/v1/2024.acl-long.364 P ix T 3: Pixel-based table-to-text generation . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6721--6736, Bangkok, Thailand. Association for Computational Linguistics

  7. [7]

    Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiao wen Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Penglong Jiao, Zhen Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li, Shua...

  8. [8]

    Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact : A large-scale dataset for table-based fact verification. In International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia

Show all 32 references
  1. [9]

    Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.300 F in QA : A dataset of numerical reasoning over financial d...

  2. [10]

    Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, and Dongmei Zhang. 2022. https://doi.org/10.18653/v1/2022.acl-long.78 H i T ab: A hierarchical table dataset for question answering and natural language generation . In Proceedings of...

  3. [11]

    Naihao Deng, Zhenjie Sun, Ruiqi He, Aman Sikka, Yulong Chen, Lin Ma, Yue Zhang, and Rada Mihalcea. 2024. https://doi.org/10.18653/v1/2024.findings-acl.23 Tables as texts or images: Evaluating the table reasoning ability of LLM s and MLLM s . In Findings of the Association for ...

  4. [12]

    Vivek Gupta, Pranshu Kandoi, Mahek Vora, Shuo Zhang, Yujie He, Ridho Reinanda, and Vivek Srikumar. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.149 T emp T ab QA : Temporal question answering for semi-structured tables . In Proceedings of the 2023 Conference on Empirical ...

  5. [13]

    Jonathan Herzig, Pawel Krzysztof Nowak, Thomas M \"u ller, Francesco Piccinno, and Julian Eisenschlos. 2020. https://doi.org/10.18653/v1/2020.acl-main.398 T a P as: Weakly supervised table parsing via pre-training . In Proceedings of the 58th Annual Meeting of the Association ...

  6. [14]

    Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...

  7. [15]

    Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, and Weizhu Chen. 2022. https://doi.org/10.18653/v1/2022.naacl-main.68 O mni T ab: Pretraining with natural and synthetic data for few-shot table-based question answering . In Proceedings of the 2022 Conference of the North A...

  8. [16]

    Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024. https://api.semanticscholar.org/CorpusID:271088459 Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models . ArXiv, abs/2407.07895

  9. [17]

    Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2023. https://api.semanticscholar.org/CorpusID:265150038 Monkey: Image resolution and text label are important things for large multi-modal models . 2024 IEEE/CVF Conferen...

  10. [18]

    Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. https://openreview.net/forum?id=O50443AsCP TAPEX : Table pre-training via learning a neural SQL executor . In International Conference on Learning Representations

  11. [19]

    Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and A. Kalyan. 2022. https://api.semanticscholar.org/CorpusID:252595921 Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning . ArXiv, abs/2209.14610

  12. [20]

    Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.483 SCITAB : A challenging benchmark for compositional reasoning and claim verification on scientific tables . In Proceedings of the 2023 Conference on Empiri...

  13. [21]

    Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.89 ToTTo : A controlled table-to-text generation dataset . In Proceedings of the 2020 Conference on Empirical Methods i...

  14. [22]

    Panupong Pasupat and Percy Liang. 2015. https://doi.org/10.3115/v1/P15-1142 Compositional semantic parsing on semi-structured tables . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natur...

  15. [23]

    Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805

  16. [24]

    Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Ke-Yang Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. https://api.semanticscholar.org/Corpus...

  17. [25]

    Wilcoxon

    Frank. Wilcoxon. 1945. https://api.semanticscholar.org/CorpusID:53662922 Individual comparisons by ranking methods . Biometrics, 1:196--202

  18. [26]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng...

  19. [27]

    Team Glm Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Ming yue Liu, Minlie H...

  20. [28]

    Zhehao Zhang, Xitao Li, Yan Gao, and Jian-Guang Lou. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.132 CRT - QA : A dataset of complex reasoning question answering over tabular data . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing...

  21. [29]

    Mingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She, Zheng Lin, Wenbin Jiang, and Weiping Wang. 2024. https://doi.org/10.18653/v1/2024.acl-long.493 Multimodal table understanding . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volum...

  22. [30]

    Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. CoRR, abs/1709.00103

  23. [31]

    Wei Zhou, Mohsen Mesgar, Heike Adel, and Annemarie Friedrich. 2024. https://doi.org/10.18653/v1/2024.naacl-long.137 FREB - TQA : A fine-grained robustness evaluation benchmark for table question answering . In Proceedings of the 2024 Conference of the North American Chapter of...

  24. [32]

    Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji rong Wen. 2023. https://api.semanticscholar.org/CorpusID:260887838 Large language models for information retrieval: A survey . ArXiv, abs/2308.07107

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.