REVIEW 32 references
Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering
T0 review · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A controlled study shows that for small models, the best table representation depends on table size and question complexity, and the proposed FRES selection rule improves accuracy by about 10 points on average.
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
This paper builds a benchmark of 1600 instances from six existing TQA datasets, split into four controlled settings based on question complexity (retrieval versus reasoning) and table size (small versus big). It evaluates six small model pairs (4B to 12B parameters) and one large 72B pair. For the large model, images consistently beat text. For the small models, the best input changes: big tables do best with text passed to an MLLM; small tables with reasoning questions do best when text and image are both passed; small tables with retrieval questions only need text.
The authors then define FRES, a simple rule that reproduces these choices using table size and a question-type classifier. On four test sets, FRES beats the 'pass everything' baseline by an average of about 10 exact-match points, and it reduces the number of input tokens. The gains are mostly driven by a fine-tuned model, TableLlaVA, where feeding both text and image on large tables hurt accuracy substantially.
Extended reading notes
Core claim
The paper's central claim is that the optimal table representation for small TQA models varies by table size and question complexity: 'with big tables, using text representations with MLLMs leads to the best performance... with small tables, if the question type is reasoning, providing MLLMs with both table representations results in optimal performance.' This pattern is then packaged into FRES, which the paper says gives 'an average of 10% exact match gain compared to baseline approaches.'
Load-bearing premise
The load-bearing premise is that the question-complexity classifier, taken from the authors' own prior work (Zhou et al., 2024) and reused here with a Qwen-2-72B backbone, reliably separates retrieval from reasoning questions (reported accuracy 93% on 200 instances). This classifier labels both the benchmark settings in Section 2.3 and the FRES decisions in Section 4. If it is biased on the test sets, the observed conditional patterns and the FRES gains could be partly an artifact of mislabeling, not a property of inputs and models.
Editorial analysis
A structured set of objections, weighed in public.
Assumptions & free parameters
free parameters (3)
- Table size thresholds (pixels, tokens) =
2e6 pixels, 288 tokens
- Contamination cutoff =
20% accuracy under masked inputs
- FRES decision rule =
Text for big; both for small+reasoning; text for small+retrieval
assumptions (5)
- domain assumption Retrieval vs reasoning is a sufficient dichotomy for question complexity.
- domain assumption The question classifier from Zhou et al. (2024) reliably labels retrieval/reasoning at 93% accuracy.
- domain assumption The six small model pairs are representative of small open-weight MLLMs/LLMs.
- domain assumption Masking out questions or tables adequately detects pre-training contamination.
- domain assumption MMTab average size statistics generalize to the TQA datasets.
Cite this review
Pith. "Pith review of Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering." pith.science (2026). https://pith.science/paper/LDTPMLSD
@misc{pith2026250514131,
author = {Pith},
title = {Pith review of: Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering},
year = {2026},
howpublished = {\url{https://pith.science/paper/LDTPMLSD}},
note = {Machine review of arXiv:2505.14131}
}
read the original abstract
In table question answering (TQA), tables are encoded as either texts or images. Prior work suggests that passing images of tables to multi-modal large language models (MLLMs) performs comparably to or even better than using textual input with large language models (LLMs). However, the lack of controlled setups limits fine-grained distinctions between these approaches. In this paper, we conduct the first controlled study on the effectiveness of several combinations of table representations and models from two perspectives: question complexity and table size. We build a new benchmark based on existing TQA datasets. In a systematic analysis of seven pairs of MLLMs and LLMs, we find that the best combination of table representation and model varies across setups. We propose FRES, a method selecting table representations dynamically, and observe a 10% average performance improvement compared to using both representations indiscriminately.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Hassan Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Singh Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, S \'e bastien Bubeck, Martin Cai, Caio C'esar Teodoro Mendes, Weizhu Chen, Vishrav Chaudhary, Parul Chopra, Allison Del Giorno, Gustavo de Rosa, Matthe...
arXiv 2024
-
[4]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774
arXiv 2023
-
[5]
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna, Devendra Singh Chaplot, Jessica Chudnovsky, Saurabh Garg, Th \'e ophile Gervet, Soham Ghosh, Am'elie H'eliou, Paul Jacob, Albert Q. Jiang, Timoth \'e e Lacroix, Guillaume Lample, Diego de Las Casas, Thibaut Lavril, Teven Le Scao, Andy Lo, William Marshall, Louis Martin, Arthur Mensch, Pavankumar Reddy Mudd...
arXiv 2024
-
[6]
I \ n igo Alonso, Eneko Agirre, and Mirella Lapata. 2024. https://doi.org/10.18653/v1/2024.acl-long.364 P ix T 3: Pixel-based table-to-text generation . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 6721--6736, Bangkok, Thailand. Association for Computational Linguistics
-
[7]
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, Keyu Chen, Xin Chen, Xun Chen, Zehui Chen, Zhi Chen, Pei Chu, Xiao wen Dong, Haodong Duan, Qi Fan, Zhaoye Fei, Yang Gao, Jiaye Ge, Chenya Gu, Yuzhe Gu, Tao Gui, Aijia Guo, Qipeng Guo, Conghui He, Yingfan Hu, Ting Huang, Tao Jiang, Penglong Jiao, Zhen Jin, Zhikai Lei, Jiaxing Li, Jingwen Li, Linyang Li, Shua...
arXiv 2024
-
[8]
Wenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang, Hong Wang, Shiyang Li, Xiyou Zhou, and William Yang Wang. 2020. Tabfact : A large-scale dataset for table-based fact verification. In International Conference on Learning Representations (ICLR), Addis Ababa, Ethiopia
2020
Show all 32 references
-
[9]
Zhiyu Chen, Wenhu Chen, Charese Smiley, Sameena Shah, Iana Borova, Dylan Langdon, Reema Moussa, Matt Beane, Ting-Hao Huang, Bryan Routledge, and William Yang Wang. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.300 F in QA : A dataset of numerical reasoning over financial d...
2021 doi
-
[10]
Zhoujun Cheng, Haoyu Dong, Zhiruo Wang, Ran Jia, Jiaqi Guo, Yan Gao, Shi Han, Jian-Guang Lou, and Dongmei Zhang. 2022. https://doi.org/10.18653/v1/2022.acl-long.78 H i T ab: A hierarchical table dataset for question answering and natural language generation . In Proceedings of...
2022 doi
-
[11]
Naihao Deng, Zhenjie Sun, Ruiqi He, Aman Sikka, Yulong Chen, Lin Ma, Yue Zhang, and Rada Mihalcea. 2024. https://doi.org/10.18653/v1/2024.findings-acl.23 Tables as texts or images: Evaluating the table reasoning ability of LLM s and MLLM s . In Findings of the Association for ...
2024 doi
-
[12]
Vivek Gupta, Pranshu Kandoi, Mahek Vora, Shuo Zhang, Yujie He, Ridho Reinanda, and Vivek Srikumar. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.149 T emp T ab QA : Temporal question answering for semi-structured tables . In Proceedings of the 2023 Conference on Empirical ...
2023 doi
-
[13]
Jonathan Herzig, Pawel Krzysztof Nowak, Thomas M \"u ller, Francesco Piccinno, and Julian Eisenschlos. 2020. https://doi.org/10.18653/v1/2020.acl-main.398 T a P as: Weakly supervised table parsing via pre-training . In Proceedings of the 58th Annual Meeting of the Association ...
2020 doi
-
[14]
Albert Qiaochu Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L'elio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, T...
2023 arXiv
-
[15]
Zhengbao Jiang, Yi Mao, Pengcheng He, Graham Neubig, and Weizhu Chen. 2022. https://doi.org/10.18653/v1/2022.naacl-main.68 O mni T ab: Pretraining with natural and synthetic data for few-shot table-based question answering . In Proceedings of the 2022 Conference of the North A...
2022 doi
-
[16]
Feng Li, Renrui Zhang, Hao Zhang, Yuanhan Zhang, Bo Li, Wei Li, Zejun Ma, and Chunyuan Li. 2024. https://api.semanticscholar.org/CorpusID:271088459 Llava-next-interleave: Tackling multi-image, video, and 3d in large multimodal models . ArXiv, abs/2407.07895
2024 arXiv
-
[17]
Zhang Li, Biao Yang, Qiang Liu, Zhiyin Ma, Shuo Zhang, Jingxu Yang, Yabo Sun, Yuliang Liu, and Xiang Bai. 2023. https://api.semanticscholar.org/CorpusID:265150038 Monkey: Image resolution and text label are important things for large multi-modal models . 2024 IEEE/CVF Conferen...
2023
-
[18]
Qian Liu, Bei Chen, Jiaqi Guo, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, and Jian-Guang Lou. 2022. https://openreview.net/forum?id=O50443AsCP TAPEX : Table pre-training via learning a neural SQL executor . In International Conference on Learning Representations
2022
-
[19]
Pan Lu, Liang Qiu, Kai-Wei Chang, Ying Nian Wu, Song-Chun Zhu, Tanmay Rajpurohit, Peter Clark, and A. Kalyan. 2022. https://api.semanticscholar.org/CorpusID:252595921 Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning . ArXiv, abs/2209.14610
2022 arXiv
-
[20]
Xinyuan Lu, Liangming Pan, Qian Liu, Preslav Nakov, and Min-Yen Kan. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.483 SCITAB : A challenging benchmark for compositional reasoning and claim verification on scientific tables . In Proceedings of the 2023 Conference on Empiri...
2023 doi
-
[21]
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.89 ToTTo : A controlled table-to-text generation dataset . In Proceedings of the 2020 Conference on Empirical Methods i...
2020 doi
-
[22]
Panupong Pasupat and Percy Liang. 2015. https://doi.org/10.3115/v1/P15-1142 Compositional semantic parsing on semi-structured tables . In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natur...
2015 doi
-
[23]
Gemini Team, Rohan Anil, Sebastian Borgeaud, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, Katie Millican, et al. 2023. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805
2023 arXiv
-
[24]
Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Ke-Yang Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. 2024. https://api.semanticscholar.org/Corpus...
2024 arXiv
-
[25]
Wilcoxon
Frank. Wilcoxon. 1945. https://api.semanticscholar.org/CorpusID:53662922 Individual comparisons by ranking methods . Biometrics, 1:196--202
1945
-
[26]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jin Xu, Jingren Zhou, Jinze Bai, Jinzheng...
2024 arXiv
-
[27]
Team Glm Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Ming yue Liu, Minlie H...
2024 arXiv
-
[28]
Zhehao Zhang, Xitao Li, Yan Gao, and Jian-Guang Lou. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.132 CRT - QA : A dataset of complex reasoning question answering over tabular data . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing...
2023 doi
-
[29]
Mingyu Zheng, Xinwei Feng, Qingyi Si, Qiaoqiao She, Zheng Lin, Wenbin Jiang, and Weiping Wang. 2024. https://doi.org/10.18653/v1/2024.acl-long.493 Multimodal table understanding . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volum...
2024 doi
-
[30]
Victor Zhong, Caiming Xiong, and Richard Socher. 2017. Seq2sql: Generating structured queries from natural language using reinforcement learning. CoRR, abs/1709.00103
2017 arXiv
-
[31]
Wei Zhou, Mohsen Mesgar, Heike Adel, and Annemarie Friedrich. 2024. https://doi.org/10.18653/v1/2024.naacl-long.137 FREB - TQA : A fine-grained robustness evaluation benchmark for table question answering . In Proceedings of the 2024 Conference of the North American Chapter of...
2024 doi
-
[32]
Yutao Zhu, Huaying Yuan, Shuting Wang, Jiongnan Liu, Wenhan Liu, Chenlong Deng, Zhicheng Dou, and Ji rong Wen. 2023. https://api.semanticscholar.org/CorpusID:260887838 Large language models for information retrieval: A survey . ArXiv, abs/2308.07107
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.