REVIEW 4 major objections 4 minor 8 references
Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper argues that a model that reads ordinary analog clocks well is only matching training patterns, not learning the time-telling rule.
desk verdict A useful perturbation study showing GPT-4.1 is brittle at analog-clock reading, but the 'memorization, not learning' conclusion needs a human baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The working instrument is a synthetic analog-clock dataset in which the same time is rendered three ways: a normal dial, a distorted dial, and a dial with thin arrow-headed hands. By keeping the clock-reading task fixed and changing only visual surface features, the dataset separates generalization from memorization and lets the authors isolate two error channels: misreading a hand's direction, measured as angular error per hand, and confusing which hand is the hour, minute, or second hand. A second probe uses isolated hour hands pointing at each of the 60 tick marks, with normal versus modified hands, to test whether hand shape impairs direction perception on its own. Together these instruments show that hand-role confusion, not pure direction error, is the dominant cause of the catastrophic drop on arrow-handed clocks, and that even when that confusion is removed the model still fails to transfer the benefit of fine-tuning.
What would settle it
Have human readers tell the time on the same 150 synthetic clock images for all three variants. If human mean absolute error on the deformed and arrow-handed clocks stays near zero while the model's is in the hundreds or thousands of seconds, the memorization conclusion is supported; if humans also make large errors on those variants, the conclusion collapses because the images are not trivially readable. A second check would be to fine-tune on clocks with hand shape and thickness randomized independently of hand role and then test on a never-seen hand style; a model that has learned the abstract rule should transfer, while a pattern-memorizer should not.
Extended reading notes
Core claim
The central discovery is that the MLLM's apparent competence on normal analog clocks does not survive small visual transformations that leave the underlying time information unchanged. In the paper's measurements, GPT-4.1 attains a mean absolute error of 232.48 seconds on the 150 normal clock images, 1380.69 seconds on distorted-dial clocks, and 3726.93 seconds on clocks with thinner, arrow-headed hands. After fine-tuning with additional synthetic samples the corresponding errors are 33.82, 784.41, and 2893.57 seconds, so the intervention helps most where the model was already good and least where it was struggling. Error decomposition shows that direction errors dominate on normal and deformed clocks, while on arrow-handed clocks the model frequently swaps the roles of hour and minute hands. Even after excluding those swaps, fine-tuning reduces error by only 32.5 percent, leading the authors to conclude that the model has memorized training patterns rather than learned to tell time, and that its failures stem from low directional sensitivity plus overfitting to hand appearance.
Load-bearing premise
The interpretation that the model is memorizing rather than learning depends on the unmeasured premise that a person can read the deformed and arrow-handed clocks without difficulty; if those images are perceptually ambiguous even for humans, the error gaps could just reflect stimulus difficulty rather than a failure to abstract the time-telling rule.
Editorial extensions
If this is right
- Fine-tuning on more clock images will not by itself make a model robust to other clock styles; the improvement on normal clocks is about 85 percent error reduction, versus about 43 percent on deformed clocks and about 22 percent on arrow-handed clocks.
- A model can identify a hand's direction correctly in isolation yet fail on a complete clock, so fixing spatial perception alone is insufficient because functional role assignment is a separate failure point.
- Small local changes to a familiar input can cause large accuracy drops even when the underlying task is unchanged, implying that MLLM performance on familiar-looking image types may overstate their competence on visually shifted inputs.
- The fine-tuning bottleneck worsens when multiple sources of interference combine, so collecting larger datasets for every visual variant is not a scalable route to general multimodal reasoning.
Reading between the lines
- If the paper's interpretation is right, a controlled human test on the exact same 150 clock images would separate perceptually ambiguous stimuli from genuine memorization; near-human accuracy on the deformed and arrow-handed clocks would support the memorization claim, while large human errors would undermine it.
- The same controlled-deformation methodology could be applied to other seemingly trivial visual tasks, such as reading analog pressure gauges or interpreting protractors, to test whether MLLMs are brittle to style changes whenever spatial-attribute integration is required.
- A training intervention that randomizes hand shape, thickness, length, and color independently of hand function during fine-tuning would reveal whether the model can learn a function-independent rule; the paper's findings predict such randomization is necessary for generalization.
- A sharper memorization test would use clock styles entirely absent from training data, such as clocks with no tick marks or with reversed handedness; a true rule-learner should transfer, while a pattern-memorizer should fail.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper investigates whether GPT-4.1 has learned a general rule for reading analog clocks or has merely memorized patterns from training images. The authors construct a synthetic clock dataset, measure GPT-4.1's mean absolute error (MAE) on normal clocks (232.48 seconds), deformed "Dalí-style" dials (1380.69 seconds), and clocks with thinner arrow-headed hands (3726.93 seconds), and then fine-tune the model on 300 samples per clock type. After fine-tuning, MAEs improve to 33.82, 784.41, and 2893.57 seconds, respectively. The authors interpret the large performance drop on visually altered clocks as evidence that the model has not learned to tell time but has memorized patterns, and they diagnose two failure modes: confusion between hand functions and reduced directional sensitivity. A single-hand probing experiment and a comparison of confused versus non-confused samples are used to support this interpretation. The manuscript is a short empirical case study with an open dataset and code repository.
Significance. If the central claim survives additional controls, the paper would be a useful and timely case study of MLLM brittleness in a simple spatial reasoning task, showing that apparent competence on synthetic in-distribution images does not imply abstraction and that fine-tuning provides limited transfer to perceptual variants. The strengths are the clear task design, the release of the dataset and code, the before/after fine-tuning comparison, and the decomposition of errors into hand-function confusion versus directional error. However, the paper's headline conclusion is currently underdetermined: the inference from large MAE gaps to 'memorization, not learning' requires a human baseline or perceptual sanity check, and the reported numbers lack uncertainty quantification. The significance of the result therefore depends on the outcome of those controls, rather than being established by the present data.
major comments (4)
- [Section 2] The central conclusion that the MLLM 'has not learned to tell the time but rather memorized patterns' relies on the premise stated in Section 2 that 'A person would be able to tell the time in these clocks with no difficulty.' No human baseline or perceptual sanity check is reported for the deformed and arrow-handed stimuli. Without such a control, the MAE increases (232.48s to 1380.69s and 3726.93s) are equally consistent with the modified images being more ambiguous even for humans, for example because arrowheads obscure hand direction or deformed dials alter tick alignment. This is not a circularity issue, but the memorization-versus-abstraction interpretation is underdetermined. A concrete fix is to measure human MAE on the same image sets, or at least to run a perceptual task confirming that the altered images are readable by humans; if human performance also degrades on these stimuli, the central interpretation must be revised.
- [Section 2] All headline MAE numbers come from a single evaluation set of 150 random times, with no confidence intervals, repeated trials, or per-sample variance reported. Given the large spread within the data (Table 1 shows MAEs of 10088.9s for confused arrow-hand samples versus 882.2s for non-confused samples), the claimed 16x degradation between normal and arrow-handed clocks needs uncertainty quantification. Bootstrap confidence intervals over the 150 samples, or repeated evaluation with different random seeds, are necessary to establish that the gaps are not driven by a small number of outliers.
- [Section 2] The fine-tuning protocol is under-specified. The text states that the model was fine-tuned with 'another set of 300 random samples from each of the datasets' and then evaluated on 'the 150 examples,' but it does not state whether the 150 evaluation times are disjoint from the 300 training times, which base model checkpoint was used for fine-tuning, or what hyperparameters (number of steps, learning rate, LoRA versus full fine-tuning) were used. If the evaluation times overlap with the fine-tuning examples, the post-fine-tuning improvements (33.82s, 784.41s, 2893.57s) could reflect memorization rather than generalization. Please report the exact split and fine-tuning configuration.
- [Section 4 and Table 1] The exclusion of six samples because 'the reason of the error could not be clearly determined' is a post hoc filtering step with no stated criterion or inter-annotator agreement. This can bias the comparison between the 'confused' and 'not confused' groups, and the resulting group sizes (45, 99, and 6 excluded) are small. In addition, the single-hand probe in Section 4 tests only the hour hand on 60 images per hand type and reports MAEs of 6.5 degrees versus 8.1 degrees without confidence intervals; this is insufficient to rule out a directional-perception contribution to the full-clock failure. Please provide the classification protocol, the reasons for exclusion, and uncertainty estimates for the angular error comparisons.
minor comments (4)
- [Section 3] The text 'As shown in Figure 2, this issue accounts for all the errors in the normal clock dataset' appears to reference the wrong figure: Figure 2 is a scatter plot of predicted versus true time, whereas the error-type analysis is presented in Table 1 and Figure 4.
- [Section 4] The second table caption repeats the first table's caption ('Table 2: Analysis of samples with and without clock hand function confusion'); the second table should be captioned to describe the angular errors per hand role.
- [Section 5] There is a typo in 'thar are trivial for humans'; it should read 'that are trivial for humans.'
- [Figure 5] Figure 5 is described as an overview of performance but the caption does not specify what the axes, bars, and any error markers represent; please add a self-contained caption.
Circularity Check
No significant circularity: the paper's claims are empirical interpretations of measured MAE differences, not constructions from fitted parameters or self-citations.
full rationale
The paper's derivation chain is a straightforward empirical experiment: generate synthetic clock images, query GPT-4.1, measure MAE on normal, deformed, and arrow-handed clocks, fine-tune on additional synthetic samples, and re-measure. The claim that the model 'has not learned to tell the time but rather memorized patterns' is an inference from the observed MAE increases (232.48s to 1380.69s to 3726.93s) and from the fine-tuning results, not an equation or fitted quantity that is forced by construction. No parameter is fitted to a subset and then reported as a prediction; the 150 test times are separate from the 300 fine-tuning samples, and the single-hand probe (MAE 6.5 degrees vs 8.1 degrees) is an independent control rather than a renormalized version of the headline result. The references to prior clock-reading work and to rule-learning-versus-memorization studies are external citations used for motivation, and none is load-bearing in a way that reduces the conclusion to a self-citation. The weakest point, noted by the reader, is the unmeasured assertion that a person would read the deformed clocks 'with no difficulty'; this concerns evidential support for the memorization interpretation, not circularity. The paper does not derive its conclusion from its own inputs by definition, and no fitted parameter is renamed as a prediction. Therefore the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Synthetic clocks with deformed dials and arrowhead/thinner hands are valid instances of the same time-telling task that a human reader can solve.
- ad hoc to paper A performance drop under visual perturbation evidences memorization rather than a perceptual generalization failure of the visual encoder.
- domain assumption Fine-tuning with 300 samples per clock type is a fair test of whether fine-tuning can teach generalization across clock styles.
Cite this review
Pith. "Pith review of Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?." pith.science (2026). https://pith.science/paper/B75C34Z6
@misc{pith2026250510862,
author = {Pith},
title = {Pith review of: Have Multimodal Large Language Models (MLLMs) Really Learned to Tell the Time on Analog Clocks?},
year = {2026},
howpublished = {\url{https://pith.science/paper/B75C34Z6}},
note = {Machine review of arXiv:2505.10862}
}
read the original abstract
Multimodal Large Language Models which can answer complex questions on an image struggle to tell the time on analog clocks. This is probably due to the lack of images with clocks at different times in their training set. In this work we explore this issue with one of the latest MLLMs: GPT-4.1 to understand why MLLMs fail to tell the time and whether fine-tuning can solve the problem. The results show how models are making progress in reading the time on analog clocks. But have they really learned to do it, or have they only learned patterns in their training datasets? In this work we put the models to the test with different clocks to illustrate the limitations of MLLMs to abstract and generalize.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
Harnessing the power of llms in practice: A survey on chatgpt and beyond.ACM Trans
Jingfeng Yang, Hongye Jin, Ruixiang Tang, Xiaotian Han, Qizhang Feng, Haoming Jiang, Shaochen Zhong, Bing Yin, and Xia Hu. Harnessing the power of llms in practice: A survey on chatgpt and beyond.ACM Trans. Knowl. Discov. Data, 18(6), April 2024
work page 2024
-
[2]
Moma: Multimodal llm adapter for fast personalized image generation
Kunpeng Song, Yizhe Zhu, Bingchen Liu, Qing Yan, Ahmed Elgammal, and Xiao Yang. Moma: Multimodal llm adapter for fast personalized image generation. InComputer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part XL, page 117–132, Berlin, Heidelberg, 2024. Springer-Verlag
work page 2024
-
[3]
Yan Rong, Shan Yang, Guangzhi Lei, and Li Liu. Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Human-like Audiobook Generation.arXiv e-prints, page arXiv:2504.11002, April 2025
arXiv 2025
-
[4]
A survey on multimodal large language models.National Science Review, 11(12):nwae403, 11 2024
Shukang Yin, Chaoyou Fu, Sirui Zhao, Ke Li, Xing Sun, Tong Xu, and Enhong Chen. A survey on multimodal large language models.National Science Review, 11(12):nwae403, 11 2024
work page 2024
-
[5]
Rohit Saxena, Aryo Pradipta Gema, and Pasquale Minervini. Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs.arXiv e-prints, page arXiv:2502.05092, February 2025
arXiv 2025
-
[6]
Zheng Yuan, Hongyi Yuan, Chengpeng Li, Guanting Dong, Keming Lu, Chuanqi Tan, Chang Zhou, and Jingren Zhou. Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.arXiv e-prints, page arXiv:2308.01825, August 2023
arXiv 2023
-
[7]
Lee, Kangwook Lee, and Dimitris Papailiopoulos
Nayoung Lee, Kartik Sreenivasan, Jason D. Lee, Kangwook Lee, and Dimitris Papailiopoulos. Teaching arithmetic to small transformers. 2024. Publisher Copyright: © 2024 12th International Conference on Learning Representa- tions, ICLR 2024. All rights reserved.; 12th International Conference on Learning Representations, ICLR 2024 ; Conference date: 07-05-20...
work page 2024
-
[8]
Do PhD-level LLMs Truly Grasp Elementary Addition? Probing Rule Learning vs
Yang Yan, Yu Lu, Renjun Xu, and Zhenzhong Lan. Do PhD-level LLMs Truly Grasp Elementary Addition? Probing Rule Learning vs. Memorization in Large Language Models.arXiv e-prints, page arXiv:2504.05262, April 2025. 6
arXiv 2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.