REVIEW 2 major objections 5 minor 46 references
Deep Ancient Roman Republican Coin Classification via Feature Fusion and Attention
T0 review · 2 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Reverse-motif recognition identifies ancient Roman Republican coins with 98.5% top-1 accuracy.
desk verdict The RRCD dataset is the real contribution and the 98.5% main-split number is credible, but the §5.6 generalization claim rests on a permissive proxy that collapses unseen coin types into coarse object buckets. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the fusion of two feature maps, $\alpha_I$ from DenseNet161 and $\beta_I$ from ResNet50, produced by the two networks before their final layers. Instead of a full bilinear outer product, which would explode memory, CoinNet uses compact bilinear pooling: each feature map is passed through a randomized count-sketch projection and the two sketches are convolved, equivalently $\varphi(\alpha,\beta,u,v)=F^{-1}(F(\varphi(\alpha,u,v))\odot F(\varphi(\beta,u,v)))$, so pairwise interactions are captured in a low-dimensional representation. A group of residual blocks then refines the fused vector, an $\ell^2$ normalization stabilizes it, and a learned soft-attention map over the $14\times14$ spatial grid emphasizes informative regions before the final classifier. The paper's ablation shows the accuracy stays near 98.4–98.5% when the backbone pair is replaced, which it interprets as evidence that the fusion, residual, and attention stages, rather than the specific encoder, carry the discriminative information.
What would settle it
Retrain or reuse the described CoinNet on RRCD-Main, then test it on the 128 RRCD-Disjoint classes with the exact motif class per image as the label rather than the grouped main object. If accuracy on biga, quadriga, and curule-chair images drops toward chance (roughly 1/128, or at least far below the reported 68–96%), the generalization claim is refuted.
Extended reading notes
Core claim
CoinNet's central claim is that a compact bilinear fusion of two deep feature maps, followed by residual groups and soft attention, is sufficient to classify eroded and stylistically varied Roman Republican reverse motifs at near-ceiling accuracy. On RRCD-Main (100 classes, roughly 70% training after augmentation), it reaches 98.5% top-1 accuracy, with precision 0.907 and recall 0.951, outperforming VGG (97.4%), NASNet (97.8%), and the spatial bag-of-visual-words baselines (up to 84.4%). The paper also claims strong generalization: using a protocol that treats any training class sharing the main object as a correct match, CoinNet scores 96.56% on 81 biga classes, 68.15% on 35 quadriga classes, and 79.28% on 12 curule-chair classes, all unseen, beating VGG and NASNet by more than 30% in several comparisons.
Load-bearing premise
The system's strongest reported evidence rests on a proxy: a disjoint test coin is deemed correct if it falls into any one of the training classes whose reverse motif contains the same main object, so 'generalization' means recognizing the object, not the unseen class.
Editorial extensions
If this is right
- If 98.5% reverse-motif accuracy holds, a practical coin-attribution tool can narrow a query to the small subset of catalog classes showing that motif, leaving legend and symbol reading for final disambiguation.
- The disjoint-set results imply that models trained on well-preserved museum images can still recognize motif styles from auction photos, provided the evaluation groups classes by shared object.
- The insensitivity to backbone choice means the method can be reimplemented with lighter encoders without expected loss, which matters for deployment on modest hardware.
- The released 18,285-image, 228-class RRCD gives later work a fixed benchmark for coarse-grained ancient coin recognition.
Reading between the lines
- A stricter test—assigning each RRCD-Disjoint image its actual unseen motif class and evaluating 128-way classification—would probably reduce the reported margins, because the paper's protocol counts any class containing the same main object as correct; this is a natural next experiment.
- Because motif and legend are complementary cues and the paper notes legends are more worn, combining reverse-motif recognition with a separate legend reader could yield finer-grained coin attribution than either cue alone.
- The architecture's invariance to encoder choice suggests the fusion and attention stages are where the discriminative signal lives; an ablation replacing compact bilinear pooling with average-pooled concatenation would make that dependency explicit.
- The protocol itself, where a prediction is correct if it shares a semantic object with the true coin, could be reused as a benchmark for object-level attribute recognition rather than exact coin-type classification.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RRCD, a dataset of 18,285 images of ancient Roman Republican coin reverses spanning 228 motif classes (100 in RRCD-Main and 128 in RRCD-Disjoint), and proposes CoinNet, a network that fuses DenseNet161 and ResNet50 feature maps via compact bilinear pooling, applies residual groups and spatial attention, and classifies reverse motifs. On a 30%-70% train-test split of RRCD-Main, CoinNet reports 98.5% top-1 accuracy, outperforming fine-tuned VGG (97.4%), NASNet (97.8%), and earlier BoVWs methods. The paper also claims strong generalization on RRCD-Disjoint, where test images share only a coarse object label (biga, quadriga, curule chair) with the training set, and reports large margins over baselines under a proxy that counts a prediction as correct if it falls into any training class with the same coarse object.
Significance. The RRCD dataset is a potentially valuable resource for cultural-heritage and fine-grained recognition research, and the paper's main accuracy result on RRCD-Main is well supported by the reported experiments. The use of compact bilinear pooling with multi-network feature fusion is a sensible technical direction, and the ablation study in Table 3 shows robustness across feature extractors. The paper also makes model and dataset availability claims, which supports reproducibility. However, the generalization claim is overstated because the disjoint-set evaluation collapses 128 unseen coin classes into three coarse object buckets, so the reported >30% margins do not demonstrate recognition of unseen coin types at the class level. This does not undermine the main accuracy figure, but it does affect one of the paper's stated contributions.
major comments (2)
- [Section 5.6, Table 5] The disjoint-test evaluation counts a test image as correctly classified if the model outputs any training class sharing the same coarse object (biga, quadriga, or curule chair). Because the 128 RRCD-Disjoint classes are not individually labeled, this proxy collapses all variants of each object into a single semantic bucket and measures object-level agreement, not recognition of unseen coin types. The >30% margins over VGG and NASNet therefore do not support the claimed generalization to 'completely disjoint test sets' at the coin-class level. I recommend either annotating the disjoint images with their Crawford class labels and reporting fine-grained top-1 accuracy, or clearly reframing the claim as object-level generalization.
- [Section 5.2, Table 2] The reported improvements over VGG (1.1%) and NASNet (0.7%) come from a single 30%-70% split with no variance estimates, confidence intervals, or multiple runs. Given the small margins, the claim that CoinNet significantly outperforms these baselines on RRCD-Main is not statistically substantiated. I ask for repeated runs (or at least a per-class breakdown and statistical test) to strengthen the central accuracy claim.
minor comments (5)
- [Section 4.2, Figure 6] The dimension flow in Figure 6 (e.g., '14 x 14 x 16k' and '14 x 14 x 2k') is not explained in the text; please clarify whether compact bilinear pooling is applied per spatial location and what the output dimensionality of the residual group is.
- [Section 5.5, Table 4] The text says vocabulary sizes are 'empirically selected' but does not state which setting from Table 4 is used for the BoVWs and RT results in Table 2; please specify the chosen vocabulary size explicitly.
- [Section 4.2, Equation (6)] The loss function uses both qi(y) and pi(y) in Eq. (6) but then refers to qi(x) in the following sentence; please make the notation consistent.
- [Section 5.7] The Limitations subsection discusses only low-resolution and blur effects; it would be appropriate to also acknowledge the coarse-label proxy used in the disjoint-set evaluation, since that limits the interpretation of the generalization results.
- [Table 2] The precision and recall figures are reported without specifying whether they are macro-averaged or micro-averaged; please clarify.
Circularity Check
No circularity: the central accuracy claim is an empirical benchmark result, and the disjoint-set protocol is a coarse proxy rather than a circular derivation.
full rationale
The paper's central claim is an empirical classification result: CoinNet obtains 98.5% top-1 accuracy on a held-out split of RRCD-Main (Table 2), with training and testing performed under a standard cross-entropy loss on fixed architecture components (compact bilinear pooling, residual groups, attention). No parameter is fitted to the reported test labels, and the accuracy does not reduce by construction to an input definition. The only self-citations, Anwar et al. [2,3], are used as baselines for comparison, not as load-bearing premises for CoinNet's design or expected performance. The BoVWs vocabulary sizes in Table 4 are empirically selected for the classical baselines only and do not enter the proposed network, so they do not define the target result. The Section 5.6 disjoint-set evaluation uses an explicitly stated workaround in which a disjoint test image of a biga is counted correct if the model outputs any of the 12 training classes containing a biga; this is a coarse evaluation proxy that limits the strength of the generalization claim, but it is not a circular derivation, because the model's outputs are still produced by the trained network and the proxy is disclosed rather than hidden. The reported generalization margins may be overstated in granularity, but that is a correctness/validity concern, not a self-referential reduction. Therefore no pattern of circularity, fitted-input-called-prediction, or self-citation-load-bearing reasoning is present.
Assumptions & free parameters
free parameters (5)
- Initial learning rate =
1e-2
- Weight decay =
1e-4
- Training split ratio =
30% train, 70% test
- BoVWs vocabulary size =
1k, 5k, 10k, 15k; best at 5k for BoVWs and 1k for RT
- Residual blocks per group =
4
assumptions (4)
- domain assumption Reverse motifs are the primary discriminative cue for classifying Roman Republican coins.
- domain assumption The class labels in RRCD are correct according to Crawford's reference book.
- domain assumption ImageNet pre-trained weights transfer to ancient coin images.
- standard math Count sketch projection and the convolution theorem hold as described.
Cite this review
Pith. "Pith review of Deep Ancient Roman Republican Coin Classification via Feature Fusion and Attention." pith.science (2026). https://pith.science/paper/KR56ZCOW
@misc{pith2026190809428,
author = {Pith},
title = {Pith review of: Deep Ancient Roman Republican Coin Classification via Feature Fusion and Attention},
year = {2026},
howpublished = {\url{https://pith.science/paper/KR56ZCOW}},
note = {Machine review of arXiv:1908.09428}
}
read the original abstract
We perform the classification of ancient Roman Republican coins via recognizing their reverse motifs where various objects, faces, scenes, animals, and buildings are minted along with legends. Most of these coins are eroded due to their age and varying degrees of preservation, thereby affecting their informative attributes for visual recognition. Changes in the positions of principal symbols on the reverse motifs also cause huge variations among the coin types. Lastly, in-plane orientations, uneven illumination, and a moderate background clutter further make the classification task non-trivial and challenging. To this end, we present a novel network model, CoinNet, that employs compact bilinear pooling, residual groups, and feature attention layers. Furthermore, we gathered the largest and most diverse image dataset of the Roman Republican coins that contains more than 18,000 images belonging to 228 different reverse motifs. On this dataset, our model achieves a classification accuracy of more than \textbf{98\%} and outperforms the conventional bag-of-visual-words based approaches and more recent state-of-the-art deep learning methods. We also provide a detailed ablation study of our network and its generalization capability. Models and Datasets available at https://github.com/saeed-anwar/CoinNet
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
M. H. Crawford, Roman republican coinage, V ol. 1, Cambridge University Press, 1974. 2, 3, 8, 9
work page 1974
-
[2]
A Bag of Visual Words Approach for Symbols-Based Coarse-Grained Ancient Coin Classification
H. Anwar, S. Zambanini, M. Kampel, A bag of visual words approach for symbols-based coarse-grained ancient coin classification, arXiv preprint arXiv:1304.6192 (2013). 2, 6, 8, 22
work page Pith review arXiv 2013
- [3]
-
[4]
Arandjelovic, Automatic attribution of ancient roman imperial coins, in: CVPR, 2010, pp
O. Arandjelovic, Automatic attribution of ancient roman imperial coins, in: CVPR, 2010, pp. 1728–1734. 2, 6, 7, 11
work page 2010
- [5]
-
[6]
J. Kim, V . Pavlovic, Discovering characteristic landmarks on ancient coins using convolutional networks, Journal of Electronic Imaging 26 (2015). 2, 8, 11
work page 2015
- [7]
-
[8]
M. Reisert, O. Ronneberger, H. Burkhardt, An efficient gradient based reg- istration technique for coin recognition, in: Proceedings of the Muscle CIS Coin Competition Workshop, Berlin, Germany, 2006, pp. 19–31. 5
work page 2006
Show all 46 references
-
[9]
B.-Y . Feng, M. Ren, X.-Y . Zhang, C. Y . Suen, Automatic recognition of serial numbers in bank notes, Pattern recognition 47 (8) (2014) 2621–2634. 5
2014
-
[10]
N ¨olle, H
M. N ¨olle, H. Penz, M. Rubik, K. Mayer, I. Holl¨ander, R. Granec, Dagobert-a new coin recognition and sorting system, in: DICTA, 2003, pp. 45–55. 5
2003
-
[11]
Huber-M ¨ork, S
R. Huber-M ¨ork, S. Zambanini, M. Zaharieva, M. Kampel, Identification of ancient coins based on fusion of shape and local features, Mach. Vision and App. 22 (6) (2011) 983–994. 5 29
2011
-
[12]
Zaharieva, M
M. Zaharieva, M. Kampel, S. Zambanini, Image based recognition of ancient coins, in: CAIP, 2007, pp. 547–554. 6
2007
-
[13]
Kavelar, S
A. Kavelar, S. Zambanini, M. Kampel, Word detection applied to images of ancient roman coins, in: International Conference on Virtual Systems and Multimedia, 2012, pp. 577–580. 6, 7, 11
2012
-
[14]
Zambanini, A
S. Zambanini, A. Kavelar, M. Kampel, Classifying ancient coins by local feature matching and pairwise geometric consistency evaluation, in: ICPR, 2014, pp. 3032–3037. 6, 7, 8, 11
2014
-
[15]
Kampel, M
M. Kampel, M. Zaharieva, Recognizing ancient coins based on local fea- tures, in: ISVC, 2008, pp. 11–22. 6, 11
2008
-
[16]
D. G. Lowe, et al., Object recognition from local scale-invariant features., in: ICCV , 1999, pp. 1150–1157. 7
1999
-
[17]
H. Bay, A. Ess, T. Tuytelaars, L. Van Gool, Speeded-up robust features (surf), Comput. Vis. Image Underst. 110 (3) (2008) 346–359. 7
2008
-
[18]
Mikolajczyk, C
K. Mikolajczyk, C. Schmid, A performance evaluation of local descriptors, Trans. Pattern Anal. Mach. Intell. 27 (10) (2005) 1615–1630. 7
2005
-
[19]
Zambanini, M
S. Zambanini, M. Kampel, Coarse-to-fine correspondence search for classi- fying ancient coins, in: ACCV , 2012, pp. 25–36. 7, 11, 15
2012
-
[20]
C. Liu, J. Yuen, A. Torralba, Sift flow: Dense correspondence across scenes and its applications., Trans. Pattern Anal. Mach. Intell. 33 (5) (2011) 978–
2011
-
[21]
Zambanini, M
S. Zambanini, M. Kampel, A local image descriptor robust to illumination changes, in: SCIA, 2013, pp. 11–21. 7
2013
-
[22]
Arandjelovi ´c, Reading ancient coins: automatically identifying denarii using obverse legend seeded retrieval, in: ECCV , 2012, pp
O. Arandjelovi ´c, Reading ancient coins: automatically identifying denarii using obverse legend seeded retrieval, in: ECCV , 2012, pp. 317–330. 7
2012
-
[23]
Kavelar, S
A. Kavelar, S. Zambanini, M. Kampel, Reading the legends of roman repub- lican coins, J. Comput. Cult. Herit. 7 (1) (2014) 5:1–5:20. 7
2014
-
[24]
J. Kim, V . Pavlovic, Improving ancient roman coin recognition with align- ment and spatial encoding, in: ECCV , 2014, pp. 149–164. 7, 11
2014
-
[25]
Csurka, C
G. Csurka, C. Dance, L. Fan, J. Willamowski, C. Bray, Visual categorization with bags of keypoints, in: Workshop on statistical learning in computer vision, ECCV , V ol. 1, 2004, pp. 1–22. 8
2004
-
[26]
Hollstein, Die stadtr ¨omische M¨unzpr¨agung der Jahre 78-50 v
W. Hollstein, Die stadtr ¨omische M¨unzpr¨agung der Jahre 78-50 v. Chr. zwis- chen politischer Aktualit ¨at und Familienthematik: Kommentar und Bibli- ographie, Tuduv-Verlag-Ges., 1993. 9
1993
-
[27]
Woytek, Arma et nummi - forschungen zur r ¨omischen finanzgeschichte und m¨unzpr¨a-gung der jahre 49 bis 42 v
B. Woytek, Arma et nummi - forschungen zur r ¨omischen finanzgeschichte und m¨unzpr¨a-gung der jahre 49 bis 42 v. chr (1993). 9
1993
-
[28]
Zambanini, A
S. Zambanini, A. Kavelar, M. Kampel, Improving ancient roman coin clas- sification by fusing exemplar-based classification and legend recognition, in: ICIAP, 2013, pp. 149–158. 11
2013
-
[29]
Kavelar, S
A. Kavelar, S. Zambanini, M. Kampel, K. V ondrovec, K. Siegl, The ilac- project: Supporting ancient coin classification by means of image analysis, International Archives of the Photogrammetry, Remote Sensing and Spatial 31 Information Sciences - XXIV International CIPA Symposi...
2013
-
[30]
J. B. Tenenbaum, W. T. Freeman, Separating style and content with bilinear models, Neural Computation 12 (6) (2000) 1247–1283. 18
2000
-
[31]
Q. Sun, Y . Liu, T.-S. Chua, B. Schiele, Meta-transfer learning for few-shot learning, in: CVPR, 2019, pp. 403–412. 18
2019
-
[32]
G.-S. Xie, L. Liu, X. Jin, F. Zhu, Z. Zhang, J. Qin, Y . Yao, L. Shao, Atten- tive region embedding network for zero-shot learning, in: CVPR, 2019, pp. 9384–9393. 18
2019
-
[33]
Y . Gao, O. Beijbom, N. Zhang, T. Darrell, Compact bilinear pooling, in: CVPR, 2016, pp. 317–326. 19
2016
-
[34]
Fukui, D
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, M. Rohrbach, Mul- timodal compact bilinear pooling for visual question answering and visual grounding, in: Conference on Empirical Methods in Natural Language Pro- cessing, 2016, pp. 457–468. 19
2016
-
[35]
Charikar, K
M. Charikar, K. Chen, M. Farach-Colton, Finding frequent items in data streams, in: International Colloquium on Automata, Languages, and Pro- gramming, 2002, pp. 693–703. 19
2002
-
[36]
N. Pham, R. Pagh, Fast and scalable polynomial kernels via explicit fea- ture maps, in: International Conference on Knowledge Discovery and Data mining, 2013, pp. 239–247. 19 32
2013
-
[37]
Huang, Z
G. Huang, Z. Liu, L. Van Der Maaten, K. Q. Weinberger, Densely connected convolutional networks, in: CVPR, 2017, pp. 4700–4708. 20, 26
2017
-
[38]
K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recogni- tion, in: CVPR, 2016, pp. 770–778. 20, 26
2016
-
[39]
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large- scale hierarchical image database, in: CVPR, 2009, pp. 248–255. 20, 23
2009
-
[40]
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, Y . Bengio, Show, attend and tell: Neural image caption generation with vi- sual attention, in: ICML, 2015, pp. 2048–2057. 20
2015
-
[41]
Zhang, K
Y . Zhang, K. Li, K. Li, L. Wang, B. Zhong, Y . Fu, Image super-resolution using very deep residual channel attention networks, in: ECCV , 2018, pp. 286–301. 20
2018
-
[42]
Z. Yang, X. He, J. Gao, L. Deng, A. Smola, Stacked attention networks for image question answering, in: CVPR, 2016, pp. 21–29. 20
2016
-
[43]
Bottou, Large-scale machine learning with stochastic gradient descent, in: Proceedings of COMPSTAT’2010, Springer, 2010, pp
L. Bottou, Large-scale machine learning with stochastic gradient descent, in: Proceedings of COMPSTAT’2010, Springer, 2010, pp. 177–186. 22
2010
-
[44]
B. Zoph, V . Vasudevan, J. Shlens, Q. V . Le, Learning transferable architec- tures for scalable image recognition, in: CVPR, 2018, pp. 8697–8710. 23, 24, 25
2018
-
[45]
Simonyan, A
K. Simonyan, A. Zisserman, Very deep convolutional networks for large- scale image recognition, ICLR (2015). 23, 24, 26 33
2015
-
[46]
R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, D. Batra, Grad-cam: Visual explanations from deep networks via gradient-based lo- calization, in: ICCV , 2017, pp. 618–626. 24, 25 34
2017
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.