{"id":"e86e6828-57d5-4cd4-ad99-040ff93738f3","arxiv_id":"1908.07465","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Visual fingerprints of figures in arXiv papers yield inter-field distances that correlate with citation and text distances, and neural-network diagram counts rose before citation counts for key deep learning papers.","lead":"This paper tests whether the pictures and diagrams in scientific papers can map how research fields relate to each other, just like text and citations do. It finds that visual patterns in figures separate academic disciplines about as well as citations or jargon, and may reveal new trends earlier than citation counts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline correlation r=0.706 is selected by maximizing over PCA dimension and cluster count k on the evaluation benchmark itself; without cross-validation, the claim that figures are as effective as text or citations is not established.","rationale":"The reader's verdict of CONDITIONAL is appropriate. The reader's weakest_assumption identified the same load-bearing concern: free parameters (PCA dimension and cluster count k=4) are chosen to maximize agreement with citation distance, with no cross-validation or out-of-sample check. My analysis confirms this is the central threat to the headline claim. The paper's own Table 2 shows that the average correlation over k is substantially lower than the reported maximum, indicating that the reported r=0.706 is an optimistic selection. A nested cross-validation test would provide an honest estimate. I do not see a need to change the verdict; the concern is fixable with additional analysis, and the paper's core idea remains plausible. I focus on parameter overfitting rather than the ResNet transfer assumption because the transfer concern is partly mitigated by the interpretable clusters found in Section 5.3; the lack of cross-validation directly undermines the quantitative comparison that supports the main claim.","tokens_in":13887,"tokens_out":5970,"duration_ms":62713,"concrete_test":"Re-run the Mantel test with nested cross-validation: split the set of arXiv disciplines into training and held-out halves, select PCA dimension and k on the training half only (maximizing Mantel r against citation distance), then compute the Mantel correlation on the held-out half. Repeat across random splits and report the mean and 95% CI. If the held-out r falls below the jargon-citation correlation (0.697) or the Table 2 average (~0.66), the headline claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative evidence is the Mantel correlation between visual distance and citation distance (r=0.706, Table 3). Section 5.1 describes selecting both hyperparameters (PCA dimension and number of clusters k) to maximize correlation with citation distance: 'higher values (200+) produced stronger correlations' and 'Maximum correlation occurred at k = 4 in most experiments.' The final pipeline uses k=4 and a high dimension, so the reported r is a maximum over a grid of parameter values evaluated on the same benchmark used for evaluation. Table 2 shows that average correlations across k=2..30 are only 0.66-0.67 for dimensions 256/320, while the maximum reaches 0.793. Since the citation-distance benchmark is used both to tune and to evaluate, the comparison against jargon-citation (r=0.697) is biased in favor of visual distance. No cross-validation, bootstrap, or nested selection is reported. If the true out-of-sample correlation is closer to the 0.66 average, visual distance would no longer be 'as effective' as the text baseline, and the headline claim would fail. This is the most load-bearing weakness because the main conclusion rests on this single correlation value.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a method for delineating scientific knowledge domains from the visual content of figures in the arXiv literature. Figures are embedded with a pretrained ResNet-50, projected with PCA, and clustered with K-means; each arXiv category is then represented by a histogram of cluster memberships, and pairwise Euclidean distances between these histograms are compared with citation-based and jargon-based distance matrices using Mantel tests. The main quantitative result is a correlation of r=0.706 between visual and citation distance, which the authors interpret as evidence that figures are as effective as text or citations for distinguishing communities of practice. The paper also analyzes disagreement between visual and citation distance (e.g., Computation and Language), and presents a case study in which the prevalence of neural-network diagrams grows before the citation counts of influential deep-learning papers, suggesting that visual signals may be leading indicators of idea adoption.","tokens_in":14128,"tokens_out":4496,"duration_ms":47415,"significance":"If the core claim held, this would be a meaningful contribution to the science-of-science literature by adding a third measurable communication channel, with potential applications in classification, recommendation, and studies of scientific information flow. The paper's strengths include the scale of the corpus (over 5 million figures), a transparent and simple pipeline, and the explicit comparison against established textual and citation baselines. However, the central quantitative claim is currently weakened by a circular evaluation loop: the free parameters (PCA dimension and number of clusters) are selected by maximizing the correlation with the very same citation-distance benchmark that is later reported as validation. The average correlation across the parameter grid is 0.66-0.67, which falls below the jargon-citation correlation of 0.697, so the headline conclusion is fragile without out-of-sample evaluation.","major_comments":[{"comment":"The reported headline correlation r=0.706 between visual and citation distance is not an honest out-of-sample estimate. Section 5.1 and Table 2 describe selecting both the PCA dimension and the number of clusters k by maximizing the Mantel correlation with citation distance (e.g., 'Maximum correlation occurred at k = 4'), and the same citation-distance matrix is then used as the evaluation benchmark in Table 3. This creates a circular loop in which the benchmark is used both for tuning and for validation. The average correlations across k=2..30 in Table 2 are 0.66-0.67 for the higher dimensions, which is actually below the jargon-citation correlation of 0.697 reported in Table 3. Without cross-validation, a held-out benchmark, or a nested selection procedure, the claim that visual distance is 'as effective' as text or citations is not established.","section":"Section 5.1 and Table 3"},{"comment":"The H3 leading-indicator claim is supported only by a visual comparison of unnormalized time series. The top chart shows the raw count of papers containing neural-network diagrams in four arXiv categories, while the bottom chart shows the citation counts of six selected influential papers. These are not comparable quantities: the figure count is not normalized by the total number of papers published in those categories per year, and the citation counts belong to only a handful of papers rather than to the field as a whole. The observed increase in figure usage could therefore reflect the overall growth of deep learning submissions or of arXiv itself. No formal lead-lag test, such as Granger causality or cross-correlation with a proper normalization, is performed. The conclusion in Section 6 that 'figures in the scientific literature are leading indicators of citations' is consequently not supported by the evidence presented.","section":"Section 5.5 and Figure 9"},{"comment":"The method depends on the assumption that ImageNet-pretrained ResNet-50 features transfer to scientific figures, an assumption the paper states but does not quantitatively validate. Section 3.2.1 argues that early layers identify 'simple patterns that are sufficiently general' for artificial images as well, but this remains an assertion. The later cluster analysis (Section 5.3) shows that clusters are dominated by broad figure types (diagram, table, plot, photo), which provides indirect, post hoc support. However, because the entire visual-signature construction rests on this transfer, the paper should offer a direct validation, for example by measuring cluster agreement against human labels on a sample of scientific figures, or by showing that the figure-type classifier performance on scientific figures is acceptable. Without such a check, the risk remains that the visual signatures capture artifacts of natural-image features rather than meaningful scientific visual communication.","section":"Section 3.2.1"}],"minor_comments":[{"comment":"There is an unexplained discrepancy between Table 2 and Table 3: Table 2 reports maximum correlations of 0.793 at dimension 256/320 with k=4, while Table 3 reports r=0.706 for the visual-citation comparison. The authors should explain how the final pipeline differs from the parameter-search experiments, or why the reported result is lower than the maximum.","section":"Section 5.1 vs Table 3"},{"comment":"The Mantel test p-values are reported as 0.0001 and 0.0002, but the number of permutations used to obtain these p-values is not stated. Please specify the permutation procedure and the number of randomizations.","section":"Table 3"},{"comment":"The precision and recall values (0.741 and 0.827) are reported only for neural-network diagrams; please provide the per-class precision, recall, and F1 for the embedding-visualization class as well, since both classes are used in the subsequent time-series analysis.","section":"Section 5.5"},{"comment":"The three line charts in Figure 9 have different y-axis scales and are not normalized by field size, making the visual comparison of trends potentially misleading. At minimum, the authors should discuss this limitation; ideally, they should normalize the figure count and the citation counts by the number of papers in each category per year.","section":"Figure 9"},{"comment":"The manuscript contains multiple OCR artifacts and typos (e.g., 'bibilometrics' in the introduction, 'Sea/t_tle' in the affiliation block, 'pa/t_terns' in several places), and the ACM template still carries placeholder metadata ('Conference’17', 'DOI: 475/123'). The authors should proofread the PDF and ensure the final version is clean.","section":"Throughout"},{"comment":"Table 1 lists only the final learning rate, decay, epochs, batch size, and loss, but the text says the authors tuned depth, layer dimensions, and dropout. Please report the search space or at least the range of values considered for each hyperparameter.","section":"Table 1 and Section 3.3"},{"comment":"It is unclear whether the full 5-million-figure corpus or only the 1.5M-figure sample from the parameter experiment is used for the final clustering and signature construction. Please clarify the exact data flow used to produce the results in Tables 2 and 3.","section":"Sections 3.2.2-3.2.4"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a novel and promising direction, and the authors have done a substantial empirical investigation. My main concern is the circular evaluation: tuning hyperparameters on the benchmark and then reporting the tuned correlation as validation. This is fixable with a nested or held-out evaluation, but given that the gap between r=0.706 and r=0.697 is very small, an honest out-of-sample estimate might well overturn the paper's central claim. I would also ask the authors to tighten the H3 analysis, which currently over-interprets unnormalized time series. The manuscript appears to be a semi-final draft of a conference paper; the authors should update the metadata and formatting before resubmission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is the first paper I know of that tries to delineate scientific fields from the visual content of figures at scale, and that alone makes it worth a look. The authors build visual signatures by embedding figures with ResNet-50, PCA, k-means, and then compare pairwise field distances to citation and text distances. They report r=0.706 with citation distance, slightly beating the text-jargon distance r=0.697, and conclude figures are as effective as text or citations for separating communities.\n\nWhat is genuinely good: the pipeline is simple and reproducible in principle; the cluster analysis shows each of the four k-means clusters maps to a dominant figure type (diagrams, tables, plots, photos), which is a nice sanity check; and the case study on Computation and Language—where tables dominate and photos are rare—is a concrete example of how visual signals capture something text and citations miss. The idea that visual style is a distinct channel for studying idea flow is plausible and worth pursuing.\n\nThe soft spots are serious, though. The headline r=0.706 is selected, not measured: Section 5.1 explicitly picks PCA dimension and k to maximize correlation with citation distance, and Table 2 shows the average correlation across k=2..30 is around 0.66, with the maximum at k=4 reaching 0.79. Using the same benchmark for tuning and evaluation inflates the reported number, and the comparison with jargon distance (which was not tuned) is unfair. Without cross-validation or an out-of-sample parameter selection, we cannot trust the 'as effective as text' claim.\n\nThe leading-indicator section (H3) is also underpowered: raw counts of neural-network diagrams in a handful of hand-picked CS subfields, compared to citation counts of six famous papers, is not a rigorous lead-lag analysis. The two-year gap they observe may be real, but the current evidence doesn't establish it.\n\nNo code or data are released, which makes verification harder. That said, the authors are transparent about their parameter choices; the problem is methodological, not hidden.\n\nFor whom? Bibliometrics and science-of-science researchers will want to know about this as a first attempt at a visual modality. I would send it to a serious referee, but with the expectation of major revisions: add cross-validation, report a fair comparison, and rework the lead-lag analysis. As is, I'd be hesitant to cite the headline numbers.\n\nRecommendation: engage, but treat the central correlation as an upper bound until it survives out-of-sample testing.","headline":"First large-scale attempt to delineate scientific fields from figures, but the headline correlation is fitted to the benchmark and the leading-indicator claim is speculative.","tokens_in":14674,"tokens_out":2640,"would_cite":true,"duration_ms":24811,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that scientific figures can differentiate communities of practice as effectively as text or citation patterns, establishing a third measurable channel for mapping knowledge domains.","keywords":["VizioMetrics","science of science","bibliometrics","scientometrics","visual signatures","knowledge domains","figure-type classification","concept adoption"],"falsifier":"Re-run the pipeline on a held-out slice of the figure corpus (fields or years not used in tuning), fix the PCA dimension and cluster count before looking at citation distances, and compute the visual–citation correlation; if it collapses toward zero while the text and citation measures stay stable, the reported r = 0.706 reflects overfitting rather than a stable visual signal.","tokens_in":13674,"feed_emoji":"🖼️","tokens_out":10365,"duration_ms":96463,"temperature":0.7,"pith_summary":"The paper argues that the images inside scientific papers carry a structured signal about which research community produced them, alongside the usual signals of text and citations. It builds a \"visual signature\" for each discipline from the distribution of figure types, then shows that distances between these signatures match citation distances (r = 0.706) and jargon distances (r = 0.531) across the major fields in the preprint corpus. Because visual distance correlates with, but does not repeat, text and citation structure, the authors claim figures can serve as a third channel for mapping knowledge domains and for detecting when a community actually adopts a method. If right, this would let classification, recommendation, and science-mapping tools make use of pictures rather than only words and links.","feed_headline":"Figures rival text and citations for mapping science","feed_subtitle":"Visual signatures from 5 million scientific figures separate fields and flag idea adoption before citations do.","key_machinery":"The central object is the visual signature: a normalized histogram over k = 4 clusters formed from PCA-reduced feature vectors of all figures in a field, where the features come from a pre-trained ResNet-50. The four clusters turn out to be dominated by diagrams, tables, plots, and photos, so the signature summarizes a discipline's visual output as a small probability distribution over figure types. Euclidean distance between signatures supplies the visual distance, and the Mantel test [32], a standard test of correlation between distance matrices, compares the resulting distance matrix against citation and jargon distance matrices. The same feature-extractor and cluster machinery, plus a classifier for two specific figure types, lets the authors track how neural-network diagrams and embedding visualizations propagate through computer science over time.","core_discovery":"The paper's central claim is that a field's visual style is a reliable marker of its intellectual territory, comparable in power to its jargon or its citation links. The authors encode 5,009,523 figures from 1,343,669 preprints with a pre-trained ResNet-50 [18], reduce the resulting vectors with PCA, cluster the corpus into four visual types, and form a normalized histogram per discipline. Euclidean distance between these \"visual signatures\" correlates with average shortest citation path at r = 0.706 (p = 0.0001) and with jargon distance at r = 0.531 (p = 0.0002); the visual–citation correlation is slightly higher than the jargon–citation correlation (r = 0.697). The disagreements are interpretable: Computation and Language is citation-close to other computer-science fields but visually distinct because of its table-heavy figures. The case study of neural network and embedding diagrams shows figure usage rising sharply around 2015, about two years before citation counts to influential deep-learning papers surge, which the authors read as evidence that figures are leading indicators of genuine concept adoption.","pith_inferences":["Because producing a method-specific figure is costlier than adding a keyword or citation, visual signals may be a more conservative measure of true adoption; the paper's observed lead over citations is consistent with that interpretation, though the paper does not directly test the cost mechanism.","The fact that k = 4 clusters sufficed suggests the scientific literature may draw on only a handful of visual genres, which would make visual signatures a cheap feature for large-scale bibliometric pipelines.","A direct extension would be to plug visual signatures into a recommendation system and compare recommendation quality against text-only and citation-only recommenders; the paper motivates this but does not build it.","The same classifier approach could map the spread of other visual conventions, such as phylogenetic trees, loss-curve panels, or t-SNE plots, to reveal how visualization practices themselves diffuse between communities."],"forward_implications":["Visual distance can be used as a standalone similarity measure for clustering and delineating scientific fields, with accuracy comparable to citation- and text-based measures.","Where visual distance and citation distance disagree, the gap identifies fields with distinct visual cultures, such as Computation and Language's reliance on comparison tables, giving science mapping a new axis of analysis.","Figure-type classifiers can track adoption of specific methods, and the presence of a method-specific figure is a better marker of genuine integration than merely mentioning the method or citing a key paper.","Figure usage leads citation counts by roughly two years in the deep-learning case, so visual signals could serve as early indicators of where a field is heading.","The method transfers to other visual corpora such as patents, web pages, and news, where similar visual-domain structure may exist."],"supporting_citations":[{"why":"Supplies the jargon-distance metric and the cultural-holes approach that the paper compares visual distance against and extends.","marker":"[43]"},{"why":"Supplies the average-shortest-path algorithm used to compute citation distance between fields.","marker":"[10]"},{"why":"Supplies the Mantel test, the matrix-correlation test used for all distance-matrix comparisons.","marker":"[32]"},{"why":"Supplies ResNet-50, the pretrained network that turns each figure into a 2048-dimensional feature vector.","marker":"[18]"},{"why":"Supplies the ImageNet corpus and pretraining task that the feature extractor depends on.","marker":"[8]"},{"why":"Establishes the prior Viziometrics link between visual information and scientific impact and provides the five-category figure taxonomy.","marker":"[28]"},{"why":"Provides the CNN classifier used to label cluster members as diagram, plot, table, photo, or equation when interpreting clusters.","marker":"[29]"},{"why":"Supplies the taxonomy used to label neural network diagrams for training the figure-type classifier.","marker":"[38]"},{"why":"Serves as the reference AlexNet figure and as one of the influential papers in the adoption-timing case study.","marker":"[23]"}],"fun_headline_variants":["Visual style maps science as well as citation networks","Figures flag new ideas before citations catch on","Visual signatures rival text for classifying fields","Picture-based map of science matches citation geography","Scientific images trace knowledge domains across fields"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that image features trained on natural photographs group scientific figures into meaningful visual types, and that the PCA dimension and cluster count, which were chosen to maximize agreement with the citation benchmark, are not overfitting that benchmark; if either fails, the visual signatures are noise and the reported correlations could be spurious.","fun_headline_variants_meta":{"raw":{"variants":["Visual style maps science as well as citation networks","Figures flag new ideas before citations catch on","Visual signatures rival text for classifying fields","Picture-based map of science matches citation geography","Scientific images trace knowledge domains across fields"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1279,"prompt_tokens":936,"completion_tokens":343,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":552,"completion_tokens_details":{"reasoning_tokens":278}},"tokens_in":552,"tokens_out":343,"duration_ms":4390,"temperature":1.0,"reasoning_tokens":278,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:46:06.157075+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the pipeline on a held-out slice of the figure corpus (fields or years not used in tuning), fix the PCA dimension and cluster count before looking at citation distances, and compute the visual–citation correlation; if it collapses toward zero while the text and citation measures stay stable, the reported r = 0.706 reflects overfitting rather than a stable visual signal.","supporting_citations":[{"cited_title":"Vilhena, J","cited_arxiv_id":null,"evidence_quote":"Supplies the jargon-distance metric and the cultural-holes approach that the paper compares visual distance against and extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the average-shortest-path algorithm used to compute citation distance between fields."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Mantel test, the matrix-correlation test used for all distance-matrix comparisons."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the ImageNet corpus and pretraining task that the feature extractor depends on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the prior Viziometrics link between visual information and scientific impact and provides the five-category figure taxonomy."},{"cited_title":"Sean Yang, Jevin West, and Bill Howe","cited_arxiv_id":null,"evidence_quote":"Provides the CNN classifier used to label cluster members as diagram, plot, table, photo, or equation when interpreting clusters."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the taxonomy used to label neural network diagrams for training the figure-type classifier."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Serves as the reference AlexNet figure and as one of the influential papers in the adoption-timing case study."}],"review_version":1}