{"id":"535b064e-77c2-4ba6-912c-820e330a7057","arxiv_id":"1908.04088","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The 50-year evolution of IJHCS and CHI shows stable core topics (AI and HCI), strong growth in mobile computing, machine learning, and security, and persistent geographic concentration.","lead":"This paper uses publication and citation data to trace how the journal IJHCS and the CHI conference evolved over 50 years, covering countries, citations, and research topics. It finds a stable core of AI and HCI topics in IJHCS, rapid growth of CHI, and a research landscape dominated by a small group of countries.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Stable 'DNA of IJHCS' claim rests on topic percentages that mix direct matches with super-topic inferences; classifier artifact not ruled out.","rationale":"The reader's CONDITIONAL verdict is appropriate and should stand. This stress-test targets the paper's most celebratory and quotable claim: that IJHCS has retained a stable core of AI and HCI topics for 50 years. The CSO classifier's superTopicOf inference is a concrete mechanism by which that stability could be produced artificially, because high-level labels like 'Artificial Intelligence' are added to any paper whose recognized sub-topic sits beneath them in the ontology. The paper does not report direct versus inferred tags, does not validate the classifier on historical texts, and does not quantify how many old IJHCS records lack abstracts, all of which are needed to rule out the artifact. This is a correctness risk rather than a disagreement with the field's consensus. The geopolitical and citation-flow analyses are more independent of the ontology and are supported by released code and data; they are not undermined by this concern. The proposed test is feasible with the provided artifacts and would either confirm the stability claim or reveal that it needs substantial qualification. Because the issue is testable and addressable, rejection is not warranted; the conditional verdict already signals that such validation is required before the paper is treated as a definitive quantitative history. Agreement with the reader is partial: the reader's stated weakest assumption was MAG metadata completeness, while this concern centers on the topic-classification pipeline, though the two interact because missing abstracts force title-only classification and make inferred-topic inflation more severe.","tokens_in":18028,"tokens_out":6207,"duration_ms":71081,"concrete_test":"Use the released GitHub/Zenodo code to rerun the CSO classifier on the IJHCS corpus with superTopicOf inference disabled, reporting direct-match prevalence for 'Artificial Intelligence' and 'Human-Computer Interaction' for 1969-1988 versus 2009-2018. If direct-match AI prevalence falls substantially in the recent period while total (inferred) prevalence stays flat, the stability conclusion is an artifact. Complement this with a manually labeled random sample (e.g., 100 papers per decade) to compute precision/recall of high-level tags on historical papers, and report missing-abstract rates by decade to check whether title-only classification biases the comparison.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.5.2 describes a classifier that annotates a paper by Levenshtein-matching n-grams to CSO labels and then adds every super-topic via superTopicOf, tidying synonyms via relatedEquivalent. The paper reports a mean of 13.9 topics per paper, so most labels are inferred rather than directly evidenced. Under this pipeline, a 1969 paper on expert systems and a 2018 paper on deep learning both receive 'Artificial Intelligence' as an inferred super-topic. The headline conclusion that 'the DNA of the journal has not changed much' (Section 5, Figures 12-13) is drawn from exactly these high-level percentages: AI is 64.1% in 1969-1988 and 62.5% in 2009-2018, while HCI grows to 77.8%. The paper never separates direct topic matches from super-topic inferences, nor validates the classifier on historical IJHCS text, where missing abstracts (MAG property gaps acknowledged in Section 2.2) make title-only classification more likely. Consequently, the observed stability of the AI/HCI core may be an artifact of CSO's transitive closure rather than an editorial constant. The trend lists also rely on small absolute counts (e.g., IJHCS topics with 5-10 papers in 2018) without uncertainty intervals, but the load-bearing issue is the inferred-topic confound for the DNA claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a quantitative, macro-level history of the International Journal of Human-Computer Studies (IJHCS) and the CHI conference, using Microsoft Academic Graph (MAG) metadata over 1969–2018 (CHI from 1982). It combines three analyses: (i) a scientometric description of publication counts, top institutions, cited/citing venues, and reference-memory patterns; (ii) a geopolitical analysis using author affiliations, country rankings with Spearman's rho, first-author patterns, and a proposed \"knowledge debit\" ratio; and (iii) a research-topic analysis using the Computer Science Ontology (CSO) and the CSO Classifier to tag papers with topics and identify rising trends in 2009–2018. The paper's central descriptive claims are that IJHCS has retained a stable core of AI and HCI topics (\"the DNA of the journal has not changed much\"), that CHI has grown from a small meeting to over 1,200 papers per year, and that both venues show highly concentrated, slowly changing geopolitical structures. The authors make their data and code publicly available, and the paper is framed as a contribution to the IJHCS 50th-anniversary special issue.","tokens_in":18248,"tokens_out":2790,"duration_ms":30106,"significance":"If the empirical claims hold, the paper provides a useful, reproducible descriptive account of five decades of HCI-related research, with particular value in its longitudinal scope and its parallel treatment of a journal and a conference. Its strengths are the public release of datasets and code, the use of a transparent extraction pipeline from MAG, and the explicit research questions. The citation and geopolitical analyses follow established methods in spatial scientometrics and appear largely sound. The main risk is in the topic-analysis leg: the classifier is developed and maintained by the same research group, the 0.94 similarity threshold is carried over from prior work without validation on historical IJHCS text, and the reported paper-level topic counts include inferred super-topics. Because the \"stable DNA\" claim is drawn directly from those high-level topic percentages, the paper's central interpretive claim is not yet fully supported. With additional validation and better uncertainty reporting, the paper would be a valuable historical record for the HCI and science-of-science communities.","major_comments":[{"comment":"The \"DNA of the journal has not changed much\" claim (Section 5, based on Figures 12–13) rests on high-level topic percentages, but the CSO Classifier enriches every directly matched topic with all of its super-topics via superTopicOf, yielding an average of 13.9 topics per paper. The paper never reports what fraction of the AI and HCI tags are direct matches versus transitive super-topic inferences, nor does it validate the classifier on historical IJHCS texts where missing abstracts (acknowledged in Section 2.2) make title-only classification more likely. Without separating direct matches from inferred ones, the near-constant AI and HCI percentages could be an artifact of CSO's transitive closure rather than evidence of editorial continuity. I recommend reporting direct-match and inferred-topic percentages separately and validating the classifier against a manually labeled sample stratified by decade.","section":"Section 2.5.2 and Section 3.3"},{"comment":"The knowledge debit formula is ambiguous: the numerator and denominator are written as \"contributions_citing\" and \"contributions_cited_by,\" but it is not clear whether these count contributions by authors from country c in papers that cite the venue, contributions of papers published in the venue that cite country c, or some other combination. The surrounding text describes an imbalance between citing a venue and being cited by it, so the formal definition should be spelled out with explicit sets. In addition, the paper does not state how zero denominators (countries that are never cited by the venue) are handled beyond being colored black, nor how missing affiliation data in MAG affect the ratio. This metric is used to support the \"knowledge generation is confined to a small number of countries\" claim, so its definition should be precise and its sensitivity to missing data assessed.","section":"Section 2.4, Eq. (1)"},{"comment":"The geopolitical rankings and citation-flow results depend entirely on MAG's affiliation fields and citation lists, but the paper acknowledges only qualitatively that \"some publications in MAG may lack some of these properties.\" No coverage statistics are reported, such as the fraction of IJHCS and CHI papers with at least one affiliation ID per decade, or the fraction of cited references with resolved venue and country information. Since the Spearman rho values (near 0.9 for CHI) and the country rankings are central to the \"closed to newcomers\" claim, the authors should quantify how much missingness varies over time and whether the trend toward higher rho could reflect improved metadata coverage in later years rather than a genuinely more static landscape.","section":"Section 2.2 and Section 3.2"},{"comment":"The rising-topic analysis uses post hoc groupings by 2018 publication counts (e.g., >=60, >=20, >=10, >=5 for CHI), then selects 10 topics per group after \"reviewing the resulting lists with domain experts\" and discarding redundant topics. The criteria and the identities of the experts are not described, and the reported counts for IJHCS are very small (e.g., topics with 5–10 papers in 2018). No uncertainty intervals or significance tests are provided, so the upward trends in Figures 14–16 may not be robust to a few reclassifications or to MAG metadata noise. I recommend reporting the full selection procedure, the completeness of the topic lists, and either confidence intervals from a resampling procedure or at least the raw counts and total paper counts per year.","section":"Section 2.5.3 and Figures 14–16"}],"minor_comments":[{"comment":"The caption contains a typo: \"IIJHCS\" should be \"IJHCS.\"","section":"Figure 1 caption"},{"comment":"The caption states \"Main research topics in IJHCS during the 1969-2018 period,\" but the surrounding text and the figure title indicate the intended period is 1969–1988; please correct the caption.","section":"Figure 12 caption"},{"comment":"The phrase \"Virtual Words\" appears in both the figure and the text; this should read \"Virtual Worlds.\"","section":"Figure 16 and Section 3.3"},{"comment":"The term \"self-sustaining countries\" is used in Figure 8 but not defined in the text; please define it explicitly as countries with at least five papers whose authors are all from the same country.","section":"Section 2.4"},{"comment":"The paper does not state the date or version of the MAG snapshot used. Since MAG is updated over time and the analysis is meant to be reproducible, the exact snapshot and download date should be reported.","section":"Section 2.2"},{"comment":"The sentence \"The CSO Classifier ... was shown to generate excellent results [10, 15]\" cites only the authors' own prior work. Please provide external validation or a clear statement of the classifier's performance on a held-out benchmark relevant to this dataset.","section":"Section 2.5.2"},{"comment":"The reference-memory heatmaps would benefit from a color scale and a description of how the number of citations is normalized (if at all), because the claim that IJHCS has a \"broader attention span\" depends on how the heatmaps are read.","section":"Section 3.1, Figures 5a–5d"},{"comment":"The concluding sentence describing the story of IJHCS as \"amazing vision, sustained excellence, and great success\" is celebratory rather than analytical; consider softening this to match the paper's otherwise descriptive tone.","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The central issue for me is the dependence of the \"stable DNA\" claim on the authors' own CSO and CSO Classifier, with no external benchmark on historical IJHCS text. This is a correctness-risk concern rather than a circularity accusation, but given that the same group maintains the ontology and the classifier, an independent or at least transparent validation step would substantially strengthen the paper. I would also encourage the editor to ensure that the specialization/celebratory framing of Section 5 does not overshadow the need for quantitative uncertainty reporting in the topic trends."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a competent scientometric study with real value. It offers the first combined longitudinal description of IJHCS and CHI from 1969/1982 to 2018, and the authors share code and data openly. The geopolitical and citation-flow analyses work well and are largely independent of the authors' own tools. The knowledge-debit maps and country-ranking stability series are useful additions, and they support the familiar picture of a concentrated, slow-changing geography.\n\nThe soft spot is the topic analysis. The CSO classifier tags papers by string matching, then adds every super-topic from the ontology. The paper reports 13.9 topics per paper, so most labels are inferred, not direct. The headline claim that the journal's DNA has not changed much is drawn from high-level percentages: AI is around 64% in both early and recent periods. Without separating direct and inferred labels, or validating the classifier on this corpus (missing abstracts in older MAG records make title-only classification likely), that stability could be an artifact of the taxonomy's transitive closure rather than a property of the journal. The trending topics also depend on post hoc grouping and expert selection, with small absolute counts and no uncertainty intervals. The authors disclose the data gaps in MAG but do not quantify how missing affiliations or citations affect their rankings.\n\nThe stress-test concern holds up. It does not sink the paper, because the geopolitical and citation claims do not rely on the classifier, and those are probably the most durable parts. But the 'remarkable vision' conclusion in Section 5 is overstated given the tool-chain issue.\n\nIf this lands on my desk, I would send it to peer review: the dataset and code are reproducible, the method is transparent, and the weaknesses are addressable with sensitivity analysis or a careful discussion. A reviewer should ask for that. I would cite this for the geographic and citation results, not for the topic stability claim.\n\nRecommendation: engage with it, with a required revision on the topic analysis.","headline":"Useful descriptive scientometrics with a solid geopolitical core, but the 'stable DNA' claim leans on an unvalidated topic classifier that infers broad super-topic labels.","tokens_in":18832,"tokens_out":2746,"would_cite":true,"duration_ms":29189,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A bibliometric study of 50 years of IJHCS and CHI shows a stable journal core, a much larger conference, and a nearly frozen country ranking.","keywords":["Science of Science","Scientometrics","Spatial Scientometrics","Bibliographic Data","Scholarly Data","Human-Computer Interaction","IJHCS","CHI"],"falsifier":"Recompute the country rankings and knowledge-debit values on a manually verified subset of, say, five thousand IJHCS and CHI papers whose affiliations and references are checked against publisher records; if the Spearman rho values for country rankings drop below 0.7 or the top-10 country lists change substantially, the claim of a static, concentrated geopolitical structure would fail.","tokens_in":17773,"feed_emoji":"📊","tokens_out":5436,"duration_ms":57286,"temperature":0.7,"pith_summary":"This paper attempts to establish a data-driven account of 50 years of human-computer interaction research by tracing everything published in IJHCS (including its predecessor, the International Journal of Man-Machine Studies) and in the CHI conference, plus the papers they cite and that cite them. The authors claim that IJHCS has preserved a stable core of artificial intelligence and HCI topics across five decades, which they call the journal's DNA, while the topics it draws on and speaks to have shifted around that core. They also claim that CHI has grown from a small specialist meeting to more than 1,200 papers a year, that IJHCS cites a broader set of venues over a longer time span than CHI does, and that the country rankings for both venues are highly concentrated and nearly static, with CHI's Spearman rank correlation approaching 0.9. If these claims hold, the field's history looks less like replacement and more like accumulation around a stable core, with access concentrated among a small set of countries.","feed_headline":"IJHCS kept its AI and HCI core for 50 years","feed_subtitle":"A quantitative study of IJHCS and CHI finds stable topics, a fast-growing conference, and a nearly frozen country ranking.","key_machinery":"The analysis rests on three instruments. One is a large, openly licensed scholarly metadata graph that supplies the publication sets, author affiliations, and citation links; completeness of that graph is the load-bearing assumption behind every computed ranking and ratio. Another is a large taxonomy of computer science research topics together with an automated classifier that tags each paper by matching n-grams from titles and abstracts to topic labels and then lifts broader super-topics such as HCI and artificial intelligence. The third is a set of quantitative measures: Spearman's rank correlation for country-ranking stability, contribution counts for geopolitical presence, and the knowledge-debit ratio, defined as the number of citing contributions a country makes toward a venue divided by the number of cited-by contributions it receives. Together these instruments convert raw metadata into the paper's descriptive claims.","core_discovery":"The central claim, stated on the paper's own terms, is that IJHCS's intellectual identity has remained remarkably stable: artificial intelligence, HCI, and knowledge-based systems were core topics in the 1969-1988 period and remain core today, with HCI and user interfaces strengthening over the last decade. CHI, in contrast, expanded dramatically in output but became more self-referential and more concentrated in its referencing, with CHI authors mostly citing recent CHI, UIST, and CSCW papers and with the yearly correlation of country rankings climbing toward 0.9. The paper also introduces a knowledge-debit measure: when a country's papers cite a venue more often than the venue's papers cite that country, the country accumulates a knowledge debit toward the venue. The results show that countries in Asia, South America, and the Middle East carry large debits toward both IJHCS and CHI, indicating participation in the conversation without reciprocal influence.","pith_inferences":["The stable-DNA claim could be tested more strictly by applying topic models directly to full texts; the paper's classifier relies on n-gram similarity to a fixed taxonomy, which may miss conceptual continuity expressed in changing vocabulary.","The journal-versus-conference contrast in citation span may be a general pattern rather than an IJHCS/CHI quirk; comparing similar journal-conference pairs in other fields would show whether long reference memories are intrinsic to journals or particular to this community.","The knowledge-debit measure could be used as a policy instrument: recomputing it after targeted calls, mentorship programs, or selection changes would yield a quantitative test of whether access interventions alter reciprocal citation flows."],"forward_implications":["If IJHCS's core is stable, the journal's future is plausibly continuous with its past: AI and HCI remain anchors, and new topics are absorbed around them rather than replacing them.","If CHI's country rankings really hover near 0.9, then efforts to increase geographic diversity face a structural headwind, not simply a pipeline problem.","If knowledge-debit imbalances are real, then institutions in high-debit countries are more likely to appear as citing partners than as first authors, suggesting a specific form of unequal participation.","If CHI's referencing is short-memory and concentrated in a few venues, then citation-based assessments of the field will systematically favor recent work and underrepresent the older literature that IJHCS-style venues still engage with."],"supporting_citations":[{"why":"Frames the Science of Science approach that motivates the study's research questions.","marker":"[1]"},{"why":"Supplies the scholarly metadata graph from which the paper extracts all publications, affiliations, and citations.","marker":"[2]"},{"why":"Provides the geopolitical framework and country-ranking method that the paper adapts for IJHCS and CHI.","marker":"[3]"},{"why":"Supplies the large topic taxonomy used to classify the papers into research areas.","marker":"[4]"},{"why":"Defines Spearman's rank correlation, the measure used to quantify country-ranking stability.","marker":"[8]"},{"why":"Describes the automatic mapping-study methodology the paper follows for topic trend analysis.","marker":"[10]"},{"why":"Describes the specific classifier version used to tag IJHCS and CHI papers with research topics.","marker":"[16]"}],"fun_headline_variants":["50 years of HCI: IJHCS stable, CHI self-inflates","CHI grew huge but cites itself: 50-year citation analysis","Knowledge debit: Asia and South America fuel HCI citations","Geography of HCI: country rankings barely moved in 50 years","IJHCS stayed true to AI and HCI for five decades"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central premise is that the scholarly metadata graph's records of which papers belong to each venue, which authors and countries are listed, and which citation links exist are complete and accurate enough over five decades that the rankings and ratios reflect the real field rather than metadata gaps.","fun_headline_variants_meta":{"raw":{"variants":["50 years of HCI: IJHCS stable, CHI self-inflates","CHI grew huge but cites itself: 50-year citation analysis","Knowledge debit: Asia and South America fuel HCI citations","Geography of HCI: country rankings barely moved in 50 years","IJHCS stayed true to AI and HCI for five decades"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001116,"raw_usage":{"total_tokens":4620,"prompt_tokens":892,"completion_tokens":3728,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":3635}},"tokens_in":508,"tokens_out":3728,"duration_ms":25416,"temperature":1.0,"reasoning_tokens":3635,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:51:13.321307+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute the country rankings and knowledge-debit values on a manually verified subset of, say, five thousand IJHCS and CHI papers whose affiliations and references are checked against publisher records; if the Spearman rho values for country rankings drop below 0.7 or the top-10 country lists change substantially, the claim of a static, concentrated geopolitical structure would fail.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames the Science of Science approach that motivates the study's research questions."},{"cited_title":"(Paul), Wang, K.: An Overview of Microsoft Academic Service (MAS) and Applications","cited_arxiv_id":null,"evidence_quote":"Supplies the scholarly metadata graph from which the paper extracts all publications, affiliations, and citations."},{"cited_title":"Data Science","cited_arxiv_id":null,"evidence_quote":"Provides the geopolitical framework and country-ranking method that the paper adapts for IJHCS and CHI."},{"cited_title":"In: International Semantic Web Conference 2018","cited_arxiv_id":null,"evidence_quote":"Supplies the large topic taxonomy used to classify the papers into research areas."},{"cited_title":"The American journal of psychology","cited_arxiv_id":null,"evidence_quote":"Defines Spearman's rank correlation, the measure used to quantify country-ranking stability."},{"cited_title":"Data Science","cited_arxiv_id":null,"evidence_quote":"Describes the automatic mapping-study methodology the paper follows for topic trend analysis."},{"cited_title":"In: In International Semantic Web Conference (P&D/Industry/BlueSky)","cited_arxiv_id":null,"evidence_quote":"Describes the specific classifier version used to tag IJHCS and CHI papers with research topics."}],"review_version":1}