{"id":"512b7c84-4a5c-488f-a71c-ced9ad8f7e32","arxiv_id":"1908.01874","paper_version":3,"verdict":"UNVERDICTED","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A proposal for an interactive graph that connects machine-learning methods by which components they inherit from earlier methods.","lead":"This white paper describes Backronym, a proposed website that maps which machine-learning methods build on which earlier methods. The authors argue that such an inheritance graph could help researchers find related work and ideas more quickly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The core utility claim is unmeasured: the graph's value depends on hand-annotated 'Based on' edges that the authors concede are imperfect, and no evaluation of navigation benefit is provided.","rationale":"The reader's weakest assumption is precisely the accuracy and completeness of the hand-annotated 'Based on' edges, and the paper's own Section 2 admits that the annotation process was imperfect. That assumption is load-bearing for the abstract's claim that the graph allows researchers to process less information and notice overlooked models: if edges point to the wrong parent methods, or if many relevant methods are missing, the graph will mislead rather than inform. The paper provides no measurement of edge quality and no user study, so the central claim is currently untestable. The reader's verdict of UNVERDICTED is appropriate, and no additional objection beyond the absence of evidence is needed. I agree with the reader's identification of the weakest assumption; my concrete test would turn that assumption into a checkable claim. The paper does have some positive features: it is an honest proposal, it describes a concrete artifact, and it includes a community-editing mechanism that could in principle improve annotation quality over time. Those features do not, however, substitute for evidence that the current graph delivers the promised research benefit.","tokens_in":2644,"tokens_out":2346,"duration_ms":24968,"concrete_test":"Run a task-based evaluation with at least 20 ML researchers: randomly assign participants to use either Backronym or a conventional search/citation baseline to find a method that could improve a given architecture; measure time, number of papers examined, and novelty of discovered components. In parallel, have the authors of 50 randomly selected papers confirm or correct their 'Based on' entries and compute inter-annotator agreement (e.g., Cohen's kappa) between the original annotation and the author correction. If agreement is low or Backronym shows no benefit over baseline, the central claim fails; if agreement is high and the benefit is significant, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the inheritance graph lets researchers process less information and notice overlooked models—can only hold if the 'Based on' edges are accurate and sufficiently complete to support navigation. That condition is unsecured. Section 2 states that the analysis process 'is very far from ideal,' that many areas were unfamiliar, and that abstracts were sometimes substituted for method descriptions. The paper presents no dataset, no inter-annotator agreement measure, no error analysis, and no task-based evaluation showing users find useful methods faster with Backronym than without. Because the claimed benefit is a behavioral/utility outcome, hand-constructed edges and an interactive visualization are not by themselves evidence for it. This is not an internal contradiction; the paper is explicitly a white paper/proposal. But it makes the central claim unverdictable: there is currently no way to tell whether the graph helps or misleads researchers. The proposed recommendation system intensifies, rather than resolves, the dependence on accurate edges. The strongest defensible reading is therefore UNVERDICTED, not ACCEPT or REJECT.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This white paper introduces BACKRONYM, an interactive graph of machine learning methods connected by hand-annotated \"Based on\" inheritance edges. The authors argue that citation graphs do not reflect which methods are directly used in an architecture, and they propose a graph of method-level inheritance to help researchers process less information and discover previously overlooked models. The graph is built from approximately 250 NeurIPS 2019 papers plus 250 related papers, with metadata such as method name, subject area, and a \"Based on\" list of parent methods. The paper describes the annotation process, the interactive 3D visualization, and future plans for community editing and a recommendation system. No evaluation of the graph's utility is reported.","tokens_in":2810,"tokens_out":4125,"duration_ms":40312,"significance":"If the inheritance graph genuinely accelerated literature navigation and cross-idea discovery, it would be a valuable community resource for the growing machine learning literature. The authors deserve credit for releasing an interactive tool and for articulating a real limitation of citation graphs for tracing method reuse. However, the paper's central claim is behavioral and unmeasured: there is no user study, no quantitative evaluation, and no falsifiable prediction to support the assertion that the graph reduces information load or improves research. The utility of the graph rests on hand-curated edges whose accuracy and completeness are conceded to be imperfect. As a result, the current contribution is a design proposal and a dataset artifact, not a validated research result.","major_comments":[{"comment":"The central claim that the inheritance graph \"allows conducting research, processing much less information, and pay attention to previously unnoticed models\" is a causal, behavioral assertion. The paper reports no user study, no controlled task, no comparison against baseline navigation methods, and no quantitative measure of how the graph affects research outcomes. As written, the contribution is unverifiable. The authors should either provide an evaluation of the navigation benefit (for example, a task where users locate method variants with and without BACKRONYM) or explicitly reframe the claim as a design hypothesis, making the paper a proposal rather than a validated result.","section":"Abstract and Section 1"},{"comment":"The reliability of the hand-annotated \"Based on\" edges is load-bearing for every benefit claimed. Yet the authors state in Section 2 that \"The way we analyzed papers is very far from ideal\" and that abstracts were sometimes substituted for method descriptions. No inter-annotator agreement measure, error analysis, or comparison against a gold standard is provided. Without such evidence, the graph may mislead as easily as it helps, and the claimed navigation advantage is unsupported. At a minimum, the authors should quantify annotation uncertainty by independently re-annotating a random sample of papers and report agreement rates; the full 10-column dataset should also be made publicly available so others can audit the edges.","section":"Section 2"},{"comment":"The proposed recommendation system that \"will recommend him the most useful paper or method\" intensifies the dependence on accurate \"Based on\" edges, but no evaluation protocol for such recommendations is described. If this is part of the contribution, the authors need to specify a concrete baseline and metric, such as retrieval of known related methods or a user study comparing search time. Otherwise, it should be clearly labeled as speculation rather than a near-term capability, so readers can judge which parts of the paper are established.","section":"Section 3 (Future Work)"}],"minor_comments":[{"comment":"The sentence \"To represent such type of information as a table or a list it not optimal\" contains grammatical errors and should be revised to something like \"To represent such type of information as a table or a list is not optimal.\"","section":"Section 1"},{"comment":"The phrase \"research’s on a new conference do not inherit most of the improvements\" is unclear and grammatically incorrect; consider rewriting to \"researchers presenting at a new conference often do not inherit most of the improvements proposed a year ago.\"","section":"Section 1"},{"comment":"The wording \"In just a few months, we took out about 250 articles\" is informal; \"took out\" should be replaced with \"collected\" or \"curated.\"","section":"Section 2"},{"comment":"The paper references Figures 1, 2, and 3, but no figure captions are provided in the text. In a journal submission, every figure must have an explicit caption explaining what is shown.","section":"Section 2"},{"comment":"The dataset is described as a 10-column matrix, but the matrix itself is not included as a supplementary file. Since the graph is the main artifact, making the underlying data available is necessary for reproducibility and external verification.","section":"Section 2"},{"comment":"The phrase \"motivate more peoples to add their models\" should be \"motivate more people to add their models.\"","section":"Section 3"},{"comment":"The paper does not compare BACKRONYM to existing literature-discovery tools such as citation-based overlays, Semantic Scholar, or Connected Papers. A short discussion of the incremental contribution over these tools would help position the work.","section":"General"},{"comment":"The term \"BACKRONYM\" is used as the name of the graph, but the paper never explains why it is an acronym or what it abbreviates; consider defining the term explicitly if the name is intentional.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The manuscript reads more like a project white paper than a research article. If the journal considers resource/tool papers, the central utility claims still require evaluation or scoping down. A user study of navigation benefit and an annotation-quality analysis would be the minimum evidence needed to turn this into a publishable research paper. The authors should also check whether the paper fits the journal's scope, as it lacks a technical contribution in the form of a new algorithm or system architecture."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a project pitch, not a research result, but the underlying idea—a hand-built inheritance graph of ML methods—is genuinely worth a look. The authors say themselves that their analysis is 'very far from ideal,' and there is no evaluation to back the claim that this graph helps researchers process less information. So take the novelty seriously, the evidence claim with a grain of salt.\n\nWhat's new: the notion of 'Based on' edges between methods, as opposed to citation edges. That is a real distinction, and the paper is right that citation graphs don't capture method-level reuse. For a small sample (about 250 NeurIPS 2019 papers plus 250 predecessors), they have hand-annotated these edges and made an interactive visualization. The paper is transparent about where annotation was difficult, including using abstracts as fallback. That transparency is a point in its favor.\n\nWhat's missing: any way to tell whether the graph actually helps. No user study, no quantitative evaluation, no released dataset, no inter-annotator agreement. The central claim is a behavioral one, and it is simply asserted. The 'recommendation' future work in Section 3 would only amplify the dependence on edge quality, which the authors concede is imperfect. The five references are all standard and fine, but there is no comparison to existing tools like Papers with Code; that would have sharpened the positioning.\n\nWho this is for: someone curious about the idea of method-inheritance graphs, or a community that wants to contribute to it. Not a citable source for any empirical claim. As a peer-review candidate, I'd treat it as a workshop-style white paper, not a full submission—it needs at least the dataset and an evaluation, or a rigorous analysis of edge accuracy, before it can be reviewed as a research contribution.","headline":"A likable white paper with a genuinely interesting idea, but no data and no evaluation; the utility claim is entirely on trust.","tokens_in":3285,"tokens_out":2484,"would_cite":false,"duration_ms":72636,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Backronym maps which ML methods build on which","keywords":["machine learning","inheritance graph","method components","literature navigation","interactive visualization","Backronym","research tool"],"falsifier":"Take a random sample of, say, fifty methods represented in the Backronym graph and have two independent ML researchers annotate each method's architectural bases without seeing the graph; if the agreement between the two annotators, or between them and the graph's edges, is low, the inheritance links cannot support the proposed discovery benefit.","tokens_in":2455,"feed_emoji":"🕸️","tokens_out":4450,"duration_ms":43748,"temperature":0.7,"pith_summary":"This paper proposes that the machine-learning literature be organized not by citations or topic area, but by inheritance: directed links from each method to the methods and components it is built on. The author hand-annotated roughly 250 papers from one conference and 250 of the papers they build on, and released the result as an interactive graph called Backronym. The intended payoff is that a researcher who uses, say, a convolutional network can see more advanced descendants at a glance and replace their component without reading hundreds of papers. The paper does not evaluate whether this actually speeds research; it argues the annotation layer is valuable and should be community-maintained.","feed_headline":"Backronym maps which ML methods build on which","feed_subtitle":"Follow one edge from a method you know to its advanced descendants, without reading hundreds of papers.","key_machinery":"The central object is the Backronym graph, an interactive directed graph whose nodes are named methods (with metadata: paper, authors, method name, acronym, subject area, description) and whose edges are hand-annotated 'Based on' links indicating that one method directly uses another method or one of its components in its architecture. For example, an Adversarial Autoencoder receives an edge to Autoencoder and to Discriminator because it is built from those pieces. The graph is rendered as a 3D force-directed visualization, and its 'Subject area' column lets users form subgraphs showing which methods of one field are used in other fields. The mechanism that is supposed to carry the benefit is the 'skipping connections' use case: a user follows an inheritance edge from a known method to a descendant and swaps the descendant in.","core_discovery":"The central claim is that a graph of method-to-method inheritance, where an edge means that one method is directly based on or includes another method or a component of it, captures information about ML research that citation graphs miss, and that seeing these inheritance paths can direct researchers to methods that were lost among many papers. The paper calls this graph Backronym. The claim is demonstrated through the construction of a small annotated prototype and examples such as Adversarial Autoencoder inheriting from Autoencoder and Discriminator; the broader assertion is that this representation supports more creative research and eventually machine-generated recommendations for how to extend a model.","pith_inferences":["The graph's value is not argued empirically; one could test it directly by measuring whether researchers using Backronym discover and adopt a significantly broader set of base methods than researchers using keyword search alone.","A fully automated version could extract 'Based on' edges from paper texts, but the paper's own examples suggest accuracy will depend on resolving component-level reuse, not just citation context, making human annotation hard to replace.","If the graph grows, the 'skipping connections' pattern implies a concrete prediction: methods that are highly central in the inheritance graph should be the ones most often improved in later papers, so network centrality could serve as an early-warning signal for impactful work.","The historical example of a model being inspired by a training-phase concept suggests the annotation scheme may eventually need two edge types, architectural inheritance and conceptual inspiration, because they support different creative uses."],"forward_implications":["A researcher can start from a method they already know and follow inheritance edges to newer or less-known descendants, turning literature review into graph navigation.","Because each node stores the method's subject area and acronym, the same data can be sliced into per-field subgraphs, showing when a method from one area is used in another.","If authors add their own methods to the graph, the structure becomes community-maintained and can stay more accurate than a purely automatic citation graph.","Inheritance edges based on architectural components can reveal 'lost' methods, models with good results that few recent papers build on, which plain citation counts would not surface.","The same graph structure is intended to grow into a recommendation system that suggests which methods a researcher should look at next, and eventually to estimate an idea's impact on progress toward general machine intelligence."],"supporting_citations":[{"why":"Supplies the GAN example whose components Generator and Discriminator become nodes in the graph.","marker":"[1]"},{"why":"Supplies the Adversarial Autoencoder example that defines an inheritance edge to Autoencoder and Discriminator.","marker":"[2]"},{"why":"Supplies the Autoencoder node that AAE is based on, anchoring the edge semantics.","marker":"[3]"},{"why":"Supplies the CNN example used for the 'skip connections' navigation scenario.","marker":"[4]"},{"why":"Supplies the motivational anecdote that conceptual inspiration can link models even when no architectural inheritance exists.","marker":"[5]"}],"fun_headline_variants":["Backronym: a genealogy of ML methods","See which ML methods inherit from others","Backronym maps method-by-method lineage","Find new ML ideas via method inheritance","Backronym: discover hidden links in ML research"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire navigation benefit rests on the hand-annotated 'Based on' edges being complete and accurate enough to reflect which methods actually build on which, yet the paper concedes that many areas were unfamiliar to the annotator and abstracts were sometimes used instead of full descriptions.","fun_headline_variants_meta":{"raw":{"variants":["Backronym: a genealogy of ML methods","See which ML methods inherit from others","Backronym maps method-by-method lineage","Find new ML ideas via method inheritance","Backronym: discover hidden links in ML research"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000226,"raw_usage":{"total_tokens":1390,"prompt_tokens":790,"completion_tokens":600,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":406,"completion_tokens_details":{"reasoning_tokens":535}},"tokens_in":406,"tokens_out":600,"duration_ms":6018,"temperature":1.0,"reasoning_tokens":535,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:59:58.286288+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, fifty methods represented in the Backronym graph and have two independent ML researchers annotate each method's architectural bases without seeing the graph; if the agreement between the two annotators, or between them and the graph's edges, is low, the inheritance links cannot support the proposed discovery benefit.","supporting_citations":[{"cited_title":"Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, Yoshua Bengio","cited_arxiv_id":null,"evidence_quote":"Supplies the GAN example whose components Generator and Discriminator become nodes in the graph."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Autoencoder node that AAE is based on, anchoring the edge semantics."},{"cited_title":"Convolutional Neural Network","cited_arxiv_id":null,"evidence_quote":"Supplies the CNN example used for the 'skip connections' navigation scenario."},{"cited_title":"https://www.youtube.com/watch?v=Z6rxFNMGdn0 4 Images Figure 1: 3 WHITE PAPER - AUGUST 9, 2019 Figure 2: Figure 3: 4","cited_arxiv_id":null,"evidence_quote":"Supplies the motivational anecdote that conceptual inspiration can link models even when no architectural inheritance exists."}],"review_version":1}