Pith. sign in

REVIEW 4 major objections 6 minor 1 cited by

ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper claims that an LLM's commonsense knowledge can be repaired automatically by using a second model's plausibility scores to locate errors, abstracting the target knowledge into general concepts, and instantiating those concepts…

desk verdict A coherent and potentially useful commonsense KE pipeline whose main plausibility result is weakened by VERA serving as both edit selector and outcome measure; independent evidence is too thin. read the letter →

arxiv 2412.11418 v2 pith:X7LXZ742 submitted 2024-12-16 cs.CL

classification cs.CL
keywords knowledgeeditingcommonsensereasoningconceptualizationinstantiationplausibilityestimationlargelanguagemodelsautomatedverificationquestionanswering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ConKE's claim is that commonsense editing in LLMs can be made fully automated and scalable: instead of relying on manually curated gold labels, the pipeline uses a separate plausibility verifier (VERA) to detect which commonsense triples the target model generates implausibly, enriches those triples through conceptualization and instantiation so that each edit covers a family of related statements, and then writes the enriched knowledge back into the model with existing editing algorithms. The paper reports that edited LLMs generate more plausible commonsense according to VERA and to a 200-sample human review, and that the best edited backbone outperforms its unedited baseline on all five commonsense question-answering benchmarks tested (aNLI, SocialIQA, PIQA, CommonsenseQA, WinoGrande). A sympathetic reader would care because the approach attacks the two bottlenecks that have kept commonsense knowledge editing from scaling: there is no complete commonsense resource to copy from, and human annotation of erroneous beliefs is unaffordable at scale.

What carries the argument

The load-bearing machinery is the pairing of an automated plausibility judge with conceptualization–instantiation enrichment. VERA is a discriminative model that scores the plausibility of arbitrary commonsense statements on [0,1]; ConKE uses it both to select which generated triples to edit (score below 0.5) and to measure whether editing improved plausibility. Conceptualization abstracts a concrete event such as 'PersonX plays together every day' into a general concept such as 'PersonX engages in enjoyable group activities,' and instantiation produces new concrete variants, so one edit becomes a family of edits. The editing layer is a standard knowledge editor: ROME or MEMIT modify internal weights, and GRACE adds an external key–value adapter memory.

What would settle it

Have independent raters with measured inter-annotator agreement score a sample of the triples VERA flagged as implausible and a sample of post-edit generations, then compare per-item labels with VERA's scores; if agreement is near chance, or if editing the same number of randomly chosen triples produces the same downstream gains, the reported improvements are not attributable to VERA-guided conceptualized editing.

Watch

Extended reading notes

Core claim

The central claim, stated on the paper's own terms, is that conceptualization-augmented editing makes a small set of detected commonsense errors generalizable into broad corrections. ConKE first prompts the target LLM to complete (head, relation) pairs from AbstractATOMIC, then uses VERA to score the generated tails and selects those scoring below 0.5 as implausible knowledge that needs editing. For each selected triple, GPT-4o abstracts the event into a general concept and instantiates it into new concrete variants, yielding roughly 160,000 enriched commonsense triples. These are applied to the model through ROME, MEMIT, or GRACE. The paper reports improved VERA plausibility on held-out AbstractATOMIC triples, a matching trend in expert acceptance on 200 sampled generations, and downstream accuracy gains over the vanilla baseline on five commonsense QA benchmarks.

Load-bearing premise

The framework assumes VERA's plausibility scores are a trustworthy proxy for human commonsense, both when selecting which beliefs to edit and when judging whether the edits worked.

Editorial extensions

If this is right

  • Commonsense knowledge editing can run end-to-end without human-labeled errors, because VERA supplies the detection and evaluation signal.
  • One detected error becomes a family of edits, since conceptualization and instantiation expand the target triple into abstract and concrete variants, and this is what the authors claim improves generalizability.
  • The approach composes with existing editing algorithms and backbones: ROME, MEMIT, and GRACE all show plausibility gains across four LLMs.
  • If the paper is right, improving internal plausibility transfers to commonsense question answering: the best edited model outperforms the vanilla baseline on aNLI, SocialIQA, PIQA, CommonsenseQA, and WinoGrande.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because VERA supplies both the signal that selects edits and the metric that reports plausibility gains, part of the reported improvement may be alignment with VERA rather than with human commonsense; the downstream benchmark gains are the more independent evidence, though they are reported for only the best backbone and without error bars.
  • The authors do not test whether the pipeline depends on the specific scorer; swapping VERA for a different plausibility model or for human-annotated implausible triples would show whether the mechanism transfers.
  • The paper's own limitations section notes that edits can cascade through related concepts, that repeated updates risk knowledge drift, and that commonsense lacks stable ground truth; these caveats qualify how far the results extend beyond the controlled evaluation.
  • Because the editing format is open-ended rather than restricted to (head, relation, tail), the same detection–enrichment–edit loop could plausibly carry to other knowledge domains, provided a suitable plausibility scorer exists.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper proposes ConKE, an automated knowledge-editing pipeline for commonsense: VERA scores are used in Section 3.1 to select implausible (head, relation, tail) generations from an LLM; GPT-4o conceptualization and instantiation in Section 3.2 expand the triples to be edited; and MEMIT, ROME, or GRACE in Section 3.3 apply the edits. Evaluation in Section 4 measures a plausible rate on AbstractATOMIC test generations using VERA, adds a 200-sample human assessment, reports downstream accuracy on five QA benchmarks for the best model, and presents an ablation comparing VERA scores with and without conceptualization. The claimed contributions are a fully automated and scalable commonsense knowledge-editing framework and improved plausibility and QA performance.

Significance. If the central claim holds, ConKE would be a useful contribution: it replaces human annotation for detecting implausible commonsense knowledge, uses conceptualization to increase edit generalizability, and evaluates three editors across four backbones. The release of data, code, and models is a strength, as is the use of external downstream benchmarks. However, the evidence as presented does not yet establish the central plausibility claim: the primary outcome metric, VERA, is also the model used to decide which knowledge to edit, so part of the measured gain may be optimization toward VERA's threshold. The independent human check is small, lacks agreement statistics, and is reported only descriptively, while the downstream results are given for a single unnamed best model without variance. These gaps are fixable but currently leave the headline claim under-supported.

major comments (4)
  1. [Section 3.1 and Section 4.1] VERA is used both to select which triples to edit (score below 0.5 in Section 3.1) and to compute the primary outcome metric (plausible ratio above 0.5 in Section 4.1), so the reported plausibility gains are partly optimization toward VERA's threshold rather than independent evidence of improved commonsense. The 200-sample human evaluation in Section 4.1 is reported only as "similar improvements" with no inter-annotator agreement, no per-item acceptance counts, and no correlation with VERA scores; it is too small and too thinly described to break the feedback loop. Please add a larger human evaluation with agreement statistics, or an independent plausibility metric not involved in edit selection.
  2. [Section 4.2 and Figure 3] The downstream claim is based on a single "best LLM" that is never named, with one bar per benchmark, no error bars, no number of seeds, and no per-backbone or per-editor breakdown. Since Figure 2 shows that the editors behave differently, a single best-model comparison cannot support the abstract's claim of stronger performance across multiple QA benchmarks. Report all backbones and editors with exact accuracies and variance (or at least multiple seeds), and state whether the improvements are statistically significant.
  3. [Section 4.3 and Figure 4] The ablation's outcome is again average VERA scores, so it inherits the same feedback-loop concern as Sections 3.1 and 4.1. In addition, Figure 4 shows only Mistral-7B, Llama-3-8B, and ChatGLM2-6b, although Section 3.3 lists GPT-J-6B as one of the four backbones. Please include GPT-J-6B results and an evaluation metric that is independent of the edit-selection model.
  4. [Section 3.2] The conceptualization and instantiation steps are performed by GPT-4o with no reported verification that the roughly 160,000 generated triples preserve the semantic content of the original triples or that the abstract concepts are valid generalizations. Since the method's claimed advantage rests on these augmented triples being correct and generalizable, include sample outputs and a quality check (automatic or human) of the augmented knowledge, along with the number of triples that failed generation or filtering.
minor comments (6)
  1. [Title and Section 1] The framework is called both "ConceptEdit" and "ConKE" in the Abstract, Figure 1, and the introduction; please use one consistent name throughout.
  2. [Section 1] The sentence "implausible information within the LLM. implausible information within the LLM." contains a duplicated fragment and should be corrected.
  3. [Figure 2] The figure does not state how many generations were scored, how the 200-sample subset was drawn, or whether both annotators rated all items; please provide exact values and inter-annotator agreement.
  4. [Figure 3] The figure has no numeric values, no model name, and no indication of variance; a table with exact accuracies per benchmark would be more informative.
  5. [Section 4.1] The text alternates between "plausible ratio" and "plausible rate"; please use one term consistently.
  6. [Section 3.3 and Section 4.2] The phrase "significant performance improvements" in Section 4.2 is not supported by any significance test or confidence interval; please soften the wording or add statistical evidence.

Circularity Check

1 steps flagged · score 4.0 of 10

VERA is both the edit selector and the plausibility metric, and the human check samples only VERA-passed outputs; the plausibility gain is partially optimization toward the same judge, though external QA benchmarks keep the central claim from being fully circular.

  1. fitted input called prediction [Section 3.1 (Automated Knowledge Verification) and Section 4.1 (LLMs-After-Editing Evaluation)]
    "VERA then evaluates the plausibility of the generated knowledge by producing a score in the range [0, 1], where values above 0.5 are considered plausible, and those below 0.5 are deemed implausible. ... With the generations on the testing set, we ask VERA to score them again and we calculate the plausible ratio whose scores are above 0.5. ... we sample a subset of 200 generations ... on the acceptance ratio of the plausible assertions that passed VERA’s filtering."

    The same discriminative model, VERA, defines which knowledge is 'erroneous' (score < 0.5) and is the primary outcome measure (fraction of edited-model generations with score > 0.5). Edit targets are built from the VERA-flagged heads (§3.2: 'For each triple targeted for editing, we first abstract its instances into more general concepts by prompting GPT-4o'), so MEMIT/GRACE/ROME push the model toward outputs that should score higher under the same judge. The reported gains in VERA-plausible rate are therefore partly alignment with the selection metric, not independently established commonsense improvement. The §4.1 human check does not break the loop because it samples only 'plausible assertions that passed VERA’s filtering'; VERA false positives are excluded by design.

full rationale

ConKE's central contribution is a pipeline, not a single equation, so no quantity is defined as itself. The conceptualization step draws on the authors' prior CANDLE work and AbstractATOMIC, but those are published, code-released resources and the paper does not invoke them as an uncheckable uniqueness theorem; self-citation is therefore not load-bearing circularity. The genuine concern is the dual role of VERA: it selects which commonsense outputs to edit and then serves as the headline plausibility metric. Because the 200-item expert review is drawn only from generations that already passed VERA's threshold, the human check cannot measure VERA's false-positive rate, and the paper's own Limitations section concedes 'the lack of stable ground truth for commonsense.' However, the downstream results on SocialIQA, PIQA, aNLI, WinoGrande, and CommonsenseQA are external to VERA and reported as positive, and the edited tails are generated by GPT-4o rather than by VERA itself, so the central claim has independent content. A score of 4 reflects a partially self-referential evaluation loop rather than a fully forced derivation.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The central ledger: the paper's main free parameter is the VERA plausibility threshold (0.5) that gates which triples get edited and which outputs count as plausible. The axioms are domain assumptions about VERA's validity, GPT-4o's fidelity in abstraction/instantiation, and the safety of existing KE methods. No entirely new entities are introduced; the pipeline reuses previously proposed models and datasets. The combination is new, but the components are all imported.

free parameters (1)
  • VERA plausibility threshold = 0.5
    Threshold used in Section 3.1 to classify triples as plausible or implausible; chosen by hand, not derived. It determines which knowledge gets edited and which outputs count as plausible in evaluation.
assumptions (3)
  • domain assumption VERA scores reflect human plausibility judgments across diverse commonsense contexts.
    Used to both select edits (Section 3.1) and evaluate success (Section 4.1), inherited from Liu et al. 2023 without re-validation on the edited models' outputs.
  • ad hoc to paper GPT-4o generated conceptualizations and instantiations preserve the semantic content of the original triples.
    Section 3.2 assumes the abstraction/instantiation loop enriches coverage without distorting meaning; no automatic quality check or human validation of the 160k new triples.
  • domain assumption Existing KE methods (MEMIT, ROME, GRACE) inject triples with negligible side effects on unrelated knowledge.
    Standard assumptions in KE; the paper cites prior evaluations but does not test for side effects in the commonsense setting.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning." pith.science (2026). https://pith.science/paper/X7LXZ742

@misc{pith2026241211418,
  author       = {Pith},
  title        = {Pith review of: ConKE: Conceptualization-Augmented Knowledge Editing in Large Language Models for Commonsense Reasoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X7LXZ742}},
  note         = {Machine review of arXiv:2412.11418}
}
read the original abstract

Knowledge Editing (KE) aims to adjust a Large Language Model's (LLM) internal representations and parameters to correct inaccuracies and improve output consistency without incurring the computational expense of re-training the entire model. However, editing commonsense knowledge still faces difficulties, including limited knowledge coverage in existing resources, the infeasibility of annotating labels for an overabundance of commonsense knowledge, and the strict knowledge formats of current editing methods. In this paper, we address these challenges by presenting ConceptEdit, a framework that integrates conceptualization and instantiation into the KE pipeline for LLMs to enhance their commonsense reasoning capabilities. ConceptEdit dynamically diagnoses implausible commonsense knowledge within an LLM using another verifier LLM and augments the source knowledge to be edited with conceptualization for stronger generalizability. Experimental results demonstrate that LLMs enhanced with ConceptEdit successfully generate commonsense knowledge with improved plausibility compared to other baselines and achieve stronger performance across multiple question answering benchmarks. Our data, code, and models are publicly available at https://github.com/HKUST-KnowComp/ConKE.

Figures

Figures reproduced from arXiv: 2412.11418 by the authors.

Figure 1
Figure 1. An overview of CONKE, which pipelines conceptualization and instantiation, knowledge editing, and LLM verification together for automated and scalable knowledge editing over commonsense knowledge. go beyond traditional Knowledge Editing tech￾niques by combining automated knowledge de￾tection, conceptualization, and instantiation, en￾hancing the model’s ability to generalize and adapt to diverse contexts. Experimenta… view at source ↗
Figure 2
Figure 2. Average plausible rate and expert acceptance [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 4
Figure 4. VERA evaluation scores of edited LLMs with [PITH_FULL_IMAGE:figures/full_fig_p004_4.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EcomScriptBench: A Multi-task Benchmark for E-commerce Script Planning via Step-wise Intention-Driven Product Association

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark for e-commerce script planning shows that current LLMs struggle to associate products with script steps, and that injecting purchase intentions improves their performance.

Reference graph

Works this paper leans on

57 extracted references · 22 canonical work pages · cited by 1 Pith paper

  1. [1]

    Hwang, Chandra Bhagavatula, Kathleen R

    Emily Allaway, Jena D. Hwang, Chandra Bhagavatula, Kathleen R. McKeown, Doug Downey, and Yejin Choi. 2023. https://doi.org/10.18653/V1/2023.EACL-MAIN.192 Penguins don't fly: Reasoning about generics through instantiations and exceptions . In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, EACL 2...

  2. [2]

    Chandra Bhagavatula, Ronan Le Bras, Chaitanya Malaviya, Keisuke Sakaguchi, Ari Holtzman, Hannah Rashkin, Doug Downey, Wen - tau Yih, and Yejin Choi. 2020. https://openreview.net/forum?id=Byg1v1HKDB Abductive commonsense reasoning . In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020 . OpenReview.net

  3. [3]

    Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi. 2020. https://ojs.aaai.org/index.php/AAAI/article/view/6239 PIQA: reasoning about physical commonsense in natural language . In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, The Thirty-Second Innovative Applications of Artificial Intelligence Conference, I...

  4. [4]

    Nicola De Cao, Wilker Aziz, and Ivan Titov. 2021. https://doi.org/10.18653/V1/2021.EMNLP-MAIN.522 Editing factual knowledge in language models . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021 , pages 6491--6506. Association for Compu...

  5. [5]

    Chunkit Chan, Cheng Jiayang, Weiqi Wang, Yuxin Jiang, Tianqing Fang, Xin Liu, and Yangqiu Song. 2024. https://aclanthology.org/2024.findings-eacl.47 Exploring the potential of chatgpt on sentence level relations: A focus on temporal, causal, and discourse relations . In Findings of the Association for Computational Linguistics: EACL 2024, St. Julian's, Ma...

  6. [6]

    Ernest Davis and Gary Marcus. 2015. https://doi.org/10.1145/2701413 Commonsense reasoning and commonsense knowledge in artificial intelligence . Commun. ACM, 58(9):92–103

  7. [7]

    Wenxuan Ding, Weiqi Wang, Sze Heng Douglas Kwok, Minghao Liu, Tianqing Fang, Jiaxin Bai, Xin Liu, Changlong Yu, Zheng Li, Chen Luo, Qingyu Yin, Bing Yin, Junxian He, and Yangqiu Song. 2024. https://aclanthology.org/2024.findings-emnlp.123 Intentionqa: A benchmark for evaluating purchase intention comprehension abilities of language models in e-commerce . ...

  8. [8]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al - Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Zhang, Aur \' e lien Rodriguez, Austen Gregerson, Ava Spataru, Baptiste Rozi \` e...

Show all 57 references
  1. [9]

    Schubert

    Benjamin Van Durme, Phillip Michalak, and Lenhart K. Schubert. 2009. https://aclanthology.org/E09-1092/ Deriving generalized knowledge from corpora using wordnet abstraction . In EACL 2009, 12th Conference of the European Chapter of the Association for Computational Linguistic...

  2. [10]

    Do, Sehyun Choi, Weiqi Wang, and Yangqiu Song

    Tianqing Fang, Quyet V. Do, Sehyun Choi, Weiqi Wang, and Yangqiu Song. 2023. https://doi.org/10.48550/ARXIV.2304.10392 CKBP v2: An expert-annotated evaluation set for commonsense knowledge base population . CoRR, abs/2304.10392

  3. [11]

    Tianqing Fang, Weiqi Wang, Sehyun Choi, Shibo Hao, Hongming Zhang, Yangqiu Song, and Bin He. 2021 a . https://doi.org/10.18653/v1/2021.emnlp-main.705 Benchmarking commonsense knowledge base population with an effective evaluation dataset . In Proceedings of the 2021 Conference...

  4. [12]

    Tianqing Fang, Hongming Zhang, Weiqi Wang, Yangqiu Song, and Bin He. 2021 b . https://doi.org/10.1145/3442381.3450117 DISCOS: bridging the gap between discourse knowledge and commonsense knowledge . In WWW '21: The Web Conference 2021, Virtual Event / Ljubljana, Slovenia, Apri...

  5. [13]

    Tom Hartvigsen, Swami Sankaranarayanan, Hamid Palangi, Yoon Kim, and Marzyeh Ghassemi. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/95b6e2ff961580e03c0a662a63a71812-Abstract-Conference.html Aging with GRACE: lifelong model editing with discrete key-value adaptors ....

  6. [14]

    Mutian He, Tianqing Fang, Weiqi Wang, and Yangqiu Song. 2024. https://doi.org/10.1016/J.ARTINT.2024.104149 Acquiring and modeling abstract commonsense knowledge via conceptualization . Artif. Intell., 333:104149

  7. [15]

    Xiusheng Huang, Yequan Wang, Jun Zhao, and Kang Liu. 2024. https://aclanthology.org/2024.emnlp-main.826 Commonsense knowledge editing based on free-text in llms . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, ...

  8. [16]

    Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de Las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, L \' e lio Renard Lavaud, Marie - Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, ...

  9. [17]

    Tianjie Ju, Yijin Chen, Xinwei Yuan, Zhuosheng Zhang, Wei Du, Yubin Zheng, and Gongshen Liu. 2024. https://doi.org/10.18653/V1/2024.ACL-LONG.486 Investigating multi-hop factual shortcuts in knowledge editing of large language models . In Proceedings of the 62nd Annual Meeting ...

  10. [18]

    Ching Ming Samuel Lau, Weiqi Wang, Haochen Shi, Baixuan Xu, Jiaxin Bai, and Yangqiu Song. 2024. https://arxiv.org/abs/2410.14276 Ecomedit: An automated e-commerce knowledge editing framework for enhanced product and purchase intention understanding . Preprint, arXiv:2410.14276

  11. [19]

    Chunyang Li, Weiqi Wang, Tianshi Zheng, and Yangqiu Song. 2025. https://arxiv.org/abs/2502.16169 Patterns over principles: The fragility of inductive reasoning in llms under noisy observations . Preprint, arXiv:2502.16169

  12. [20]

    Smith, Yejin Choi, and Hannaneh Hajishirzi

    Jiacheng Liu, Wenya Wang, Dianzhuo Wang, Noah A. Smith, Yejin Choi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.81 Vera: A general-purpose plausibility estimation model for commonsense statements . In Proceedings of the 2023 Conference on Empiric...

  13. [21]

    Jingping Liu, Tao Chen, Chao Wang, Jiaqing Liang, Lihan Chen, Yanghua Xiao, Yunwen Chen, and Ke Jin. 2022. https://doi.org/10.1016/j.artint.2022.103744 Vocsk: Verb-oriented commonsense knowledge mining with taxonomy-guided induction . Artif. Intell., 310:103744

  14. [22]

    Kaixin Ma, Filip Ilievski, Jonathan Francis, Yonatan Bisk, Eric Nyberg, and Alessandro Oltramari. 2021 a . https://ojs.aaai.org/index.php/AAAI/article/view/17593 Knowledge-driven data construction for zero-shot evaluation in commonsense question answering . In Thirty-Fifth AAA...

  15. [23]

    Kaixin Ma, Filip Ilievski, Jonathan Francis, Satoru Ozaki, Eric Nyberg, and Alessandro Oltramari. 2021 b . https://doi.org/10.18653/V1/2021.EMNLP-MAIN.445 Exploring strategies for generalizable commonsense reasoning with pre-trained models . In Proceedings of the 2021 Conferen...

  16. [24]

    Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov. 2022. http://papers.nips.cc/paper\_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html Locating and editing factual associations in GPT . In Advances in Neural Information Processing System...

  17. [25]

    Andonian, Yonatan Belinkov, and David Bau

    Kevin Meng, Arnab Sen Sharma, Alex J. Andonian, Yonatan Belinkov, and David Bau. 2023. https://openreview.net/forum?id=MkbcAHIYgyS Mass-editing memory in a transformer . In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2...

  18. [26]

    George A. Miller. 1995. https://doi.org/10.1145/219717.219748 Wordnet: A lexical database for english . Commun. ACM , 38(11):39--41

  19. [27]

    Gregory Murphy. 2004. The big book of concepts. MIT press

  20. [28]

    OpenAI. 2024 a . https://openai.com/index/gpt-4o-mini-advancing-cost-efficient-intelligence/ Gpt-4o mini: advancing cost-efficient intelligence . OpenAI

  21. [29]

    OpenAI. 2024 b . https://openai.com/index/hello-gpt-4o/ Hello gpt-4o . OpenAI

  22. [30]

    Hao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin, Lei Hou, Juanzi Li, Zhiyuan Liu, and Qun Liu. 2022. https://doi.org/10.18653/V1/2022.EMNLP-MAIN.335 COPEN: probing conceptual knowledge in pre-trained language models . In Proceedings of the 2022 Conference on Empirical Method...

  23. [31]

    Vyas Raina, Adian Liusie, and Mark J. F. Gales. 2024. https://aclanthology.org/2024.emnlp-main.427 Is llm-as-a-judge robust? investigating universal adversarial attacks on zero-shot LLM assessment . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language...

  24. [32]

    Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2021. https://doi.org/10.1145/3474381 Winogrande: an adversarial winograd schema challenge at scale . Commun. ACM , 64(9):99--106

  25. [33]

    Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019. https://doi.org/10.18653/v1/D19-1454 Social iqa: Commonsense reasoning about social interactions . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9t...

  26. [34]

    Haochen Shi, Weiqi Wang, Tianqing Fang, Baixuan Xu, Wenxuan Ding, Xin Liu, and Yangqiu Song. 2023. https://aclanthology.org/2023.findings-emnlp.1023 QADYNAMICS: training dynamics-driven synthetic QA diagnostic for zero-shot commonsense question answering . In Findings of the A...

  27. [35]

    Haochen Shi, Tianshi Zheng, Weiqi Wang, Baixuan Xu, Chunyang Li, Chunkit Chan, Tao Fan, Yangqiu Song, and Qiang Yang. 2025. https://arxiv.org/abs/2505.16303 Inferencedynamics: Efficient routing across llms through structured capability and knowledge profiling . Preprint, arXiv...

  28. [36]

    Yangqiu Song, Haixun Wang, Zhongyuan Wang, Hongsong Li, and Weizhu Chen. 2011. https://doi.org/10.5591/978-1-57735-516-8/IJCAI11-388 Short text conceptualization using a probabilistic knowledgebase . In IJCAI 2011, Proceedings of the 22nd International Joint Conference on Arti...

  29. [37]

    Yangqiu Song, Shusen Wang, and Haixun Wang. 2015. http://ijcai.org/Abstract/15/537 Open domain short text conceptualization: A generative + descriptive modeling approach . In Proceedings of the Twenty-Fourth International Joint Conference on Artificial Intelligence, IJCAI 2015...

  30. [38]

    Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. https://doi.org/10.18653/v1/n19-1421 Commonsenseqa: A question answering challenge targeting commonsense knowledge . In Proceedings of the 2019 Conference of the North American Chapter of the Association...

  31. [39]

    Ben Wang and Aran Komatsuzaki. 2021. GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model . https://github.com/kingoflolz/mesh-transformer-jax

  32. [40]

    Jiaan Wang, Yunlong Liang, Zengkui Sun, Yuxuan Cao, Jiarong Xu, and Fandong Meng. 2024 a . https://doi.org/10.18653/V1/2024.ACL-LONG.627 Cross-lingual knowledge editing in large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Li...

  33. [41]

    Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, and Jundong Li. 2025 a . https://doi.org/10.1145/3698590 Knowledge editing for large language models: A survey . ACM Comput. Surv. , 57(3):59:1--59:37

  34. [42]

    Weiqi Wang, Limeng Cui, Xin Liu, Sreyashi Nag, Wenju Xu, Chen Luo, Sheikh Muhammad Sarwar, Yang Li, Hansu Gu, Hui Liu, Changlong Yu, Jiaxin Bai, Yifan Gao, Haiyang Zhang, Qi He, Shuiwang Ji, and Yangqiu Song. 2025 b . https://arxiv.org/abs/2505.15196 Ecomscriptbench: A multi-t...

  35. [43]

    Weiqi Wang, Tianqing Fang, Wenxuan Ding, Baixuan Xu, Xin Liu, Yangqiu Song, and Antoine Bosselut. 2023 a . https://aclanthology.org/2023.findings-emnlp.902 CAR : Conceptualization-augmented reasoner for zero-shot commonsense question answering . In Findings of the Association ...

  36. [44]

    Weiqi Wang, Tianqing Fang, Chunyang Li, Haochen Shi, Wenxuan Ding, Baixuan Xu, Zhaowei Wang, Jiaxin Bai, Xin Liu, Jiayang Cheng, Chunkit Chan, and Yangqiu Song. 2024 b . https://arxiv.org/abs/2401.07286 CANDLE: iterative conceptualization and instantiation distillation from la...

  37. [45]

    Weiqi Wang, Tianqing Fang, Haochen Shi, Baixuan Xu, Wenxuan Ding, Liyu Zhang, Wei Fan, Jiaxin Bai, Haoran Li, Xin Liu, and Yangqiu Song. 2024 c . https://doi.org/10.48550/ARXIV.2406.10885 On the role of entity and event level conceptualization in generalizable reasoning: A sur...

  38. [46]

    Weiqi Wang, Tianqing Fang, Baixuan Xu, Chun Yi Louis Bo, Yangqiu Song, and Lei Chen. 2023 b . https://aclanthology.org/2023.acl-long.733 CAT: A contextualized conceptualization and instantiation framework for commonsense reasoning . In Proceedings of the 61st Annual Meeting of...

  39. [47]

    Weiqi Wang and Yangqiu Song. 2024. https://doi.org/10.48550/ARXIV.2406.02106 MARS: benchmarking the metaphysical reasoning abilities of language models with a multi-task evaluation dataset . CoRR, abs/2406.02106

  40. [48]

    Peter West, Ronan Le Bras, Taylor Sorensen, Bill Yuchen Lin, Liwei Jiang, Ximing Lu, Khyathi Chandu, Jack Hessel, Ashutosh Baheti, Chandra Bhagavatula, and Yejin Choi. 2023. https://aclanthology.org/2023.findings-emnlp.80 Novacomet: Open commonsense foundation models with symb...

  41. [49]

    Wentao Wu, Hongsong Li, Haixun Wang, and Kenny Qili Zhu. 2012. https://doi.org/10.1145/2213836.2213891 Probase: a probabilistic taxonomy for text understanding . In Proceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD 2012, Scottsdale, AZ, USA, ...

  42. [50]

    Baixuan Xu, Chunyang Li, Weiqi Wang, Wei Fan, Tianshi Zheng, Haochen Shi, Tao Fan, Yangqiu Song, and Qiang Yang. 2025. https://arxiv.org/abs/2505.07313 Towards multi-agent reasoning systems for collaborative expertise delegation: An exploratory design study . Preprint, arXiv:2...

  43. [51]

    Baixuan Xu, Weiqi Wang, Haochen Shi, Wenxuan Ding, Huihao Jing, Tianqing Fang, Jiaxin Bai, Xin Liu, Changlong Yu, Zheng Li, Chen Luo, Qingyu Yin, Bing Yin, Long Chen, and Yangqiu Song. 2024 a . https://aclanthology.org/2024.emnlp-main.446 MIND: multimodal shopping intention di...

  44. [52]

    Derong Xu, Ziheng Zhang, Zhihong Zhu, Zhenxi Lin, Qidong Liu, Xian Wu, Tong Xu, Wanyu Wang, Yuyang Ye, Xiangyu Zhao, Enhong Chen, and Yefeng Zheng. 2024 b . https://doi.org/10.1145/3627673.3679673 Editing factual knowledge and explanatory ability of medical large language mode...

  45. [53]

    Zonglin Yang, Xinya Du, Erik Cambria, and Claire Cardie. 2023. https://doi.org/10.18653/V1/2023.EACL-MAIN.255 End-to-end case-based reasoning for commonsense knowledge base completion . In Proceedings of the 17th Conference of the European Chapter of the Association for Comput...

  46. [54]

    Aohan Zeng, Bin Xu, Bowen Wang, Chenhui Zhang, Da Yin, Diego Rojas, Guanyu Feng, Hanlin Zhao, Hanyu Lai, Hao Yu, Hongning Wang, Jiadai Sun, Jiajie Zhang, Jiale Cheng, Jiayi Gui, Jie Tang, Jing Zhang, Juanzi Li, Lei Zhao, Lindong Wu, Lucen Zhong, Mingdao Liu, Minlie Huang, Peng...

  47. [55]

    Ningyu Zhang, Yunzhi Yao, Bozhong Tian, Peng Wang, Shumin Deng, Mengru Wang, Zekun Xi, Shengyu Mao, Jintian Zhang, Yuansheng Ni, Siyuan Cheng, Ziwen Xu, Xin Xu, Jia - Chen Gu, Yong Jiang, Pengjun Xie, Fei Huang, Lei Liang, Zhiqiang Zhang, Xiaowei Zhu, Jun Zhou, and Huajun Chen...

  48. [56]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  49. [57]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.