Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-04T01:22:59.320550Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2605.23965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-04T01:22:59.320550Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
86 of 86 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c919cafc-370e-4ea2-9a85-e961ba16780b · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs In: Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 49b0abed-c7e0-4fc8-96d2-40228109f60b · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Nguyen and Raymond Choo
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9295fe30-bd9a-4d07-b2e1-9444c3393c50 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Metamorphic testing: A new approach for generating next test cases
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 69deb20b-1e0e-4236-8e96-24b8741c32ae · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c58ff1e9-5b43-43b4-ac5c-3bb8e65fda7e · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Available: https://doi.org/10.1145/3143561
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1f3824e9-75e9-4ac1-8ae4-7384ffb3553d · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 76915796-726a-4136-acb8-cf38ac796633 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 90a9a5b3-ba65-4a4a-b215-a06901d912c0 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Transformers as soft reasoners over language, in: Proceedings of the International Joint Conference on Artificial Intelligence, pp
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a5530c2a-1dd3-4213-9539-4b343dfc168d · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Errors of measurement in statistics
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c7c15b9b-3054-4dae-bd9d-9e1031c07d2c · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Explaining Answers with Entailment Trees
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9f41835a-81b5-477f-8442-673075927669 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs DeepSeek-V3 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b456f161-9fe1-458e-93db-c0eabacf0b26 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c026078b-80bb-48cd-89d9-cf1b91f96d02 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d9936f2d-0d02-4c2b-a7c6-a41a0c1f519b · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3de5d91f-d6a0-4968-b255-bc2f7da4b8a2 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs 70293–70332
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 567ac94d-5abf-4d8e-8d78-77e0ec07e2ca · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Logical consis- tency of large language models in fact-checking, in: The Thirteenth International Conference on Learning Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5884b6ab-5c89-4837-af9e-7543782b1e93 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Intelligent Virtual Assistants with LLM-based Process Automation
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e0e37fc7-d403-45c1-b13f-8329851b54bb · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs In: Zong, C., Xia, F., Li, W., Navigli, R
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8d57016a-6882-463f-b57d-9b1e2334ba69 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Conditional andmodalreasoninginlargelanguagemodels,in:Proceedingsofthe Conference on Empirical Methods in Natural Language Processing, pp
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 761453ab-d752-41f8-abbc-37edaa13976d · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d621aa70-15c9-43f8-a41e-698253c70746 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e730553b-e7cc-46aa-b033-315f3a7d7360 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Differential optimization testing of gremlin -based graph database systems
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cf47d0db-4f5c-42ca-8187-b4678cf96a46 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs 16889–16914
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 50870515-d37f-471b-ac3c-a553fe73202a · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 141045f4-5d0d-4f43-a129-22d2832051e5 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3ba912d-575a-4c3e-902f-b763d0593ec9 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Retrieval-augmentedgenerationforknowledge-intensive nlp tasks, in: Advances in Neural Information Processing Systems, Curran Associates, Inc
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f7edbb65-328e-4a3f-9ba5-b78afdf7d2e0 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Drowzee: Metamorphic testing for fact-conflicting hallucination detection in large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 361e4689-1d8b-4abd-a9a5-acbb087746ee · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Evaluating the Logical Reasoning Abilities of Large Reasoning Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 097a5bfd-dcf1-400d-a6db-4c8784b0699b · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Logiqa: a challenge dataset for machine reading comprehension with logical reasoning, in: Proceedings of the International Joint Confer- enceonArtificialIntelligence,pp.3622–3628
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a75b04b7-fb3e-4ecb-855a-68841d6a2c7e · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7106473c-3560-4ab8-b86c-64e58fdf2631 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Beyond accuracy: Evaluating the reasoning behavior of large language models - a survey, in: First Conference on Language Modeling
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a1695432-b589-4d47-bb6a-4ea95999471a · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b3d208d2-36cf-450c-82fe-5e9768379c85 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c020cf22-54e6-4ee3-8862-54876d6efa7e · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Logic-lm: Empoweringlargelanguagemodelswithsymbolicsolversforfaithful logical reasoning, in: Proceeding of the Conference on Empirical MethodsinNaturalLanguageProcessing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 26e45428-eb44-4d4f-9b67-bfc0ff1323eb · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Thinking Assistants: LLM-Based Conversational Assistants that Help Users Think By Asking rather than Answering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e587e0a-2aa3-4a2e-a35e-6b80a02fcf30 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs L ogic B ench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 996e1536-d018-462f-9f0b-507f8d883bd1 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Wu, T., Xiang, C., Wang, J
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8aa4b073-cb70-4e8b-a1b9-d407f4d92d41 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Large language models meet symbolic provers for logical reasoning eval- uation, in: Proceeding of the International Conference on Learning Representations
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5fcefd3b-0f24-4a5c-8a0d-7b12e8d627d5 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Code Llama: Open Foundation Models for Code
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9cb7a9a4-1098-49af-8edf-c73cf1ee7baf · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Language models are greedy reasoners: a systematic formal analysis of chain-of-thought, in: The International Conference on Learning Representations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fd744f83-7ee0-457d-9ec8-0fb17e6cb5ab · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs IEEE Trans
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0ed50b9b-3293-47c4-b70e-baf32f3f4eb9 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Arelargelanguagemodelsgoodatfuzzyreasoning?, in: Proceedings of the International Conference on Computational Intelligence and Intelligent Systems, pp
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 724c07b2-cc80-4923-ac60-26a35fd5ef2c · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Hamilton
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4adfa41b-15cf-42c0-83d4-cdfeb1b706da · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs https://doi.org/10.48550/arXiv.2509
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9fc98283-8805-4ea4-8143-1546618dba56 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Challenging big-bench tasks and whether chain- of-thought can solve them, in: Findings of the Association for Com- putational Linguistics: ACL 2023, pp
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f5526ee9-e77c-431d-8194-ba31e04e1482 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Proofwriter: generating implications,proofs,andabductivestatementsovernaturallanguage, in: Findings of the Association for Computational Linguistics: ACL- IJCNLP 2021, pp
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 02b26877-f7f0-42f2-976f-8c3b70984598 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Diagnosing the first-order logical reasoning ability through logicnli, in: Proceed- ings of the Conference on Empirical Methods in Natural Language Processing, pp
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 380caf70-ba8e-454c-9677-989314203926 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bf604990-9e0c-4f85-abab-2e391aac99a9 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs L ogic A sker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b3bc9a1d-3a94-4f74-9417-c615b97c96b3 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Self-consistency improves chain of thought reasoning in language models, in: The Eleventh International Conference on Learning Representations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 19c14d6e-f270-43ce-8b49-a1f8f0a7d988 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Chain-of-thought prompting elicits reasoning in large language models, in: Proceedings of the International Conference on Neural Information Processing Systems, pp
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation be7b517e-a2dd-446d-8b21-15534da6753f · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs A systematic literature review of hallucinations in large language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b3f56e49-bdfc-4f91-a37f-677e22296d41 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Detecting and reducing the factual hallucinations of large language models with metamorphic testing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3f7ff43d-132b-46d7-abed-247cf581fe40 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Testing and validating machine learning classifiers by metamorphic testing
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9f84d98f-c6a6-48a2-a632-d19753b3573e · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Are large language models really good logical reasoners? a comprehensive evaluation and beyond
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3732a0e4-37ef-43f7-948a-ee90bf729cdc · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Benchmarking Benchmark Leakage in Large Language Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 83d93925-478d-4b5d-b2ce-c27805ddb658 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Hal- lucination detection in large language models with metamorphic relations
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation abd913fb-f043-4d25-97ef-adfad4415ac1 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Hallucinationdetectionfor llm-based text-to-sql generation via two-stage metamorphic testing
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 21f4fbcb-3ce6-4d11-88c8-564b8f49ce22 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b83b8d35-db3a-4ec7-a384-4879c1534017 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs A Survey of Large Language Model Agents for Question Answering
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9e1e57e0-b856-4fc0-9400-cc06b9d86b74 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9956994b-9f2c-4558-a83a-d3a6aa8d7f92 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs From system 1 to system 2: A survey of reasoning large language models.IEEE Trans
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7f5945d8-1ede-4078-89c4-2905ac3df064 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c5f17d58-c272-416b-95de-72b47ca8d5c6 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 69b1aff6-d1b6-4e63-861f-847b6958c78a · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Toolqa: A dataset for llm question answering with external tools
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 85edcce7-d95e-47ac-b407-f09206245b63 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7ef1f811-ff52-4f55-a59c-215fd071bb9b · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Conclusion Tom is a citizen of Washington
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3253b911-d670-4b60-96e1-83c225d1b38f · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 01205fd1-49a3-451d-ac17-135c533e52e7 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d2dd9d6a-8e58-4a87-bec2-4c5227f500cc · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 77e38a68-47dc-422f-b674-710f23379cde · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Conclusion Tom is a citizen of Washington
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cd7bfd97-dd3e-4308-937d-d9881dac4ae0 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs label".↪ The value for
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 82ae885e-8989-42a1-8258-ff64b9b769e1 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6a0f9444-211f-44b3-8d51-7b24f87d9a65 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs label".↪ The value for
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 000b52ca-ca4c-4e2f-b62c-9c349291779d · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Your evaluation must rely strictly on formal logical structure
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8c24ea4c-f861-45f9-98a6-055f733af53a · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs reasoning
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2e6c7f7b-d62c-491d-a62a-f060daf60f3f · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs both A and B
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 60127795-46bd-4943-b75f-3a2668a4379c · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Jadiel is Bitter
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a4633576-4bee-4101-a7db-2592a12c1230 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs For all x,
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1a85ccb8-eeeb-4134-908c-ffafc99f4e71 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs it is not the case that it is not the case that A
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e0775d88-1763-4823-a7fb-59e0bccf95d9 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - AND (&), OR (|), implication (->), and biconditional (<->) must be preserved.↪
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 813bd88e-4e8b-4e0b-a06a-ab2d0651e585 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Pay close attention to negation scope
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ba2d1349-3528-466a-94e0-8a4483609312 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Do not swap universal and existential quantifiers
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 06a0bdd6-6a32-4ec6-82ed-3e2c8d144191 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Keep predicate/relation identity and argument order
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8d840226-1bdd-4e3f-a094-9caea48d114b · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Do not replace placeholder symbols such as Pre1 or Con1 with guessed original meanings.↪
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0846755d-3cad-48cf-a7bb-34160c4d77c6 · outbound
LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs False".↪ If the NL sentence faithfully preserves the FOL structure, return
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.