Pith. sign in

Paper Citation Record · LEDGER

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs

As of 20 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 0 inbound Pith citation observations for arXiv:2605.23965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23965 v3

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-04T01:22:59.320550Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact24
  • verified fuzzy30
  • unresolved13
  • parse uncertain0
  • malformed identifier5
  • metadata mismatch14

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c919cafc-370e-4ea2-9a85-e961ba16780b · outbound

This paper cites In: Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs In: Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V

Reference 1

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.775751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:94a2f1f0e2480c70b339fb7d686e8fdac0cb8098839f6d47066410bad577a051

Observation 49b0abed-c7e0-4fc8-96d2-40228109f60b · outbound

This paper cites Nguyen and Raymond Choo.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Nguyen and Raymond Choo

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:21.746311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:94ad255419dfd41e064866dc87995718f00a02f647346eeb9c2cdead889df4e7

Observation 9295fe30-bd9a-4d07-b2e1-9444c3393c50 · outbound

This paper cites Metamorphic testing: A new approach for generating next test cases.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Metamorphic testing: A new approach for generating next test cases

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.019546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:a78c699112fcfbdaf550315c2f4334e7ebd7f5ef035adc3edf558b5d54414816

Observation 69deb20b-1e0e-4236-8e96-24b8741c32ae · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.029154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:b7a05135f6ae0b7f72b0e49aa60bf317551bcfffff43efd82e49cbba60e11111

Observation c58ff1e9-5b43-43b4-ac5c-3bb8e65fda7e · outbound

This paper cites Available: https://doi.org/10.1145/3143561.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Available: https://doi.org/10.1145/3143561

Reference 5

Resolution
metadata mismatch
doi, observed 2026-07-04T01:29:21.748908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:537fc54b6b51301e970501a2a5ca1190b2e5418a93351fe3c09741959cd3b886

Observation 1f3824e9-75e9-4ac1-8ae4-7384ffb3553d · outbound

This paper cites Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Multilingual Knowledge Graph Completion with Self-Supervised Adaptive Graph Alignment

Reference 6

Resolution
malformed identifier
doi_truncated, observed 2026-07-04T01:29:21.708295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:4bb9cfb10b4b850a454829d21b6c7ed3abec40e76e498f1c9d68f07939f7995b

Observation 76915796-726a-4136-acb8-cf38ac796633 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T01:29:21.721448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:a3e82b642fed83b49cc3e6aa796f454414dbe776ac53be4c69b8e9bede5753f9

Observation 90a9a5b3-ba65-4a4a-b215-a06901d912c0 · outbound

This paper cites Transformers as soft reasoners over language, in: Proceedings of the International Joint Conference on Artificial Intelligence, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Transformers as soft reasoners over language, in: Proceedings of the International Joint Conference on Artificial Intelligence, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.032267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:9cae3ce9e9da9fcc27376fa0e404e34bee60bffc8b5717225f7230e93aaf3667

Observation a5530c2a-1dd3-4213-9539-4b343dfc168d · outbound

This paper cites Errors of measurement in statistics.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Errors of measurement in statistics

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.779147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:b38941edebfc02d6f8fbcaea3c42fde208230a4feee98b78944838d8e28ae2c1

Observation c7c15b9b-3054-4dae-bd9d-9e1031c07d2c · outbound

This paper cites Explaining Answers with Entailment Trees.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Explaining Answers with Entailment Trees

Reference 10

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.807703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:8c03fe6ba62e787e65d603c5db904bf167a446f21b00f6a336e695c12376b44c

Observation 9f41835a-81b5-477f-8442-673075927669 · outbound

This paper cites DeepSeek-V3 Technical Report.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs DeepSeek-V3 Technical Report

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T01:29:21.794051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:b5a927909ba7f64c8f69358c42c44f2d143341e390a1503aab9c24e948708783

Observation b456f161-9fe1-458e-93db-c0eabacf0b26 · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:10.992729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:d7d9b0c4482a01c0ee202e3acd61e40bd91b983ba4fd8f15275ea9eb10c7b205

Observation c026078b-80bb-48cd-89d9-cf1b91f96d02 · outbound

This paper cites DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T01:29:21.787927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:6cc053b3427054a5d64dfe0b2949f6924235858a4b7fe7517e752298e98c68da

Observation d9936f2d-0d02-4c2b-a7c6-a41a0c1f519b · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.054170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:94c1d01194ec52954ab699d261a71a030e976fe9ec5029341bdd40a0747d124a

Observation 3de5d91f-d6a0-4968-b255-bc2f7da4b8a2 · outbound

This paper cites 70293–70332.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs 70293–70332

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.997837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:17d96c7cdd2a02e521eeb0f902d6dd37dae68e42aaa970d89b55dcc97e66f739

Observation 567ac94d-5abf-4d8e-8d78-77e0ec07e2ca · outbound

This paper cites Logical consis- tency of large language models in fact-checking, in: The Thirteenth International Conference on Learning Representations.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Logical consis- tency of large language models in fact-checking, in: The Thirteenth International Conference on Learning Representations

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.895199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:a0d5925202101c6be31b7fbde0ee5c46c4bb79a46a8a348541bb0c0d4d1b7baa

Observation 5884b6ab-5c89-4837-af9e-7543782b1e93 · outbound

This paper cites Intelligent Virtual Assistants with LLM-based Process Automation.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Intelligent Virtual Assistants with LLM-based Process Automation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:22.059350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:71ae72a29e8aa3a404c2a8ede2907a511027a8e3bfb885ac55f506eef4d60220

Observation e0e37fc7-d403-45c1-b13f-8329851b54bb · outbound

This paper cites In: Zong, C., Xia, F., Li, W., Navigli, R.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs In: Zong, C., Xia, F., Li, W., Navigli, R

Reference 18

Resolution
malformed identifier
doi_truncated, observed 2026-07-04T01:29:21.817585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:60c26c3ee75d4f6e4e27958ff79c55ecf306cf1e2948fc7bd99f9ea9eef65645

Observation 8d57016a-6882-463f-b57d-9b1e2334ba69 · outbound

This paper cites Conditional andmodalreasoninginlargelanguagemodels,in:Proceedingsofthe Conference on Empirical Methods in Natural Language Processing, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Conditional andmodalreasoninginlargelanguagemodels,in:Proceedingsofthe Conference on Empirical Methods in Natural Language Processing, pp

Reference 19

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.754732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:034429aa44e23d3a24ff76b1b4373ea6dd598e63b6fe3c10298ddca342524581

Observation 761453ab-d752-41f8-abbc-37edaa13976d · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.003634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:f69cb6561ae6975c9629755646902d5c912591eed483fc5534f845de6f445a48

Observation d621aa70-15c9-43f8-a41e-698253c70746 · outbound

This paper cites A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions

Reference 21

Resolution
metadata mismatch
doi, observed 2026-07-04T01:29:21.724479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:148162e3f903b0b4c0e7ffb3bfe9bc5ba06d7d91376f86adb6cbd5726042c08d

Observation e730553b-e7cc-46aa-b033-315f3a7d7360 · outbound

This paper cites Differential optimization testing of gremlin -based graph database systems.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Differential optimization testing of gremlin -based graph database systems

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:21.734484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:025d558d373205f91df4ab9524daae4308b1cfbe7fca7dd31414ea608100d529

Observation cf47d0db-4f5c-42ca-8187-b4678cf96a46 · outbound

This paper cites 16889–16914.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs 16889–16914

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.059770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:b0a5e9c5aef2e0c74acf72617e039fac8410f8550615baef14ceba5ba6ae4889

Observation 50870515-d37f-471b-ac3c-a553fe73202a · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:10.960304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:362b7b3990b355ad9ab2e4af80303f7dffda2d79a234e42c302c902b87d409b6

Observation 141045f4-5d0d-4f43-a129-22d2832051e5 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T01:29:21.767038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:82bde9dd1eb01c13d7dd80936b465336e3c0854ca064ef7801963f9bd854c38d

Observation e3ba912d-575a-4c3e-902f-b763d0593ec9 · outbound

This paper cites Retrieval-augmentedgenerationforknowledge-intensive nlp tasks, in: Advances in Neural Information Processing Systems, Curran Associates, Inc.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Retrieval-augmentedgenerationforknowledge-intensive nlp tasks, in: Advances in Neural Information Processing Systems, Curran Associates, Inc

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.979099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:9a1f7bf7e2293fe02a3e0f5622f4be45d368bb7bd9df7b56812f1acc1a7f0224

Observation f7edbb65-328e-4a3f-9ba5-b78afdf7d2e0 · outbound

This paper cites Drowzee: Metamorphic testing for fact-conflicting hallucination detection in large language models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Drowzee: Metamorphic testing for fact-conflicting hallucination detection in large language models

Reference 27

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.815203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:48326c8a59398bf28c7d0655696d581da56eeac63bd5ec47c7baa91291324889

Observation 361e4689-1d8b-4abd-a9a5-acbb087746ee · outbound

This paper cites Evaluating the Logical Reasoning Abilities of Large Reasoning Models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Evaluating the Logical Reasoning Abilities of Large Reasoning Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.764000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:b905e0ca1e0f92fc666cc55fa70d444a9c10ee3b81023c7355d9eb3d59f4aa4e

Observation 097a5bfd-dcf1-400d-a6db-4c8784b0699b · outbound

This paper cites Logiqa: a challenge dataset for machine reading comprehension with logical reasoning, in: Proceedings of the International Joint Confer- enceonArtificialIntelligence,pp.3622–3628.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Logiqa: a challenge dataset for machine reading comprehension with logical reasoning, in: Proceedings of the International Joint Confer- enceonArtificialIntelligence,pp.3622–3628

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.714904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:dd0767507f444613fcf840335f1935a4e99966eab55d46943a9d47d3964f5b81

Observation a75b04b7-fb3e-4ecb-855a-68841d6a2c7e · outbound

This paper cites Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:22.064707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:7b7c0f3f6e1641e76d7f8db6d73f0bddbadbf38ded1daf6c3478b7133257751e

Observation 7106473c-3560-4ab8-b86c-64e58fdf2631 · outbound

This paper cites Beyond accuracy: Evaluating the reasoning behavior of large language models - a survey, in: First Conference on Language Modeling.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Beyond accuracy: Evaluating the reasoning behavior of large language models - a survey, in: First Conference on Language Modeling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.953204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:d7a2b7e30acf3c9bece0a4c26aaf0d7c88695041ed9770b2b8600cf9e54dd982

Observation a1695432-b589-4d47-bb6a-4ea95999471a · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:10.975749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:8133b778349c6f643c2ba7483d2ef2741cee1c531df215f844737d2912330ead

Observation b3d208d2-36cf-450c-82fe-5e9768379c85 · outbound

This paper cites Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Xin Zhao, and Ji-Rong Wen

Reference 33

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.751947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:2ef093a23034c270bca8732d77d081faa488c446fb851deed3bd8c1398bf5d2c

Observation c020cf22-54e6-4ee3-8862-54876d6efa7e · outbound

This paper cites Logic-lm: Empoweringlargelanguagemodelswithsymbolicsolversforfaithful logical reasoning, in: Proceeding of the Conference on Empirical MethodsinNaturalLanguageProcessing.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Logic-lm: Empoweringlargelanguagemodelswithsymbolicsolversforfaithful logical reasoning, in: Proceeding of the Conference on Empirical MethodsinNaturalLanguageProcessing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.935148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:7e065be58bc8ba38adc43acc31ce786f3b1010e9425d5de7c6919e7370631473

Observation 26e45428-eb44-4d4f-9b67-bfc0ff1323eb · outbound

This paper cites Thinking Assistants: LLM-Based Conversational Assistants that Help Users Think By Asking rather than Answering.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Thinking Assistants: LLM-Based Conversational Assistants that Help Users Think By Asking rather than Answering

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:22.061952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:adae09edd402cf083714537c2e9323eca3d062029e5435bc400c46e8b5825888

Observation 4e587e0a-2aa3-4a2e-a35e-6b80a02fcf30 · outbound

This paper cites L ogic B ench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs L ogic B ench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models

Reference 36

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.773800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:4f079d34f56057cf660e20cafbcd79a950d426329dc12b729fb7f4b8498ffc30

Observation 996e1536-d018-462f-9f0b-507f8d883bd1 · outbound

This paper cites Wu, T., Xiang, C., Wang, J.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Wu, T., Xiang, C., Wang, J

Reference 37

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.821972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:61dba3bdee6d00b3ba315975c447ce6a649443bc619eb7c2582303babaf9c6ec

Observation 8aa4b073-cb70-4e8b-a1b9-d407f4d92d41 · outbound

This paper cites Large language models meet symbolic provers for logical reasoning eval- uation, in: Proceeding of the International Conference on Learning Representations.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Large language models meet symbolic provers for logical reasoning eval- uation, in: Proceeding of the International Conference on Learning Representations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.901770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:6ae272d2942aef419db5ec4816813637b622f4efc7ff1d678124cead89145428

Observation 5fcefd3b-0f24-4a5c-8a0d-7b12e8d627d5 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Code Llama: Open Foundation Models for Code

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-04T01:29:22.056569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:9c926a0b60b6a11bc6b8752a230755197d93c9427415706f0d26dc1d5bace58c

Observation 9cb7a9a4-1098-49af-8edf-c73cf1ee7baf · outbound

This paper cites Language models are greedy reasoners: a systematic formal analysis of chain-of-thought, in: The International Conference on Learning Representations.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Language models are greedy reasoners: a systematic formal analysis of chain-of-thought, in: The International Conference on Learning Representations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.953791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:a33da73880c25c539a1cbb9ede0d893c2d6eb97da5476f55670b22958658c373

Observation fd744f83-7ee0-457d-9ec8-0fb17e6cb5ab · outbound

This paper cites IEEE Trans.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs IEEE Trans

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:21.785087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:bfd0efb750dc6ba4011fb21f642914727350926fce311771dfc1d3e424aae00e

Observation 0ed50b9b-3293-47c4-b70e-baf32f3f4eb9 · outbound

This paper cites Arelargelanguagemodelsgoodatfuzzyreasoning?, in: Proceedings of the International Conference on Computational Intelligence and Intelligent Systems, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Arelargelanguagemodelsgoodatfuzzyreasoning?, in: Proceedings of the International Conference on Computational Intelligence and Intelligent Systems, pp

Reference 42

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.729099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:83a2d89a426859d8b8a02aed9cfd4e7b5e9cf17956811d93da22835b8706ba4c

Observation 724c07b2-cc80-4923-ac60-26a35fd5ef2c · outbound

This paper cites Hamilton.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Hamilton

Reference 43

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.771896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:bbd51ef6be956558a89b5a1e6c400cdabfeb583ca06372e3aa4fcd9b10233925

Observation 4adfa41b-15cf-42c0-83d4-cdfeb1b706da · outbound

This paper cites https://doi.org/10.48550/arXiv.2509.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs https://doi.org/10.48550/arXiv.2509

Reference 44

Resolution
metadata mismatch
doi, observed 2026-07-04T01:29:21.737950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:e5e84a69d098e3877ac6116f78b659c947ebad2712a4bff63224c053e130387c

Observation 9fc98283-8805-4ea4-8143-1546618dba56 · outbound

This paper cites Challenging big-bench tasks and whether chain- of-thought can solve them, in: Findings of the Association for Com- putational Linguistics: ACL 2023, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Challenging big-bench tasks and whether chain- of-thought can solve them, in: Findings of the Association for Com- putational Linguistics: ACL 2023, pp

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-07-04T22:30:11.010980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:0ad804879e94db3e340bac274b7ac5dea50cf1464e9c1addc393d3f8efaa1628

Observation f5526ee9-e77c-431d-8194-ba31e04e1482 · outbound

This paper cites Proofwriter: generating implications,proofs,andabductivestatementsovernaturallanguage, in: Findings of the Association for Computational Linguistics: ACL- IJCNLP 2021, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Proofwriter: generating implications,proofs,andabductivestatementsovernaturallanguage, in: Findings of the Association for Computational Linguistics: ACL- IJCNLP 2021, pp

Reference 46

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.819712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:05d2a68fe44893cd7c8b84e3e51e255ebdffbc6d1937ce9e53fcb8982140cd1d

Observation 02b26877-f7f0-42f2-976f-8c3b70984598 · outbound

This paper cites Diagnosing the first-order logical reasoning ability through logicnli, in: Proceed- ings of the Conference on Empirical Methods in Natural Language Processing, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Diagnosing the first-order logical reasoning ability through logicnli, in: Proceed- ings of the Conference on Empirical Methods in Natural Language Processing, pp

Reference 47

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.769610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:0216dd58a0e78c997f5f5da53e2eb7ef07990430b3ba4cb6a3147abc8db62f59

Observation 380caf70-ba8e-454c-9677-989314203926 · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.013523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:7e333634a402b856afd965a255e74efa408a14851d8d55f9dff6f5d75634e82c

Observation bf604990-9e0c-4f85-abab-2e391aac99a9 · outbound

This paper cites L ogic A sker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs L ogic A sker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models

Reference 49

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.756880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:0583ae4e587c27df13f08d946051457eb9b8874647e4c5fb56b6eebb5a45139c

Observation b3bc9a1d-3a94-4f74-9417-c615b97c96b3 · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models, in: The Eleventh International Conference on Learning Representations.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Self-consistency improves chain of thought reasoning in language models, in: The Eleventh International Conference on Learning Representations

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.023768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:3004464e8627d966cde8a431180a57e8cd57d0ab1ae3bf23fd41eb44d8194e96

Observation 19c14d6e-f270-43ce-8b49-a1f8f0a7d988 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models, in: Proceedings of the International Conference on Neural Information Processing Systems, pp.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Chain-of-thought prompting elicits reasoning in large language models, in: Proceedings of the International Conference on Neural Information Processing Systems, pp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.002217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:6a8f42679ac28ce978bb4d179a114a848fbebe48c6c6a735ef6b0250f6639d4c

Observation be7b517e-a2dd-446d-8b21-15534da6753f · outbound

This paper cites A systematic literature review of hallucinations in large language models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs A systematic literature review of hallucinations in large language models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.800900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:d03d232837ad5915c872a916ffb92d7c0b2b14b5d755c21d43d2291537d8ea83

Observation b3f56e49-bdfc-4f91-a37f-677e22296d41 · outbound

This paper cites Detecting and reducing the factual hallucinations of large language models with metamorphic testing.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Detecting and reducing the factual hallucinations of large language models with metamorphic testing

Reference 53

Resolution
verified exact
doi, observed 2026-07-04T01:29:21.781952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:6b583b194e6d959a1f225934981546e00045ac7aba2c2e3eae472ec482b32d91

Observation 3f7ff43d-132b-46d7-abed-247cf581fe40 · outbound

This paper cites Testing and validating machine learning classifiers by metamorphic testing.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Testing and validating machine learning classifiers by metamorphic testing

Reference 54

Resolution
malformed identifier
raw_fallback, observed 2026-07-04T22:30:11.026338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:876312110a9f45c8c6170c450f07b04d4033f6eabaf9611c025d4192a5791ec2

Observation 9f84d98f-c6a6-48a2-a632-d19753b3573e · outbound

This paper cites Are large language models really good logical reasoners? a comprehensive evaluation and beyond.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Are large language models really good logical reasoners? a comprehensive evaluation and beyond

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.742406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:8afc5ce32bb7732ce9dc671a045ca576f0c184c1a89a0c965369505af0363cbf

Observation 3732a0e4-37ef-43f7-948a-ee90bf729cdc · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Benchmarking Benchmark Leakage in Large Language Models

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:21.760196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:0db18b33673898e609006b8f3ef2d35f7e2a8c5900d96b7468a40f87efd771ad

Observation 83d93925-478d-4b5d-b2ce-c27805ddb658 · outbound

This paper cites Hal- lucination detection in large language models with metamorphic relations.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Hal- lucination detection in large language models with metamorphic relations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.038349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:41f61bfc7fc930d7f1a1d161ff1aa22493ee8c133279d43d07138a2d9f932069

Observation abd913fb-f043-4d25-97ef-adfad4415ac1 · outbound

This paper cites Hallucinationdetectionfor llm-based text-to-sql generation via two-stage metamorphic testing.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Hallucinationdetectionfor llm-based text-to-sql generation via two-stage metamorphic testing

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.797634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:3320ddf2d15d0d4167c32b96c6a162a7ea1fc97f81a8f2f441fcb2652669d83d

Observation 21f4fbcb-3ce6-4d11-88c8-564b8f49ce22 · outbound

This paper cites ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:22.067780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:1939c832fa9540b28fba9c55856c96d7cce70e3043e9c19a08a0631aeb16d7dd

Observation b83b8d35-db3a-4ec7-a384-4879c1534017 · outbound

This paper cites A Survey of Large Language Model Agents for Question Answering.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs A Survey of Large Language Model Agents for Question Answering

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:21.804812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:26f50a3d5ae1ac48e9063001b1a132cd68250cc2fb941d11c3852665089887ce

Observation 9e1e57e0-b856-4fc0-9400-cc06b9d86b74 · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:10.967496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:b1a844df83534475c69b6a5a1bfba171128017c03f67453183d58ec4faf9b58b

Observation 9956994b-9f2c-4558-a83a-d3a6aa8d7f92 · outbound

This paper cites From system 1 to system 2: A survey of reasoning large language models.IEEE Trans.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs From system 1 to system 2: A survey of reasoning large language models.IEEE Trans

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T01:29:21.791024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:9edc0536708eca85f95857d1ef208b0ed9d20528d918c833b25da97bf274e541

Observation 7f5945d8-1ede-4078-89c4-2905ac3df064 · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:10.979695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:3ec17868657001c745228bd8f7887181f4c60dbb5c3549587d8f73e306e96b90

Observation c5f17d58-c272-416b-95de-72b47ca8d5c6 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T01:29:21.813060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:c6e998e54117500f1270dd68d0d847448c4cb7a9dec98bd14d48a853e27f9bfa

Observation 69b1aff6-d1b6-4e63-861f-847b6958c78a · outbound

This paper cites Toolqa: A dataset for llm question answering with external tools.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Toolqa: A dataset for llm question answering with external tools

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.026957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:3aa1a1a093108aba93147e86619381ee5c9dc4212189f178d658df346f17ecf6

Observation 85edcce7-d95e-47ac-b407-f09206245b63 · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.024340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:0576731f2908ef4a358e1375352c53220fbf3341f0dd416e566a294755153123

Observation 7ef1f811-ff52-4f55-a59c-215fd071bb9b · outbound

This paper cites Conclusion Tom is a citizen of Washington.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Conclusion Tom is a citizen of Washington

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.013535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:4130a687ef5e6b90eda83c3199782bfcebfe8f13a6e77029f1bc9c686399e6de

Observation 3253b911-d670-4b60-96e1-83c225d1b38f · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.049604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:f64715b7eaafc78f4be539647a54112c01b018d16472afaf9b299da6aa15ca4f

Observation 01205fd1-49a3-451d-ac17-135c533e52e7 · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.034060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:0f19b660eb52d69e05d98a99ec591e245cb26b0b15d3f646fedd8769a7220cb5

Observation d2dd9d6a-8e58-4a87-bec2-4c5227f500cc · outbound

This paper cites an unresolved cited work.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-07-04T22:30:11.008064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:f0cf4ba28ff05befdee4a06e7302f8015f00e69be001b55ab6a2bba0f21461c4

Observation 77e38a68-47dc-422f-b674-710f23379cde · outbound

This paper cites Conclusion Tom is a citizen of Washington.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Conclusion Tom is a citizen of Washington

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.043102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:a3ed6d8facb33ab9ef2421cd79b3532f75d6eb5d2bd32d4a46bb9ee9f21da06a

Observation cd7bfd97-dd3e-4308-937d-d9881dac4ae0 · outbound

This paper cites label".↪ The value for.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs label".↪ The value for

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.910976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:de76987822e5bc1a588131ff154b8178819ad05cb0a517232043259f03febb03

Observation 82ae885e-8989-42a1-8258-ff64b9b769e1 · outbound

This paper cites reasoning.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs reasoning

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.000222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:498c28fc9d98b0795e3e313551876e6308c26c91cb7ecf30d4af5a3df874a877

Observation 6a0f9444-211f-44b3-8d51-7b24f87d9a65 · outbound

This paper cites label".↪ The value for.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs label".↪ The value for

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.913685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:2252da62441b4a11a0f332e6260fd96096038bd107daa7fcc03acb7b11ca5738

Observation 000b52ca-ca4c-4e2f-b62c-9c349291779d · outbound

This paper cites Your evaluation must rely strictly on formal logical structure.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Your evaluation must rely strictly on formal logical structure

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.039830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:39799c9290e8c5bc95a275fa53faff9ab4071f4411545b6cac7f8c8ff9d3e3ae

Observation 8c24ea4c-f861-45f9-98a6-055f733af53a · outbound

This paper cites reasoning.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs reasoning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.010621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:717f557fdcce3ed88d348cae83df8a69e3609184bf08f72b22c0fd9d7ea9d3a8

Observation 2e6c7f7b-d62c-491d-a62a-f060daf60f3f · outbound

This paper cites both A and B.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs both A and B

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.950732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:01c7069d02981ab0aad4658191a06e73534ed31d1e3b1298d757a48be30db36d

Observation 60127795-46bd-4943-b75f-3a2668a4379c · outbound

This paper cites Jadiel is Bitter.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs Jadiel is Bitter

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.976511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:6b990b57ab75fc8c9f4ce46d9122cd65b1ad4b472f1874a94a127453d871afe7

Observation a4633576-4bee-4101-a7db-2592a12c1230 · outbound

This paper cites For all x,.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs For all x,

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.015843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:4abb126873e3168a62680b55c2c418f7757a9641ed80cbbc1b200c9c04342158

Observation 1a85ccb8-eeeb-4134-908c-ffafc99f4e71 · outbound

This paper cites it is not the case that it is not the case that A.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs it is not the case that it is not the case that A

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.037419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:fe77d063862ae208c6ba3f2482966b63cc3c7ace75d6287217abe74cd6432fcd

Observation e0775d88-1763-4823-a7fb-59e0bccf95d9 · outbound

This paper cites - AND (&), OR (|), implication (->), and biconditional (<->) must be preserved.↪.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - AND (&), OR (|), implication (->), and biconditional (<->) must be preserved.↪

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.020033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:826b33bc7b47fbb99a948d431dc0348fae9fce49f03f1f86ce5470a466ed8df4

Observation 813bd88e-4e8b-4e0b-a06a-ab2d0651e585 · outbound

This paper cites - Pay close attention to negation scope.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Pay close attention to negation scope

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.916778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:ebb6ea08cd9f0823a48b1e8943df397d4324bacc0e8567fdab5feb66502aa6ed

Observation ba2d1349-3528-466a-94e0-8a4483609312 · outbound

This paper cites - Do not swap universal and existential quantifiers.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Do not swap universal and existential quantifiers

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:11.041305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:006419f06381c9c2575779a2ab25f0dfe62380ec7297f836e60a19b335caa174

Observation 06a0bdd6-6a32-4ec6-82ed-3e2c8d144191 · outbound

This paper cites - Keep predicate/relation identity and argument order.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Keep predicate/relation identity and argument order

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.937285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:ebe2eb4889b3d065e3827cb3e3416385281cb18ce04279e804397d767236e9cf

Observation 8d840226-1bdd-4e3f-a094-9caea48d114b · outbound

This paper cites - Do not replace placeholder symbols such as Pre1 or Con1 with guessed original meanings.↪.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs - Do not replace placeholder symbols such as Pre1 or Con1 with guessed original meanings.↪

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-07-04T22:30:10.995763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:283571e761e65a46984ea12784c3ba904e5a249fbf67beeac66d29bf34db5f32

Observation 0846755d-3cad-48cf-a7bb-34160c4d77c6 · outbound

This paper cites False".↪ If the NL sentence faithfully preserves the FOL structure, return.

LGMT: Logic-Grounded Metamorphic Testing for Evaluating the Reasoning Reliability of LLMs False".↪ If the NL sentence faithfully preserves the FOL structure, return

Reference 91

Resolution
malformed identifier
raw_fallback, observed 2026-07-04T22:30:10.944087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-04T01:22:59.320550Z digest=sha256:ed9aca8f80970ac7a7c13587be018fddd10549023171fc0fb82a9c778976f0a1

Pith citing papers

No inbound Pith citation observations are available.