Pith. sign in

Paper Citation Record · LEDGER

HalluScore: Large Language Model Hallucination Question Answering Benchmark

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 2 inbound Pith citation observations for arXiv:2605.17007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17007 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T20:31:20.017866Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:47:41.319887Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T14:47:41.656143Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact17
  • verified fuzzy49
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee1713f0-de1a-4a51-ae93-0f7774c7255f · outbound

This paper cites On faithfulness and factuality in ab- stractive summarization.

HalluScore: Large Language Model Hallucination Question Answering Benchmark On faithfulness and factuality in ab- stractive summarization

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.010261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:457b32c689b6b90f849d51546ed96f255b501d3e726a8b54cc9be14da0e222f8

Observation 0e6f4f07-d000-4691-b095-07c8c1432a89 · outbound

This paper cites A survey on hallucination in large lan- guage models: Principles, taxonomy, challenges, and open questions.

HalluScore: Large Language Model Hallucination Question Answering Benchmark A survey on hallucination in large lan- guage models: Principles, taxonomy, challenges, and open questions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.008460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:25b3b2eadbf2f27da8a6b881499cfcf6034e964bb2896f35c70f9d908f809d44

Observation 7d517dd1-47c4-4289-adcb-c391a79f178f · outbound

This paper cites Arahallueval: A fine-grained hallucination evaluation framework for arabic llms.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Arahallueval: A fine-grained hallucination evaluation framework for arabic llms

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.006561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:3b7bf6c84cfbb234cabf5ef6b1f5f8dc2130244ed8de5b6383162698dba3909d

Observation 3bfb508f-7658-48e4-a9fd-daa32f797051 · outbound

This paper cites Survey of hallucination in natural language generation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Survey of hallucination in natural language generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.004803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:d11ca4c04ac5a44e72aca416e9da534f049c1d88be43078d1d3195a0ed7e3a23

Observation 108dd1fb-4c12-4896-ab06-e2f25628da6e · outbound

This paper cites A sur- vey of automatic hallucination evaluation on natural language generation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark A sur- vey of automatic hallucination evaluation on natural language generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.398882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:b9707e5d7522b9723ae383605e6bdff7e3bfffd608f804250325a7acdb41d7a6

Observation 35812170-7057-4218-88ab-839f62a898d6 · outbound

This paper cites Large language models hallucination: A comprehen- sive survey.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Large language models hallucination: A comprehen- sive survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.396206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:64ef01160c440cc8bd3553ecadd6640bbb095ba5f9df88852f9569a2593b8516

Observation f77be92e-deb7-487a-b3c0-821ec7192995 · outbound

This paper cites ALLaM: Large Language Models for Arabic and English.

HalluScore: Large Language Model Hallucination Question Answering Benchmark ALLaM: Large Language Models for Arabic and English

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.393536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:e6bdf9c4f6937bf3ac05330df8a626da3f7e4827603375f195e63ea642b5ba47

Observation 7008bd6d-bab8-4001-bf65-c191ed138861 · outbound

This paper cites Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.407453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:c5d49231a0b79e808900c79af893dd720c68f9e5238c6caef66c83860a59dd94

Observation 45a8a431-084d-4c3a-ac63-156a3986a408 · outbound

This paper cites Fanar: An Arabic-Centric Multimodal Generative AI Platform.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Fanar: An Arabic-Centric Multimodal Generative AI Platform

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.404492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:0e1c3fc0ee4c56bb47cdbb98d2889b2a159e9b50a266e8c2049d5394ffbafa27

Observation e68b2474-68db-4288-b396-c2ce47cf277f · outbound

This paper cites A Survey of Large Language Models for Arabic Language and its Dialects.

HalluScore: Large Language Model Hallucination Question Answering Benchmark A Survey of Large Language Models for Arabic Language and its Dialects

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T03:06:11.362819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:bde942d113516d2f8add0a1edefeab0cc19a7f7987cf45623006de96b020eee0

Observation 088dd01e-ee3d-4d35-a82d-d45eeb4446c8 · outbound

This paper cites Evaluating ara- bic large language models: A survey of bench- marks, methods, and gaps.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Evaluating ara- bic large language models: A survey of bench- marks, methods, and gaps

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.387652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:1c0af229450ca67c13ee93bb9c6111242d96355e46dd40eb63fb31c58029f3ea

Observation a2f8ef85-b251-4744-b345-e0ce56ab6cf1 · outbound

This paper cites Arabic natural language processing: Challenges and solutions.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Arabic natural language processing: Challenges and solutions

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.993415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:2d48ed36387c423ea54e78bf70f403940683aca6478eb00a2cf6dfb7b643e82e

Observation fa137c4c-29f5-405b-a223-7f7c7c73a60c · outbound

This paper cites an unresolved cited work.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-19T20:32:45.995389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:dbc415fd8f3708a977eef09f17cad0ad819b03cdaef858f14087ae8c034f456f

Observation 52f550ee-f6e0-45c0-9f34-c3ea298bd781 · outbound

This paper cites Halwasa: Quantify and analyze halluci- nations in large language models: Arabic as a case study.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Halwasa: Quantify and analyze halluci- nations in large language models: Arabic as a case study

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.991633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:a5cfd58e8712e43fc589c839129f475ccc62bf89365b75dfedad4c03dca42590

Observation c8ef8c2d-32c9-4c8d-8b7a-c92dc04dbc7d · outbound

This paper cites HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations.

HalluScore: Large Language Model Hallucination Question Answering Benchmark HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:32:45.422579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:3786ed0d34d450ce80c6dd9b6d5c2fca7a0abd7efc2e0efdb931878bc0f99aca

Observation 4284bdd8-4407-484c-9fb7-0cd5b53c122f · outbound

This paper cites Aftina: enhanc- ing stability and preventing hallucination in ai- based islamic fatwa generation using llms and rag.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Aftina: enhanc- ing stability and preventing hallucination in ai- based islamic fatwa generation using llms and rag

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.000676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:9d6d927ff2f144ed70d65581d0c46a7f4a5416d93e695cef1f51aad2d9fcbad5

Observation e66ab6ac-9dcf-47de-a318-875b8627cc5f · outbound

This paper cites Islamiceval2025: The first shared task of capturing llms hallucination in islamic content.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Islamiceval2025: The first shared task of capturing llms hallucination in islamic content

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.987956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:38465e45bf7dd4babd7f20367adcbeb82d4f6c7e0d5b669f2d2b7ea4b5115fa2

Observation e244d198-3317-4186-b063-28b3a535b664 · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T20:32:45.413600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:117ea661621873549a393871b11c2db4d44b81213db160e8342e25897a729d4f

Observation 9f6835c7-52c9-4767-84e8-c9248efd34ee · outbound

This paper cites Analyzing llm behavior in dialogue sum- marization: Unveiling circumstantial hallucina- tion trends.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Analyzing llm behavior in dialogue sum- marization: Unveiling circumstantial hallucina- tion trends

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.986276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:fc2f9e37afb38d6546a5a7e37e7aa438d80c247db34241bcc4e6905c739fe34b

Observation 765d1820-4d41-431a-975f-1118eabfb459 · outbound

This paper cites Evaluating Hallucinations in Chinese Large Language Models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Evaluating Hallucinations in Chinese Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.384974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:a9de66d775c1643e6c1ded8dfe5b20d27e510630a658dcbe5765381ffee1767c

Observation fcce44a8-d1c1-457a-9085-c44a0e70c1d4 · outbound

This paper cites Uhgeval: Benchmarking the hallucina- tion of chinese large language models via uncon- strained generation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Uhgeval: Benchmarking the hallucina- tion of chinese large language models via uncon- strained generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.982696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:210bab77bb06b577d0dc77764879bd3ccf28670d1e5c9096312b4d7729a33f5e

Observation 45e0857c-6403-4820-ae6a-321d78502b29 · outbound

This paper cites C-faith: A chinese fine-grained benchmark for automated halluci- nation evaluation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark C-faith: A chinese fine-grained benchmark for automated halluci- nation evaluation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.980962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:34a3b1d82185b6b9ea7c49b2425dedbcb50234a6995b186b81f978f7643478ec

Observation e822f75a-c0eb-41d2-83d6-e060f18f6d17 · outbound

This paper cites Retrieve only when it needs: Adap- tive retrieval augmentation for hallucination mitigation in large language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Retrieve only when it needs: Adap- tive retrieval augmentation for hallucination mitigation in large language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.416655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:85ad0780c9535c6fe9cfc2033ba2497dba56a48ef061de634d3d7ebf0febafef

Observation 732f6da4-14ad-472f-b31c-39a3f6efdcf3 · outbound

This paper cites Exploring rag solu- tions to reduce hallucinations in llms.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Exploring rag solu- tions to reduce hallucinations in llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.984339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:0d6d98f350f29d5a6ecb68e3fef212e5404378db99e6cdb9b0438541441b15ea

Observation 32431627-2742-4a0c-8a81-000a72677b30 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Detecting hallucinations in large language models using semantic entropy

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.989848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:ad5c2ed72c46b66cf1d8f0b03df2e849b0c293d7c0779785039fe24392df30b9

Observation 05e7868a-31f7-48fa-8916-8205e410034e · outbound

This paper cites En- hancing uncertainty-based hallucination detec- tion with stronger focus.

HalluScore: Large Language Model Hallucination Question Answering Benchmark En- hancing uncertainty-based hallucination detec- tion with stronger focus

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.002530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:883832a64e69a6f823e0d7ed813b665470a54799b0f4aaa285780a4baac8c99d

Observation fb2d3bcd-8392-4535-b3a3-ca96d6c546d0 · outbound

This paper cites Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity even better.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Detecting and mitigating hallucinations in machine translation: Model internal workings alone do well, sentence similarity even better

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.975266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:c2f092e10f089209a8897c5246bfcb5954db0a0c202cbae0c44bb629380e0619

Observation d8485e3d-3a11-4ee5-a43e-0d9518e2f22c · outbound

This paper cites Leveraging graph structures to de- tect hallucinations in large language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Leveraging graph structures to de- tect hallucinations in large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.977511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:2df2299ce72d3821d6abc5163ff3f71a33274d20509dd148732cb01586a63300

Observation 9dea6609-a06f-45f6-bf2d-38569e678140 · outbound

This paper cites Halugnn: Hallucination detection in large language models using graph neural network.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Halugnn: Hallucination detection in large language models using graph neural network

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.969895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:8955a8fa6c8b9db34a757027290603cff9ab9b8f1aeeadc4f28a73fa015a783b

Observation 915cc98f-6616-4985-9cd7-50b99aeb8595 · outbound

This paper cites Hallushift: Measuring distribution shifts towards hallucination detection in llms.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Hallushift: Measuring distribution shifts towards hallucination detection in llms

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.024604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:9ded11588571ddf6b9e3f532f80ecf456cf733fe4f5a01e4494a5b4023d2e640

Observation 344b517f-b600-4b25-8dd4-66fd0219f26f · outbound

This paper cites Selfcheck- gpt: Zero-resource black-box hallucination de- tection for generative large language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Selfcheck- gpt: Zero-resource black-box hallucination de- tection for generative large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.968141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:e380e1fd21c75ad5f72a77989e1b51fdf405c16f5e4411bf55b78fb98c07a9f4

Observation 6dcab520-1454-4c62-bed1-9e2c31e52141 · outbound

This paper cites Sac3: reliable hallucination detection in black- box language models via semantic-aware cross- check consistency.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Sac3: reliable hallucination detection in black- box language models via semantic-aware cross- check consistency

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.971702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:a7e23a23ebbc4b0e38a81d1942e931d01f3709a0d94b8f99164d84af3ff0331c

Observation 4bab83ab-45ee-4848-829b-7c2ecc20ba49 · outbound

This paper cites Ai- generated news articles based on large language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Ai- generated news articles based on large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.979226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:8d3373bf9ade14b3bca520917ef1516dbfd2103c2b6a188e19c78f2af0b8637b

Observation ee8fde59-8d33-4e06-a247-8251f52e8388 · outbound

This paper cites Self- expertise: knowledge-based instruction dataset augmentation for a legal expert language model.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Self- expertise: knowledge-based instruction dataset augmentation for a legal expert language model

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.966303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:a516149c27df96b9bd0d9beb729d0cce7830f12a2aaa96f862eaea6141e81ae5

Observation 8fcd387a-6c98-440b-898b-bd3952538b60 · outbound

This paper cites A survey on rag with llms.

HalluScore: Large Language Model Hallucination Question Answering Benchmark A survey on rag with llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.973404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:74c9fbe23de52b428c6083e886c139f6097416d8be1ec888390bec06a31a402c

Observation 9eb4c08a-852f-4144-9344-6876065e7ad8 · outbound

This paper cites Chain- of-thought prompting elicits reasoning in large language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Chain- of-thought prompting elicits reasoning in large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.997193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:6ca8688927ea05da3f2bee04669c527105dbd3e3325764d0d0d559d7b3faa06f

Observation 4bc90e49-ac48-4db5-b679-254a5ab2acd8 · outbound

This paper cites Chain- of-verification reduces hallucination in large lan- guage models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Chain- of-verification reduces hallucination in large lan- guage models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.962495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:70740cd67b58cfe0d4a820f155ee56b7fe4ddda795d60dc9b47a32cea03fd725

Observation da7a8fee-6ddd-4625-89c9-ec350d4eaa90 · outbound

This paper cites Mitigating Large Language Model Hallucination with Faithful Finetuning.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Mitigating Large Language Model Hallucination with Faithful Finetuning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.431359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:b6249959a6eb2e996b89e9eca9a5b089d4d15e9ce5b31c770bd5ca7d20974d44

Observation a2be840f-0364-4a9b-8c33-d35a1f38f977 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:32:45.419449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:f54bd52fd109ece568a61a3c24a014ce9e7d1314d82584f6f33c36e915a6f115

Observation 80be3781-54ac-46db-9c66-78c9bffd56bc · outbound

This paper cites Feqa: A ques- tion answering evaluation framework for faith- fulness assessment in abstractive summariza- tion.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Feqa: A ques- tion answering evaluation framework for faith- fulness assessment in abstractive summariza- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.960680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:c8092902b9c27aac8623e7c494864a438f81504b970616171098679dddff2842

Observation 3493b1ca-a1f0-4307-b9c2-6b0ae8d5f3a8 · outbound

This paper cites Evaluating the factual consistency of abstractive text summarization.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Evaluating the factual consistency of abstractive text summarization

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.964306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:f887ef7129c4c9de64b9763aef757141c720d10f3a8139bc0a7eb93e280b889c

Observation 9795a8df-c8a7-4829-b5a5-a710bb42fcd2 · outbound

This paper cites Factscore: Fine-grained atomic evalu- ation of factual precision in long form text gen- eration.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Factscore: Fine-grained atomic evalu- ation of factual precision in long form text gen- eration

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:45.998948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:824f2e7620dd4e955aa57604969ba1603719212eb0c8451d85e14c935efd414e

Observation cec35b2f-6dfc-457a-8da9-9672e421a43d · outbound

This paper cites How reliable are automatic eval- uation methods for instruction-tuned llms?.

HalluScore: Large Language Model Hallucination Question Answering Benchmark How reliable are automatic eval- uation methods for instruction-tuned llms?

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.012052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:4e37c13c133a4d9346d11fa3b780068bc4e65449192fda3593057bd7e0b7c4cb

Observation c8fa316e-4275-4c73-af19-1907bff5e92e · outbound

This paper cites Truthfulqa: Measuring how models mimic human false- hoods.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Truthfulqa: Measuring how models mimic human false- hoods

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.046200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:e1ee5c2dde92cbf0fd943035492d0c1a1ba23459c2d0fe4836d260a3ff787590

Observation 50996ebd-c92f-4f9b-bf00-ef85b2626835 · outbound

This paper cites Freshllms: Refreshing large language models with search engine augmentation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Freshllms: Refreshing large language models with search engine augmentation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.048077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:78f93943758001b79fff35d1f6dadd5a75f1041dd20a5ac57d471a02084a2cf1

Observation 8b55cd7f-f84b-48a9-9cba-e2c77c9beaab · outbound

This paper cites Assessing the factual accuracy of generated text.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Assessing the factual accuracy of generated text

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.037179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:d7a5a327356df995033040f96e57d265893ff03dc8964275480462a9d8676caa

Observation 1f16aa9a-a40c-42ca-84b3-35d368f2d36d · outbound

This paper cites Generativeaiforislamictexts: Theeman framework for mitigating gpt hallucinations.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Generativeaiforislamictexts: Theeman framework for mitigating gpt hallucinations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.039013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:83b330e0372acb6d27df0185aea3cfff22cd5fb1871229c85456c921ca3d91ef

Observation 111d5b67-34b1-4a61-8fed-4a6cfd66c4f5 · outbound

This paper cites Mitigating llm hal- lucinations in quranic content: An agentic ap- proach using deployable language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Mitigating llm hal- lucinations in quranic content: An agentic ap- proach using deployable language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.040806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:97536a693f8d5474fc5ae124650227cd3e0a076d6f36c5a06e42d544b2bd8aaa

Observation 112953d5-8999-4a6e-8327-c6b27314b492 · outbound

This paper cites SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes.

HalluScore: Large Language Model Hallucination Question Answering Benchmark SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.390226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:07f3b147069c49f72afaf39273c60ca27793f0151564c5d1b4f13903bebd97bc

Observation ed6af985-1c9e-4b1f-9015-3dbe75261a1c · outbound

This paper cites Halomi: A manually anno- tated benchmark for multilingual hallucination and omission detection in machine translation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Halomi: A manually anno- tated benchmark for multilingual hallucination and omission detection in machine translation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.049786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:59024b448e33922f6253ab3ec0553702785b72d1ed34a93387f8e42a400ab541

Observation 6ed730fe-c1c8-41ec-81ea-d1f6ad11c87c · outbound

This paper cites Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:32:45.428432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:e8b76cf31fb14aa803096939b98ccb101d1e22cbcaccc624129a79c857d697cd

Observation a71aa15b-e1ec-4b42-bc2d-a1b438c8d02f · outbound

This paper cites Hot- potqa: A dataset for diverse, explainable multi- hop question answering.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Hot- potqa: A dataset for diverse, explainable multi- hop question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.029774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:f9991c6f5e2f8a3364dbca1bfe8e51cf492117e861972218707c1aebc1fc6d59

Observation 16475eef-9b6b-49f9-a441-84d848b51268 · outbound

This paper cites Triviaqa: A large scale distantly super- vised challenge dataset for reading comprehen- sion.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Triviaqa: A large scale distantly super- vised challenge dataset for reading comprehen- sion

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.031635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:ba01db009354bf880cfccd79c5c992e5d44152598c4b52078308df615f761aba

Observation 52f2babf-f010-475f-a35c-a432d35dbf70 · outbound

This paper cites Medhallu: A comprehen- sive benchmark for detecting medical hallucina- tions in large language models.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Medhallu: A comprehen- sive benchmark for detecting medical hallucina- tions in large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.033725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:485a37c72e84aa7de69c607d70be1db472246df7ec534f8c845f810754335297

Observation 77476ed3-a272-4f28-b306-c70fd2f09690 · outbound

This paper cites Defan: Definitive answer dataset for llm hallucination evaluation.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Defan: Definitive answer dataset for llm hallucination evaluation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.035447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:30c6af67c41f183763af69a85c95caa04e97337877ec290c8774ba69ccc584e6

Observation 82dba318-3575-4609-bf53-d8c0cc031e17 · outbound

This paper cites Naseej launches its innovative arabic ai language model “noon.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Naseej launches its innovative arabic ai language model “noon

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.026395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:0b2e593547480a97c64f740fc411303e95b0721bd84f4fa06c5021904e3d1911

Observation 29b5d582-a1e3-40b0-a1d5-a4903d8d0e0a · outbound

This paper cites Introducing claude sonnet 4.5.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Introducing claude sonnet 4.5

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.028020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:0c3b7f0d1da9d16a90676ef01f0f9ed50c038b0a5dee4839633375fa9286eeca

Observation 50ba481a-9c98-4ea7-a466-81b453853f01 · outbound

This paper cites DeepSeek-V3 Technical Report.

HalluScore: Large Language Model Hallucination Question Answering Benchmark DeepSeek-V3 Technical Report

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:32:45.433934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:3745b2cf48d93a97c32de75cb5b5dd55c15539023fbee79ec34ab0523b94bdab

Observation 0fa784af-b4e4-439f-8aac-f74c763ccd72 · outbound

This paper cites [Online].

HalluScore: Large Language Model Hallucination Question Answering Benchmark [Online]

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.042683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:487a98f4ff489d43bc660af9c348c0bc4fb26e339a3f3d1a0f72d00612378afb

Observation 3a169845-6b2f-4138-b956-fd2368894682 · outbound

This paper cites GPT-4 Technical Report.

HalluScore: Large Language Model Hallucination Question Answering Benchmark GPT-4 Technical Report

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:32:45.401489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:21fa2f7c543f0fcaf40845c1ce95475eccf0ccfa8d7d27bbeecc73ed0be32d2d

Observation c0031f15-fee3-425e-aa18-f2d29992e9be · outbound

This paper cites OpenAI GPT-5 System Card.

HalluScore: Large Language Model Hallucination Question Answering Benchmark OpenAI GPT-5 System Card

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:32:45.425340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:29c0d02dc044321d097488afc4956f09a539559500b9f2742ed635423439c240

Observation 066260b9-1c48-4b2d-869a-edb5f4840def · outbound

This paper cites Llama-4-maverick- 17b-128e-instruct-fp8.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Llama-4-maverick- 17b-128e-instruct-fp8

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.017577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:6f215d052bc4c1036f0f766f3d1d96621c31deb8e4c67392e7d5791e25b64ccd

Observation b6a81e73-62c6-41fa-ba3e-6c588ff67bca · outbound

This paper cites Qwen3-next- 80b-a3b-instruct.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Qwen3-next- 80b-a3b-instruct

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.019311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:9357bec91eeff9bbabc432c9526fa0b4a30cd667c67ef5cf755f77078ac7463f

Observation dc8e0658-9aad-4438-8202-da09e65ad465 · outbound

This paper cites Qwen3-235b-a22b-instruct-2507-fp8.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Qwen3-235b-a22b-instruct-2507-fp8

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.020921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:9dfb193cfce574788d9ed6c668872855253a087f04b7a8868b92ae26e52e093a

Observation 23745810-8085-4d9d-a51b-68c60228742a · outbound

This paper cites System card: Claude opus 4 and claude sonnet 4.

HalluScore: Large Language Model Hallucination Question Answering Benchmark System card: Claude opus 4 and claude sonnet 4

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.013805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:b412e3b4f4136815c5bd58efad7f169dcdafeb3c3d8211c9cfa5bd714d0bdc9b

Observation 9f7e693e-a8f3-4b50-b1a6-9efe2974ab9a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

HalluScore: Large Language Model Hallucination Question Answering Benchmark DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-19T20:32:45.410606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:004038fa364cff9d142426e52b176edeb1485eda505b22848d447a48261ab52a

Observation 833ce9a3-d8ec-4e64-af6d-0e093ef7b48d · outbound

This paper cites Openai o3 and o4-mini system card.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Openai o3 and o4-mini system card

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.015758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:2edf68d0d21c8972007f496d3608e0e9ad857837b83b92738020a1b043a37d1a

Observation 5ca92cf6-8b2b-4d57-9b75-5b8505d8f55e · outbound

This paper cites The double-edged sword of anthro- pomorphism in llms.

HalluScore: Large Language Model Hallucination Question Answering Benchmark The double-edged sword of anthro- pomorphism in llms

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.022652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:44fb029f508ab74ae04d51d09c7ed67dbec5f583c7d4aa3ba1a9779fb6564308

Observation eb2336b8-516e-4607-a6d2-847058e4d86a · outbound

This paper cites Breaking the illusion: Revisiting llm anthropomorphism.

HalluScore: Large Language Model Hallucination Question Answering Benchmark Breaking the illusion: Revisiting llm anthropomorphism

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T20:32:46.044418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T20:31:20.017866Z digest=sha256:48f4d018ef44f3d9719902c756ec0db8817b2b84f17832da1805d81e06858b54

Pith citing papers

Observation aa867811-6cb4-460f-a4e5-cced2d308534 · inbound

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering cites this paper.

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering HalluScore: Large Language Model Hallucination Question Answering Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T01:54:05.218315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:54:05.218315Z digest=sha256:a9c597b5174eb0d257cd571424327a9a9315383c605fe000fca6bd23fe9d3060

Observation 376be2f7-2ebf-42f3-a560-69c2d37534d2 · inbound

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification cites this paper.

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification HalluScore: Large Language Model Hallucination Question Answering Benchmark

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T14:47:41.661449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T14:47:41.319887Z digest=sha256:39071a5f0ae94b0c706b5b5e553e718774f2bb8689b3265ecf9d1f6cf6c99c07