Pith. sign in

Paper Citation Record · LEDGER

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.05978.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05978 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T19:49:20.020874Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact6
  • verified fuzzy30
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 49ca3f7d-9d52-4e68-9e85-9533f3f7ada9 · outbound

This paper cites Qwen2.5-VL Technical Report.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.910545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:941136f68e736307ae0f6489603daba7032b448289382513aedd0057c3748512

Observation 65a341be-a14d-4d58-a01b-06044049e8f1 · outbound

This paper cites End-to- end object detection with transformers.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention End-to- end object detection with transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.264918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:1ceb0f472fff80f0476ffa621154cb0c6f1627c07310e61b8d40245149e38738

Observation 3768151d-15c9-49b2-a0f1-695000a92cfe · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.905378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:d679c090aff2164e441a6fa87c0488e4868974dbc0eea402065b39dbb134eba8

Observation aefe5c54-efa5-487d-b89a-7d89ae8a04d2 · outbound

This paper cites BEATs: Audio pre-training with acoustic tok- enizers.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention BEATs: Audio pre-training with acoustic tok- enizers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.273355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:522beef4541dc48b0b33092363d033b85e4366d3f8b03e286d2df6c85ba6bfbe

Observation 8223a1a4-b59f-46e4-b110-a41da26b81d0 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.907904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:e775c8f2089bb7362f634fe1decf77b49e37bd0a6a68487ff6d27e9d7a897a70

Observation ef046333-4b51-45c5-a0d8-e50586390d8a · outbound

This paper cites Lookback lens: De- tecting and mitigating contextual hallucinations in large lan- guage models using only attention maps.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Lookback lens: De- tecting and mitigating contextual hallucinations in large lan- guage models using only attention maps

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.266727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:6661078199c0183658ac129ed9634aad308f9f953157516a5d0eb37491a3ca85

Observation e3d90b48-35f9-4d10-9cd0-de7588d00934 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher R´e.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Fu, Stefano Ermon, Atri Rudra, and Christopher R´e

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.292144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:e2e2f01237f66a65ebad6be0016d415b67639a945c4b61f32b743cf35aca8fda

Observation df80284b-8a5d-4368-a0b1-c228898df0d1 · outbound

This paper cites coco-gemini: Zero-shot COCO detec- tion with Gemini.https://github.com/simedw/ coco-gemini, 2025.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention coco-gemini: Zero-shot COCO detec- tion with Gemini.https://github.com/simedw/ coco-gemini, 2025

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.276645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:a37df55e758d2fcd48dbd43100f63e664cb929d4c9897ed98e7b1e8a630cdb3b

Observation 67814fe2-40d1-45b6-8552-1652a7985482 · outbound

This paper cites Multi-modal hallucination control by visual information grounding.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Multi-modal hallucination control by visual information grounding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.310543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:76d537b1c5520d61e71f3080b30b465e0d58396dd273fa8ef8d604d57dc3fadc

Observation eb0b4114-781a-463c-a197-5e1d513d4e63 · outbound

This paper cites TALL: Temporal activity localization via language query.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention TALL: Temporal activity localization via language query

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.280076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:21c2c0ecc44aeb1ecb6fdcc70ce7fd44515eb770e6e20333919e6700045965a5

Observation 52dcb6a6-ebf9-4184-aaed-7871279815dc · outbound

This paper cites Gemma 3 Technical Report.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Gemma 3 Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.902913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:df8cf350e1d10224afd84467010e579235a7f0526e6255a3350ea2f8319a0248

Observation 92ea6d40-80d9-4b7e-9eb1-b033106f1a52 · outbound

This paper cites Gemmeke, Daniel P.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Gemmeke, Daniel P

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.308839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:e4fab5b787d14ca30f54d5739ff773a975d859bef726d08a360322cc6e313a4a

Observation 3f087581-5dfa-440e-a7ca-ef0f46246b8d · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.901819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:7413d6c117cde0989b56a3805674ba0932fff212ba37a21dcc15ddb4abd472f1

Observation c0cc97b4-24f1-4ba0-ad02-2566d353699b · outbound

This paper cites DAMRO: Dive into the attention mechanism of LVLM to re- duce object hallucination.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention DAMRO: Dive into the attention mechanism of LVLM to re- duce object hallucination

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.286983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:c097fa450b5939b825de04f9cdc29ad538b524efedeacae2ad39fd222f9ab7d6

Observation be4929bb-018b-4874-93be-351e5d8346ac · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in visual question answering.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Making the V in VQA matter: El- evating the role of image understanding in visual question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.317666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:9ca8580a333c57e77792c706b7a72c27d60cad25134e63f2a45e50d873db5eef

Observation a4eff65d-50b6-47a3-b525-57bcfdd5b2ad · outbound

This paper cites an unresolved cited work.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-07-08T20:45:37.305482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:97da4952fbaafc68eb548e3d7a50c6edfba1a5551eab49b4b7adbf75e792acdd

Observation 6b426b24-5376-44d6-bb0d-fdd9f36d3f9a · outbound

This paper cites OPERA: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention OPERA: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.283399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:ceed36ee4f72f7d2a467b5dec998750bf4585ea98d9324f0f8386a340c69ef8e

Observation 54ddb34c-2b88-45bb-9f4d-3ee8ef27d513 · outbound

This paper cites Interpreting and editing vision-language rep- resentations to mitigate hallucinations.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Interpreting and editing vision-language rep- resentations to mitigate hallucinations

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.290443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:374a676048d36ec51db02d938697c2c45327cb15e459e5acb500d22b41a17d24

Observation 324b8e7c-c3b1-4a52-ad62-526928ef1beb · outbound

This paper cites Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Devils in middle layers of large vision- language models: Interpreting, detecting and mitigating ob- ject hallucinations via attention lens

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.285189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:2560d9b76490bc2f70487b73833a30af1145775ffc94843a5973f77e79c66540

Observation 229bcd6c-8178-4afd-8a8c-ce2525cfcc1c · outbound

This paper cites Shamma, Michael S.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Shamma, Michael S

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.307082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:379ab54dd6fe03d861eaebdbb28298d166fdd036f922c88ae8b10d61b079af00

Observation 05d4d1e9-b467-4b78-96f3-624df10f30b7 · outbound

This paper cites Berg, and Mohit Bansal.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Berg, and Mohit Bansal

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.293761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:5c80dd82de42c91cb97e8af249f85560658cb0c6b7b1c75a69f151d113e8b700

Observation 3c7a3251-c031-4d6a-9865-6bf63d160993 · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.268615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:7ed46e87e6ed8d605a7abc195f2b5727feabc34f9d0f4ab2ec48de7e0e7323ce

Observation 62a7be5d-a2b1-4828-9052-a9cf2542b54a · outbound

This paper cites Grounded language-image pre-training.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Grounded language-image pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.300407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:ff8dff4195cd199d635b436444f735517033c4f55838ee81dc1503763009fd00

Observation 8870ead8-958a-473e-af59-edec0dd466d6 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Evaluating object hallucination in large vision-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.281676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:5064b11a66c8590e0584c2d069b4396aab499446169571634e0198236bf4f581

Observation bac235d7-7159-4f89-868b-7e955c3dcbca · outbound

This paper cites Lawrence Zitnick.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Lawrence Zitnick

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.263281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:58c233c71222eb0f765b1cb6680eb73c17692c740ff03de6eb956d58e34992fe

Observation 8dd7a3de-7b9d-4f7f-aba3-ae3fc92ce1ae · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.312378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:4f5a95434b8f1b6fe56944e5495e04328bb0632f2bdb61b0dedd60192da62b76

Observation 535ec38c-64a7-4aeb-b415-ca87ea3989d2 · outbound

This paper cites Paying more atten- tion to image: A training-free method for alleviating halluci- nation in LVLMs.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Paying more atten- tion to image: A training-free method for alleviating halluci- nation in LVLMs

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.261517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:1f0ebbe78e11a11aca095987eaa024aef423d2a2e8657bf91899a497dd09a341

Observation c5c98b0b-11c7-46c1-a9ae-c10bdc58afcf · outbound

This paper cites MMBench: Is your multi-modal model an all-around player? InECCV, 2024.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention MMBench: Is your multi-modal model an all-around player? InECCV, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.314104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:c51df2988651966c4cf588dc41edec6332638cc1f137a07ded9a5586d9f7e292

Observation 142b4287-43b9-4fcf-9dd3-77db617e1d8e · outbound

This paper cites Simple open-vocabulary object detection with vi- sion transformers.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Simple open-vocabulary object detection with vi- sion transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.288680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:cd101c0fc9bcdfa663ae6e8a8ccdf410035076b13bcc43f9a87df056b3cf33bd

Observation 1b3ccc82-5771-4cb9-adb4-dc273502c279 · outbound

This paper cites Query-Dependent Video Represen- tation for Moment Retrieval and Highlight Detection.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Query-Dependent Video Represen- tation for Moment Retrieval and Highlight Detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.297122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:6f87fde6669639cc780f2b37737040254614e97f4478f66455c59e7efe7005ff

Observation cf384d7b-f0d5-46c9-930c-b400c45887cc · outbound

This paper cites Verjans, Phi Le Nguyen, and Vu Minh Hieu Phan.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Verjans, Phi Le Nguyen, and Vu Minh Hieu Phan

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.298811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:4c99edf20b03938c0d4af8bb87d4ca67a2a9275d95beb1ba972b7fb0243c97d4

Observation 281d80a6-1c69-4e76-8a24-9ef93bbb5bf4 · outbound

This paper cites GLSim: Detecting ob- ject hallucinations in LVLMs via global-local similarity.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention GLSim: Detecting ob- ject hallucinations in LVLMs via global-local similarity

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.278345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:dcdab43ba689b9eccb28a8863c7f3f68589a94d0bcc031a86a53bc53cad6e530

Observation 156f1f7c-eabb-4f3c-9bdb-8b84acedf8f2 · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Kosmos-2: Grounding multimodal large language models to the world

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.295387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:2c67a4d862bb7cc14a75870f3b60da5aa1c4d7e2c11a69f6f4684efc6f412f19

Observation ba38589b-a41b-4b48-ba66-9c86bb97a9ba · outbound

This paper cites Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in VLMs

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.303962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:f06063696816149e9e0126decf9fb5ccc4ab9105a71437c042398d3f245325a5

Observation 6e499d5b-46c1-46ad-aad6-93a29f9e9197 · outbound

This paper cites Effective pre- training of audio transformers for sound event detection.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Effective pre- training of audio transformers for sound event detection

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.302250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:df6ce1a65ce7668026ff098e26e5263b87d0b6f8fe4b7c9bce3e5c9266c969ca

Observation 4c1a8bf1-36d2-4fdc-beac-e8c39e32b0f5 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T19:55:33.897014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:eb6156ab5c166829fea34e05b54c560a159e3f441ad0000f068c9083f18e0c5e

Observation c01575d3-9dd5-44a0-8ce1-cc07aa9206cf · outbound

This paper cites Berg, and Tamara L.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention Berg, and Tamara L

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T20:45:37.315775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:cfda0ff1a3d36e11d82a34b1551bec853f1e5d3a58e60209b5ecbf8f8e62457e

Observation a484ad0f-67ce-4a29-a1c8-e12d11fc78e6 · outbound

This paper cites bbox_2d": [x1,y1,x2,y2],.

Propose and Attend: Training-free MLLM Grounding Confidence via Multi-Token Localized Attention bbox_2d": [x1,y1,x2,y2],

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T20:45:37.274934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-08T19:49:20.020874Z digest=sha256:5fa1b9dcc5070b7a799c601d2dd22d9dbe4546fd17e4a3c136a6fd23a1b9b9a1

Pith citing papers

No inbound Pith citation observations are available.