Pith. sign in

Paper Citation Record · LEDGER

An Introduction to Vision-Language Modeling

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2405.17247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.17247 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:37:24.093410Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

33
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3b082036-cdca-4917-a35f-22c8a6e019fe · inbound

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning cites this paper.

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning An Introduction to Vision-Language Modeling

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-23T01:52:22.906252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-23T01:51:12.410547Z digest=sha256:0933a782a42a2741691dd037d66b4e2ce7729a32326bfc67c014152b7dcdea6a

Observation 37f0f4cb-86c4-49a2-9f8e-4811e7b9aea8 · inbound

Are Vision-Language Models Ready for Dietary Assessment? Exploring the Next Frontier in AI-Powered Food Image Recognition cites this paper.

Are Vision-Language Models Ready for Dietary Assessment? Exploring the Next Frontier in AI-Powered Food Image Recognition An Introduction to Vision-Language Modeling

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T20:02:02.040008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T20:01:40.341832Z digest=sha256:fcb4d904ac2b766f9fc2a1a4b7b051b9e9ba70f55b440f688af5a34bc31019f4

Observation e477ded6-793c-414a-9b5e-fdc7025c97eb · inbound

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey cites this paper.

Alignment and Safety of Diffusion Models via Reinforcement Learning and Reward Modeling: A Survey An Introduction to Vision-Language Modeling

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T02:30:56.274186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T02:28:23.754561Z digest=sha256:42067bc4be4a1acf6eb5f8acd276b0dfe14d96fcd829e6cecebe4fc4d56f91b6

Observation 9eba8e9f-92ec-4e2a-938e-1ec9ae5e8812 · inbound

Participatory AI: A Scandinavian Approach to Human-Centered AI cites this paper.

Participatory AI: A Scandinavian Approach to Human-Centered AI An Introduction to Vision-Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T16:37:24.093410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:37:24.093410Z digest=sha256:ca4456f054339557f013195fcec04546a720163dfaf6d0d8f9957dfe7ab2699e

Observation afcc685e-3e51-4df1-a6db-8eada6c4fc77 · inbound

Benchmarking and Mitigating Sycophancy in Medical Vision Language Models cites this paper.

Benchmarking and Mitigating Sycophancy in Medical Vision Language Models An Introduction to Vision-Language Modeling

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:06:27.386473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T14:03:41.489520Z digest=sha256:bb422625de4ff597ab36c561814e7bd29a2e6802fe4b9546c77b5625d702d40d

Observation ac7ff740-6789-4f54-abf9-65dd268a643a · inbound

Benchmarking and Mitigating Sycophancy in Medical Vision Language Models cites this paper.

Benchmarking and Mitigating Sycophancy in Medical Vision Language Models An Introduction to Vision-Language Modeling

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T21:30:39.342101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T21:29:52.725437Z digest=sha256:17565e7e759bde82d1f72905dd13c7dec653f36aa1763d44d05c35fee470a546

Observation e592f8e0-fd7f-4609-988e-70e8a3a81a67 · inbound

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering cites this paper.

Spatiotemporal Knowledge Graphs as Persistent Scene Memory for Embodied Question Answering An Introduction to Vision-Language Modeling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T12:55:35.208142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:55:35.208142Z digest=sha256:c0f3e72cf47f716175e5dd9bacdc41a0d99e84458e62c1705934e9c1e464e666

Observation 4f6642ec-0864-46de-bd20-303bad458ba4 · inbound

SemanticOpt: Towards LLM-Based Semantic Black-Box Optimization cites this paper.

SemanticOpt: Towards LLM-Based Semantic Black-Box Optimization An Introduction to Vision-Language Modeling

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:15:34.195019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T20:14:30.658631Z digest=sha256:e2c7627946a032a38c94d13d3adc9217f6c30650bb9adb82e697b429462f4e00

Observation 014b06b0-984d-478c-844f-bb2a82dd5fec · inbound

Dataset Safety in Autonomous Driving: Requirements, Risks, and Assurance cites this paper.

Dataset Safety in Autonomous Driving: Requirements, Risks, and Assurance An Introduction to Vision-Language Modeling

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T23:35:31.620155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T23:32:23.036776Z digest=sha256:32fdd8bea8aa31257e6f0fd6b42801324b4f630297c91ce1c9ce7bfa39b5b462

Observation a53208cc-d842-4f79-aa9d-61583ef1bbc4 · inbound

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications cites this paper.

MM-Telco: Benchmarks and Multimodal Large Language Models for Telecom Applications An Introduction to Vision-Language Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:10:23.017736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:06:30.391838Z digest=sha256:4c2be46120bd72e31e99f42f641e417b5b46bebdb4e33aebb0c17384c57a30f2

Observation 98e4998f-f086-4e26-a339-f68753379d5e · inbound

Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios cites this paper.

Multimodal Benchmark for Safety Assessment in Industrial Inspection Scenarios An Introduction to Vision-Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T07:06:01.235221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:06:01.235221Z digest=sha256:9a120bace5a98aec824f49b1a5b0b1abde92233c7d04c7f362b88da2bc07ee59

Observation ec9c2555-0500-494c-9341-1069322bbfd9 · inbound

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification cites this paper.

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification An Introduction to Vision-Language Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-10T19:20:44.349250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:20:29.951227Z digest=sha256:e391275183a1f3b1b9f09d8c30c7c005807d21e4f4a01aaf798f087e3ee75d92

Observation 0167d732-4c69-4cfa-a2d9-826ec5b2fd21 · inbound

Training-Free Semantic Multi-Object Tracking with Vision-Language Models cites this paper.

Training-Free Semantic Multi-Object Tracking with Vision-Language Models An Introduction to Vision-Language Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.939248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:12:08.080882Z digest=sha256:c9ee4c319500dd617bd409bf4826dadbfbd24a3f0dc0cc36d8085c607915bb43

Observation 609f6644-2f72-43f1-aaf2-e9cd7e91ac6b · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm An Introduction to Vision-Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:51:05.090330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:cff5bb194654c06199b32b22adda36402f24e36240d98785f9a904fc97c1e05e

Observation d795f7b7-fd97-412b-8678-9d9a03b447c6 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm An Introduction to Vision-Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:24.996880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:ebbc45ea48f521f08955e28c4693341394a9e5a418d7a631ec795805989ab7ff

Observation a0208140-b0dc-4ef4-96d6-2bff8399a3d7 · inbound

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation cites this paper.

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation An Introduction to Vision-Language Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:51:07.535999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T14:47:15.946199Z digest=sha256:2a2c78f84adad047ad4dceb2fbde36625e5487219631601b03804680676ce56d

Observation 06721989-116b-487b-9ad6-e4eb181b9e3b · inbound

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation cites this paper.

Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation An Introduction to Vision-Language Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:45:11.956542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T00:41:00.814875Z digest=sha256:1403f477375416a8408dda01648cc44259de040c1a5c3b268f3b96caa0e7ae0b

Observation b1016b3c-194d-4f9a-aefc-441cbc6bac8f · inbound

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid cites this paper.

Structural Ranking of the Cognitive Plausibility of Computational Models of Analogy and Metaphors with the Minimal Cognitive Grid An Introduction to Vision-Language Modeling

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T16:56:06.159239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T14:33:11.033906Z digest=sha256:ab4d455024552d6f187afac6b130fcdd8b307cde01df4193a435b9603c25323d

Observation 67009cba-1eca-428b-b224-5b63e7f14042 · inbound

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs cites this paper.

RouteHijack: Routing-Aware Attack on Mixture-of-Experts LLMs An Introduction to Vision-Language Modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:16.515490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T19:22:00.217729Z digest=sha256:13fd6772120e280fec610652c5072c40f0c27c8a67a43276df0e242e6b1ae66f

Observation 6d8915a2-f183-4658-be7b-56b5b612884a · inbound

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models cites this paper.

Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models An Introduction to Vision-Language Modeling

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:41:30.843174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-07T04:07:58.031100Z digest=sha256:0dd663ba8996808f2c6ab28f77a07dbfb05472ca32e2629ebdc36854f29f21a6

Observation 64829ab6-3694-4ad2-8790-09329dab8a1a · inbound

Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought cites this paper.

Convergence and Emergence of In-Context Reinforcement Learning with Chain of Thought An Introduction to Vision-Language Modeling

Reference 191

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:31:00.589358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:15:41.980346Z digest=sha256:7901868778065d2a1b2f9ed5601bc073175c57d2f9df3da676f5547cdf7c1c7d

Observation 139fe1bd-0a75-4c58-86fb-af0687f16340 · inbound

Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning cites this paper.

Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning An Introduction to Vision-Language Modeling

Reference 220

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:35:58.657864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T01:13:34.836246Z digest=sha256:8d9e6800ccea341fbfa77515091909449c6baea311ce56c3cddcdc15baf958f6

Observation 3ed98cbe-2c1e-469d-8e94-9df909871a60 · inbound

Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning cites this paper.

Beyond Linear Attention: Softmax Transformers Implement In-Context Reinforcement Learning An Introduction to Vision-Language Modeling

Reference 223

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T23:29:13.098615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T23:26:35.072127Z digest=sha256:e1d5205f1d4b13d2208e6b51f2bb7ac552456b407d9268dfe81021413e612f38

Observation 41ef5036-b2f7-4edb-9ccd-dffbd6074067 · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production An Introduction to Vision-Language Modeling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.962691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:59874dfe28011060a35b8055df51fa6f6cbbdc4e6e4b8b7b662000b66d49429f

Observation 345adb82-6668-4a01-b947-59b133cf7e3f · inbound

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study cites this paper.

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study An Introduction to Vision-Language Modeling

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:35:04.165472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T21:34:33.366046Z digest=sha256:1c5f0840dbcfa29c4da04f914c17261834c808b27d142bef5648882016ec8c35

Observation c4ac9317-8b87-410f-85d0-08dd67859fb9 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers An Introduction to Vision-Language Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:38:11.242061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T09:34:45.186929Z digest=sha256:f587d35f312cbd421f7f67ef21f6c50aa0bddaad1f6d835345601218102ae32f

Observation 611cdf40-cefa-4b14-b0fe-b5ac6fbc5af9 · inbound

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers cites this paper.

Symmetry-Compatible Principle for Optimizer Design: Embeddings, LM Heads, SwiGLU MLPs, and MoE Routers An Introduction to Vision-Language Modeling

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.454252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:42:01.854481Z digest=sha256:f3a0d03495f608149dc82b0dba28e712b8b8698e43be97e2bbd35a5059c68cec

Observation 4ce1209c-38ff-4805-bc82-4ec33f26eabe · inbound

OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis cites this paper.

OrganicHAR: Towards Activity Discovery in Organic Settings for Privacy Preserving Sensors Using Efficient Video Analysis An Introduction to Vision-Language Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-20T08:38:08.917795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T08:37:44.181899Z digest=sha256:1148d5bb593424b97cb9073f4502687e05454290c112f00e2d52b8941f3eae22

Observation 2606c9dd-8328-468a-ab13-18bcf3c354c7 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models An Introduction to Vision-Language Modeling

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T04:59:36.551480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:3461efd40410a7eea42b71e78a0f55103f29366d32beff196b54b875826d86d0

Observation 5769037e-ae06-4c74-b7a1-3930efef51f8 · inbound

In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models cites this paper.

In Search of the Ingredients of Open-Endedness: Replicating Picbreeder with Large Vision-Language Models An Introduction to Vision-Language Modeling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T15:09:20.426018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T15:09:20.426018Z digest=sha256:27cc2064f4c8d0749c9f8e84093505e01d634047da8933f67fb6ac0aef6c4385

Observation 868b1f91-27a2-4dce-b9fd-1f3d632df916 · inbound

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization cites this paper.

MGVQ: Synergizing Multi-dimensional Sensitivity-Aware and Gradient-Hessian Fusion for Vector Quantization An Introduction to Vision-Language Modeling

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:48.331211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:36:45.807397Z digest=sha256:d217c78329d22e68a624f4ab1e105a88cab2cadd1c780da293be1580bbcb6be2

Observation 8b35ea78-0da1-4a70-b4bb-8e2abc571a02 · inbound

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering cites this paper.

Reflective Dialogue between Teacher and Solver Agents for Video Question Answering An Introduction to Vision-Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:53:28.533978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T13:51:13.952689Z digest=sha256:5174e9f02cac0d7766908c4abb2d10771f4491104b7b4ace203acc4eb53a7ab8

Observation 241d9055-55a0-43a5-a2ed-e51a8dffe4f4 · inbound

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models cites this paper.

When Meaning Travels: A Granular Lens on Hybrid-MoE's Role in Idiomatic Understanding for Language Models An Introduction to Vision-Language Modeling

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:16.768958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T15:19:26.983760Z digest=sha256:899ad69365fbb1c515a752b4569f5c1ca9082d83fabe244e6762a16333ee90a8

Observation 507bac3a-e05b-4116-b4a8-d487dc2cf2fe · inbound

Do Vision-Language Models See Dwarf Galaxies the Way We Do? cites this paper.

Do Vision-Language Models See Dwarf Galaxies the Way We Do? An Introduction to Vision-Language Modeling

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-27T20:41:14.111928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T20:38:55.736579Z digest=sha256:5cce278e5b433879288df5e07b888a5127390ca0810ebf15a0988118650000ac

Observation 1be23868-cf52-4436-9ba2-f28b4c4a4b5b · inbound

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models cites this paper.

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models An Introduction to Vision-Language Modeling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:37:14.457798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T21:58:53.702009Z digest=sha256:8ab38e376e922da9c0afc2668587df28c39bc527a791c626ee3a21af7bde1a34

Observation 6f66c2dc-4e0f-4493-9bfe-9700d3f1d54c · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving An Introduction to Vision-Language Modeling

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:23:51.147850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T05:08:46.145616Z digest=sha256:cb9bc0955ccaa95ec45aacc95583aab272aeac1c69f74c9c499f9999144432f3

Observation c5efc954-3b99-47c8-b3d0-7f7bef75eb2f · inbound

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving cites this paper.

MVPruner: Dynamic Token Pruning for Accelerating Multi-view Vision-Language Models in Autonomous Driving An Introduction to Vision-Language Modeling

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T21:37:24.221809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T21:35:33.280591Z digest=sha256:4258b91672bfacfae2fd6188f5cc3053712414850871205a7e6275a09b4676c2

Observation 23b8b264-c111-48ef-8bf9-e0996208d565 · inbound

AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation cites this paper.

AC3S: Adaptive Conditioning for 3D-Aware Synthetic Data Generation An Introduction to Vision-Language Modeling

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.416228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-01T06:12:59.736399Z digest=sha256:a2993d1a6564c37047be152a9fc4fdd4c69e9b068461040b595b8a460c8e88a4

Observation 32fb74ae-2b89-4919-99a6-1c50e2db20af · inbound

Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability cites this paper.

Telescope: Improving Zero Shot Detection of LLM Generated Content By Measuring Token Repetition Probability An Introduction to Vision-Language Modeling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-11T21:58:20.755128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T21:58:20.755128Z digest=sha256:45c63877b64187ddb40459a7c2ea1f4c6f3f963c438db5281f94cbaf8cc53cb8

Observation 5e7f0901-94d9-4cb5-aa69-c894d17338f2 · inbound

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications cites this paper.

RoboVista: Evaluating Vision Language Models for Diverse Robot Applications An Introduction to Vision-Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-11T16:28:31.916881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T16:28:31.916881Z digest=sha256:5a65065172401e618727200f1697d424ca44ecc913e137e1b43b3d46e31c3b7f

Observation 911ec530-27fe-4e8b-b0d2-036f862cecbe · inbound

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks cites this paper.

GaP: A Graph-as-Policy Multi-Agent Self-Learning Harness For Variational Automation Tasks An Introduction to Vision-Language Modeling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T07:12:02.239732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T07:12:02.239732Z digest=sha256:7772d00ea20b535f5cda2b2bba6484ead157942c5a920ef8dcc6d17bae5cb985