Pith. sign in

Paper Citation Record · LEDGER

MouSi: Poly-Visual-Expert Vision-Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2401.17221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.17221 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:37:54.726053Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.364731Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a6e169af-b8f8-4e5e-bdd5-e05a896e9cc3 · inbound

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning cites this paper.

From Holistic to Localized: Local Enhanced Adapters for Efficient Visual Instruction Fine-Tuning MouSi: Poly-Visual-Expert Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T17:37:54.726053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:37:54.726053Z digest=sha256:d7ce6403e9d27c838dfbe6757f4aec00ed8353f04c93ba30d9e09c0fef8aa56e

Observation 65923e19-8764-488a-8f45-b407ced021d8 · inbound

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion cites this paper.

Florence-VL: Enhancing Vision-Language Models with Generative Vision Encoder and Depth-Breadth Fusion MouSi: Poly-Visual-Expert Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:27:52.421540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:27:52.421540Z digest=sha256:26b7cdb83aec5f141b5d4549a40ba0afe4831f9d3d211331e6eda4fddf06123c

Observation 37462020-3f3a-46e7-8157-55b8d09095c0 · inbound

Hidden in plain sight: VLMs overlook their visual representations cites this paper.

Hidden in plain sight: VLMs overlook their visual representations MouSi: Poly-Visual-Expert Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.884425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.884425Z digest=sha256:ff80ddb990a145dbcb905d695805312516b91956332a2fdb609bc7da2fc2e111

Observation 7fac3077-35de-432e-899d-6aaaa01bad56 · inbound

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models cites this paper.

METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models MouSi: Poly-Visual-Expert Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:18:01.081840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:18:01.081840Z digest=sha256:8adf61178c24433f44b6eb27dcf31d56eb4eddb90f783d2febada00cdf8aacc8

Observation 6b37c4e7-c371-4897-80c7-549f47e882db · inbound

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence cites this paper.

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence MouSi: Poly-Visual-Expert Vision-Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:06:35.104117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-18T16:04:11.834915Z digest=sha256:7c14539de7f64d110bbb8a9bf66c6276d3f265154a1a78cd5c97d10ef74c8d55

Observation c5761389-2ece-4c18-9d6b-8e9891c02993 · inbound

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence cites this paper.

EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence MouSi: Poly-Visual-Expert Vision-Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T16:15:50.670967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:15:50.670967Z digest=sha256:212e80453afc303672bca07e15c56c4a718e12cabfe5a741ad0ea337b48babe5

Observation f27ef241-a76b-4659-9e1d-b2ea56468237 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models MouSi: Poly-Visual-Expert Vision-Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.081484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:a3096186e7a5fbd249a6874d9d20c4cdce04d6c1cf3cd377ccac7a04579e62c9

Observation 0258be39-2044-4ee6-b9ea-649ffc62b469 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models MouSi: Poly-Visual-Expert Vision-Language Models

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.366724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:d6b1cf91c27e1e1b20ffa0dc7c7cf3437e56588d899152ea4ac1b10c44e7e9e3

Observation 5a5a3714-2f73-4188-8e8b-e4171f14d540 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception MouSi: Poly-Visual-Expert Vision-Language Models

Reference 146

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:d1d8364017d85638bd4fbcee35cdfb84d65ca5b02172c890feaa1409a40e5e76