Pith. sign in

Paper Citation Record · LEDGER

Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

As of 12 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2312.16602.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.16602 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:08.385867Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T17:41:53.208887Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ab626ba4-9043-4e99-b595-da6c85bd7601 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:08.983223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:08.983223Z digest=sha256:bad977ad08c2d31c0e6615199615f86694bfd2429184b6782efc737cc9975293

Observation 76c474ea-bf81-4d73-8188-d4bcdcaed8fa · inbound

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces cites this paper.

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:08.385867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:08.385867Z digest=sha256:c9dfe39307baa39eaa3585eaaa08492c6aed3f051ed0751618c1fece361d8874

Observation c881b9e8-5aee-47a1-bd0e-54d4f3633b19 · inbound

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model cites this paper.

PAINT: Paying Attention to INformed Tokens to Mitigate Hallucination in Large Vision-Language Model Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:28:42.001774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:28:42.001774Z digest=sha256:d04e6376373956bb545225018771967d0aca48e7412b7b0ce95fdfa1d3ef5345

Observation 258c1744-55b4-4d36-945f-0568c0653ae1 · inbound

Visual Compositional Tuning cites this paper.

Visual Compositional Tuning Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T17:41:53.212540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-22T17:39:09.890605Z digest=sha256:51ed8f037759ece35ae19688cbf40540e245c37c9909478c02ede9bae48ec010

Observation 211c8548-f824-4440-b349-4ea4f90927ae · inbound

Towards Unconstrained Human-Object Interaction cites this paper.

Towards Unconstrained Human-Object Interaction Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:20:30.381115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T14:17:06.952676Z digest=sha256:b7a9d244262cdf021a7ad7cd37de4c69c02b6c4cb9fd373a935f29d0aa3df2b9

Observation e9fb204d-9419-4e09-8c04-cceeb7b8bc0c · inbound

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models cites this paper.

ASRU: Activation Steering Meets Reinforcement Unlearning for Multimodal Large Language Models Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:40.845800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T19:22:43.082083Z digest=sha256:61aebc19001b451236db35cca1bf89328fa34e4738859b5d7212ef6e8cf5efe5

Observation 9319d652-9a93-454d-abae-1b934d7a731e · inbound

Learning in Deep Networks under Dale's Constraint cites this paper.

Learning in Deep Networks under Dale's Constraint Visual Instruction Tuning towards General-Purpose Multimodal Model: A Survey

Reference 228

Resolution
unresolved
no resolver link, observed 2026-08-10T17:39:43.306915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:39:43.306915Z digest=sha256:619af3ae8acbdc7d26401215e1d587a0e2cfed829ffa08044c24d2313e8d6b7b