Pith. sign in

Paper Citation Record · LEDGER

SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2408.11813.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.11813 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:52:07.914884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:56.838355Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2fc3c403-207d-4698-97c9-7b4240bb6994 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:52:01.940574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:b1122e59d6ee0274586cedfa8a014b6de7fab51cfef9200973bad73d9476415b

Observation 15f7dcb8-de70-4ef8-a31d-a938af74c5c5 · inbound

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models cites this paper.

Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:07.914884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:52:07.914884Z digest=sha256:aafdebe22d4432b95e220baf886e8bfa99cd385c446757af1e30a810332d974a

Observation 002e6126-6d0f-4cb9-9fd6-a1e4c5f649ed · inbound

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL cites this paper.

FocusDiff: Advancing Fine-Grained Text-Image Alignment for Autoregressive Visual Generation through RL SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:23:12.573432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:23:12.573432Z digest=sha256:5c7ffe7f98ea9b4708dbd296877371815a30cc07722efeb0194657f280b73e5a

Observation 72d1f194-5539-46e4-8128-9a778d8161c4 · inbound

Towards an Explainable Comparison and Alignment of Feature Embeddings cites this paper.

Towards an Explainable Comparison and Alignment of Feature Embeddings SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:39.604539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:05:39.604539Z digest=sha256:6391c1a89621c5ec41c40af4a8daf18feeed961b860b889d06a3aeeaf6dc8465

Observation 1963e07b-c00d-4d04-a740-fe8f8324cb59 · inbound

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning cites this paper.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMs

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.839772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:52:44.785582Z digest=sha256:52f06dd4c6d079b9f263dd2953fc34e550f7df3c153dbf7fa4981779f7825db6