Pith. sign in

Paper Citation Record · LEDGER

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

As of 14 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.11167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11167 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:57:29.583850Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact9
  • verified fuzzy12
  • unresolved64
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 565c71e8-44ee-40a7-bc9c-ca5e5cff1448 · outbound

This paper cites Visual Instruction Tuning , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Visual Instruction Tuning , booktitle =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.196786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.196786Z digest=sha256:a83000429381fcb064909e28f32989782b61a8e21ad61de701bacb6e3ab221e4

Observation 885c6a52-3474-490d-95a4-3e54da950b4a · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.205718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.205718Z digest=sha256:468aab97485794c8081b4eeb1d7f94b750979af331de89835b96d39b91df6859

Observation 4b2b103e-19f2-4ef5-8ddd-af34039fd017 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.210383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.210383Z digest=sha256:57c49e70bd1a0c6b88b4ecef050eaba0189ecd84f6c0ed7c5751a129fa47c13a

Observation 0bc2528a-21ed-49d6-afd1-97a4894d6f61 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.214140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.214140Z digest=sha256:d6c6ae49ca57ef620d28f65a6592f4b00fcfa5c58abe421d8349a445a7a09dff

Observation 754d5abb-09e1-49bb-a305-b31610cae958 · outbound

This paper cites Qwen2.5-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.218561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.218561Z digest=sha256:cc65f8d6bc72a5db40708e599892af4a3b10f312ef011d5acf8750b42d8b22b2

Observation a3f5ab58-210a-4f26-be6b-b9e5e7637237 · outbound

This paper cites 2025 , eprint=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.223012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.223012Z digest=sha256:3e1d00a38f2d9c6a35d17833fb1a2198661c9d3e52bb0650e571645ed26d1a3a

Observation ec6fb014-cc3f-411c-97ab-3317bc25fe51 · outbound

This paper cites MiMo-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MiMo-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.230150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.230150Z digest=sha256:273e951f3e331c8afccc68f5e683eb298a0046705e5b84bf651f497693bb5c97

Observation c498bd22-1e36-48e9-8295-3b69342f6ca9 · outbound

This paper cites 2025 , eprint=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.234182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.234182Z digest=sha256:8b299bce204caae14e2a341cdf0ec1c7eec4e383731f895537b223c5e75cf552

Observation 56cd082d-a781-48ff-a7b9-25ea335c561a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.237647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.237647Z digest=sha256:0ea51f10dc5ff15afce2497bd8b26d42517751a0ca7f19b2e3c4646dc8cc9cd3

Observation 9700f8ad-d9d6-4c9c-9ab2-e256ead06748 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.241637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.241637Z digest=sha256:cf8d681725fb1468daa02a7f39950fbd0c5993da9462f01f52a06c106bd78f5f

Observation a98493a4-1f8a-4c78-8c33-d5e957b981c8 · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Scalable Vision Language Model Training via High Quality Data Curation , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.245635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.245635Z digest=sha256:e03357a08fbab50f083fc90ac24782604d00e5c6a95759fdc7e42dcc994f302c

Observation f4742661-3eac-49a7-96af-d507aed679e0 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.249132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.249132Z digest=sha256:9e93f20221b6618a174510bf029bde998e8e7bd426faa17212546489f1d61f8b

Observation 5e997eb9-1d21-4376-aba1-c8cb9b751f88 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.253141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.253141Z digest=sha256:2cae4cffea2255e1ce1cb3e5776b79ec23792341f281f1d903c0a8ce062457f1

Observation a1ca3212-96f5-4120-8be9-32b9fd38b029 · outbound

This paper cites Seed1.5-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.257406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.257406Z digest=sha256:dee9275ed866fa5e4e03a44baf4f196bd99fd05ec3187df9e973d29769f6a88f

Observation 606d72bd-1521-41af-86f8-ed8e16e64fee · outbound

This paper cites CoRR , volume =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoRR , volume =

Reference 16

Resolution
verified exact
doi, observed 2026-08-12T04:57:30.180987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.261748Z digest=sha256:75d95d4852be129ee08a5aba328e1114d23d6b9a66b90877f73953d18830c08d

Observation ea48c056-da6f-48a1-ab50-01bb588aa137 · outbound

This paper cites Computer Vision -.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Computer Vision -

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.265494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.265494Z digest=sha256:ba6057bf6d7a02930b35573176070c949fbc25b775dbadbb245c13a419509d57

Observation a47e8c63-bfea-4588-8ce7-f1fe15e5d248 · outbound

This paper cites Qwen2.5 Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2.5 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.269711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.269711Z digest=sha256:b077a996bf6c6b3ed53f0730c4ea52bf83a8a57e088650b0aef27f75f96ffce1

Observation de5b4bd7-3978-4b42-83fc-81c5c01d00f2 · outbound

This paper cites Qwen3 Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.273623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.273623Z digest=sha256:6338f46e2288c4f3e14e76358adf39f57f46a12151a106916f1259e737a3cc4d

Observation 8679e8f7-be7a-4ec3-9f82-ee1df8f835b9 · outbound

This paper cites The Llama 3 Herd of Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.277639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.277639Z digest=sha256:32100604c9f4fe6f844953d02e4bfc0551b14f9ad288ae661894b94c0f24269a

Observation cd8b26b5-f973-4828-8103-9c4e6686aec1 · outbound

This paper cites Making the.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Making the

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.281596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.281596Z digest=sha256:22660f212e0c7765cc7afe54e173f8f3d7d81eff793ad591980f3748659b7969

Observation 7843997f-c70e-4364-835e-31a78e905f08 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.285486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.285486Z digest=sha256:a9a2a113fb69c0937a346efe7a039926eca7f0fd24f06e1490a7bc0e3a54b246

Observation 8fb8e444-3744-443c-9fa4-582fcb27a03e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.289625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.289625Z digest=sha256:2ec277982eba2c83da7ffae60fa49e20be1030f16ba1b5decb7c30feb92381d5

Observation f478145a-65b0-4b2f-92cb-dec2857d45c9 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.294805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.294805Z digest=sha256:a94eb745ceb457da30433706dd0bea02e51173d935e13c55d9cfdd7524c543a7

Observation cc75e00e-fa7a-443b-9502-7c12dd2e591b · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.299431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.299431Z digest=sha256:6b5ad939a37a97585fd9c7342373204b927d72be237abe386365ff87ace94666

Observation b6f4399b-c5c1-438f-832f-2d47f8d26dc6 · outbound

This paper cites Forty-first International Conference on Machine Learning,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Forty-first International Conference on Machine Learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.301283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.307220Z digest=sha256:22c402837ea65eccbbe17d092fbfe080d1139f8dc265de27ebc74b187ae37386

Observation 92af91a1-ab07-4a2f-bbd3-d324f400ce42 · outbound

This paper cites Hudson and Christopher D.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hudson and Christopher D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.311156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.311156Z digest=sha256:fc06fbbbb1f29565f6e9b469e30c8e7f65169527dffe77930829bcf98e5b71e0

Observation b2889eba-9548-4e16-b9c6-e9ec1d9db9fb · outbound

This paper cites A Diagram is Worth a Dozen Images , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment A Diagram is Worth a Dozen Images , booktitle =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.314674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.314674Z digest=sha256:6658e02fde65db6da1f3dc947392ad7679ee604ad0d542206eaae6e6ab710010

Observation 5d21991e-be69-43e5-b0e1-c1de80f27ee0 · outbound

This paper cites Joty and Enamul Hoque , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Joty and Enamul Hoque , editor =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.318875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.318875Z digest=sha256:4cd406207549a985c1370ef8fb1236c7d940cbd3d059cfad9ec537b6528ff836

Observation 966b00db-cd54-4d65-85b5-0de2840a07ed · outbound

This paper cites Cambrian-1:.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Cambrian-1:

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.289350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.322900Z digest=sha256:8ab1534493d808e3b3baa16e6e8ff3e92954fbe8dcc6de1780b0f558becfffbd

Observation e8eb2ba9-63e7-4ac4-aa06-97e6ee4e2dfd · outbound

This paper cites OCRBench: on the hidden mystery of.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment OCRBench: on the hidden mystery of

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.327016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.327016Z digest=sha256:68f57da7e5cc7c763178d19a66de5fb810c6d7b89c304f5e1ef1dd53f219a1ad

Observation 0bb53f02-c088-4c53-9237-85ab45e99862 · outbound

This paper cites 2019 , url =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2019 , url =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.330804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.330804Z digest=sha256:822ce053f60c5e282150f87f5bd1dc2a3ba4fbf4e5f7082761a8ff31e0e37686

Observation 8d61bc53-0d30-4d5b-8979-cb784e26e15d · outbound

This paper cites Berg , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Berg , editor =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.337861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.337861Z digest=sha256:5528ef7207f38094b74c06fcc791564a878a11900ef20ec67cbf472b9d129396

Observation 438c1bf1-98ec-4d3d-8c16-65f9f746a306 · outbound

This paper cites Yuille and Kevin Murphy , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Yuille and Kevin Murphy , title =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.341502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.341502Z digest=sha256:3e53a02c64e90c45326ccac62fb11f1926bfe1df0f94e40097faec08be86ceec

Observation e74cd858-eb1d-44c7-9060-f68719cf461e · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.345672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.345672Z digest=sha256:981ca05b3d89206002baadcabe51ba5ffacf2b457e670aa9952f51bfd9c01c90

Observation 205dfa8c-497f-424e-a738-adee591aca4e · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.349323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.349323Z digest=sha256:e0bcabdcbe2919d6726453a0b7068368999ed0bbff6fbd1163a0a978fce5c6f2

Observation 45e93531-70cd-4e47-bbed-9e0b49aca790 · outbound

This paper cites ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.352991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.352991Z digest=sha256:541c3a22c0155113e66f79c5612d203c952d48e18ab6d6f3ce5f0eed6c184310

Observation 3cec23b2-d15e-479f-8f48-d39036e9083a · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.356681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.356681Z digest=sha256:2fbc2106120dfa5a473f759cfe3667f58b2c75d095cea32ce7efce99274d7896

Observation cef80f05-23ca-4fd2-b9f9-f0cdad2e25d2 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.277887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.360560Z digest=sha256:7fc398c66f46edab0010439a66edc18301ff22253619ae09977c70c7a85a35f9

Observation 89047adb-f0dd-4db6-8626-d5c982287beb · outbound

This paper cites DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.366022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.366022Z digest=sha256:704a99e44321228061186b976d93eeeba36a74063122157e985bc87f782185a6

Observation 21da1ec2-b606-4317-af1b-4c006199bc6b · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.370959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.370959Z digest=sha256:76f212ff0ae4c73ebc3e1cf4f03971a648dd9d7b734886cb17981d76aa771fca

Observation 4e2ded75-9897-483d-bdaa-3244b3cc8133 · outbound

This paper cites Computer Vision -.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Computer Vision -

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.374937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.374937Z digest=sha256:177312728daf42a05314a5f3987ced4de9a0f3109e35488e3b4d4a9921d59ca3

Observation 59fce482-b53e-4896-a479-f67ead3f281a · outbound

This paper cites CoRR , volume =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoRR , volume =

Reference 45

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.921963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.380107Z digest=sha256:bbcaccba9567d1e542f47f0000893c628817ebd8542591e6515ef86cf946be75

Observation b417e9be-da9b-45fe-ade6-3cc9286b023f · outbound

This paper cites Microsoft.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Microsoft

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.385103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.385103Z digest=sha256:4295412f26102561a3f2b5176b120318c73d980bd57622f6882df4aef07a6309

Observation c5ddacd0-7702-4c31-95f8-6ca162d3a25d · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.389041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.389041Z digest=sha256:5f9d8dac161dadaec781bfa20bf420ef1b75e6852091ef6924e4eea17cb28eb8

Observation ce3f8042-9f48-4740-8fa3-884b0755b6c6 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.396828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.396828Z digest=sha256:bf765c4f4ae44f82bf772a5b7890d2af411089ebfb1a8d7575482b63da8e41f5

Observation 9d10c119-6fdb-454c-b553-625a839350ea · outbound

This paper cites 2024 , url =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2024 , url =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.402678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.402678Z digest=sha256:faf8e0f850aadd3b81b5cf0fe08f9e5c1d7da690a169bede57745db8a8bfffc3

Observation abbab471-c565-464f-aea3-0451272a2df6 · outbound

This paper cites Segment Anything , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Segment Anything , booktitle =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.406419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.406419Z digest=sha256:5618cca6e0c8431386df8893746c9a5a2504cde00efc6cb2ebc23e89e902e1fc

Observation bb5db8b4-3b52-4168-a186-825279c19c33 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-12T04:57:29.410596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.410596Z digest=sha256:2ee3c74db8e5eb7ea05523da2d072ef7860f7c345275ed98b8765fa2aba10107

Observation 70fb8413-9ed0-4208-8925-81225e0307ff · outbound

This paper cites 2019 International Conference on Document Analysis and Recognition,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2019 International Conference on Document Analysis and Recognition,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.414823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.414823Z digest=sha256:72260483a20e9c7aebb89cc24935f9d6d05736031e02425f97ec9786e69748da

Observation 0be0bab2-6ab8-41c5-b7ff-f0acd559e871 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.418656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.418656Z digest=sha256:c9f23d5295419d3334c230bd8f237e67747ce3a4d7f9719e791ba2815c809f2d

Observation 99e3d290-888c-4409-8749-adef6a03714e · outbound

This paper cites Price and Scott Cohen and Christopher Kanan , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Price and Scott Cohen and Christopher Kanan , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.422714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.422714Z digest=sha256:f951adaf2612a6b2894ba036730006121827e0e7e90b3b8298550e0a20bbe880

Observation e793620b-01d4-4fc2-8d3d-0b07548dbd58 · outbound

This paper cites OCR-Free Document Understanding Transformer , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment OCR-Free Document Understanding Transformer , booktitle =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.427268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.427268Z digest=sha256:d1531488ac7f48543d215368af997dd9f8f86141cd094f66464462a76b50bb7c

Observation d7e7e17a-16cd-40c2-9c9c-6fce36402e1b · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:57:31.266267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.431339Z digest=sha256:e359b30da3730d38a9b360b336e8d7664ef80519ec369de792da24de13bde6dd

Observation 2ff79465-f090-4877-939c-03b535a9919f · outbound

This paper cites Grounding.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Grounding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.435187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.435187Z digest=sha256:63a9a2cc50473a433e9ed47ea261f379c94bb5f132c69f3183d8be732869aff4

Observation d872d31f-f3b5-4831-a098-3ecb84aa09a2 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The Thirteenth International Conference on Learning Representations,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.438815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.438815Z digest=sha256:18c25a432f358f310b9fabbd5f3704cfbb4bcbc077639b84818937c0a84be1f5

Observation 5f8c0848-552d-4c53-b33b-54ac20f4bec2 · outbound

This paper cites Forty-first International Conference on Machine Learning,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Forty-first International Conference on Machine Learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.246523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.446143Z digest=sha256:dcc7d2fd9dd3e74ac15aaf43e815acdafc62d5c9cda08a9bc4ae9c889f7204d8

Observation bdaf6862-6507-4e7c-9463-79024530986e · outbound

This paper cites Hinton , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hinton , editor =

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.234332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.450985Z digest=sha256:f8ddefa2d80efa7c70c908e1f715c76daea2fbbc4b944a65c6913e65949b379d

Observation 9e0ab83a-63a3-4426-b20c-d62751cdb8e8 · outbound

This paper cites SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s

Reference 61

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.788892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.456006Z digest=sha256:4b5f134ebdd84835d416c02fd52c6ff0cbebe534824470107fd14893630e9607

Observation d6ac425d-f669-4b50-ae94-43c4517bc5f7 · outbound

This paper cites Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.460876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.460876Z digest=sha256:d791986c89ae14f0a5962b675b421c63598089242baf091e9934ac2464381a98

Observation 13a8c10a-4f3b-4043-88ba-1996568f87b1 · outbound

This paper cites GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =

Reference 64

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.762427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.469014Z digest=sha256:e3928abb27abfaaea152996ab70a4ef4191083dc8fe1a3dd2893944005523ec4

Observation e4d0ef44-8a33-4824-a5f3-a1f312028ad2 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 65

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.750175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.473156Z digest=sha256:5066947afd9d73c216340f499097bc928859b50556655fcaba6ebe2b23080ded

Observation d34a3720-006b-476d-9616-28a8ab2b83c5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.477504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.477504Z digest=sha256:86961cd82b686b008762937357026f63dbcc72d2399a4989e5929767bb54db29

Observation 771488ca-ee58-44b5-878e-c446e0a4dfe8 · outbound

This paper cites ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =

Reference 68

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.726467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.486341Z digest=sha256:e15b6da6839ab606a4d21f584b8af185d475c88a4a6fb60a83a216946a3a96b9

Observation 73485882-dd2d-4a8e-930b-2214a5f87f1f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.490215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.490215Z digest=sha256:802b74451e0db262070213db821ba8632afd7b4e58b2439bc3e37168b9c31be2

Observation b0b25b03-1f68-4821-9d4a-7fad0d63e5f6 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.222667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.494547Z digest=sha256:5318fd0322078cd8f34a5fcb2e9d9d0510648d8edfa422b9296298062d42052e

Observation 2d1c01d8-6b9d-4668-901f-e007c922e92b · outbound

This paper cites ICCV , year=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ICCV , year=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.210661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.498269Z digest=sha256:6692137d60ffc9425091ab4f8703d5c4ee46ef1d2e2264c469b9589df337088f

Observation 2046d4a3-1d46-4e01-8fb0-c5a24677d9fc · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:57:31.197617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.502369Z digest=sha256:e287d78f30fba27dea1cee53d75f9ab5dd2205b114015a1ecf35fa7e9e03b93c

Observation 9aa41628-5252-4051-b715-16d7d06c68d3 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.505985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.505985Z digest=sha256:f07daa8d3d2503471394575a881abe7580abd0caffd442831d97232bcc76989f

Observation baf50106-d779-424c-acb3-61de245ec292 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Learning Transferable Visual Models From Natural Language Supervision , booktitle =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.509977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.509977Z digest=sha256:c50404f1b9b6bc33e71adc891d3714ae9a66d2e5a0b2385b2845559fff1c655b

Observation 1b9d230d-3562-416b-80aa-8a8b8503e51a · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.513508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.513508Z digest=sha256:b9a20547692b8b972d1b8e160af06c5095e1f9f3ea33fcc1e1fd089883735ade

Observation 2d5521f7-924d-4bba-b023-f15988c62f71 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =

Reference 76

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.677408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.518355Z digest=sha256:c9f63a5446eb7a6158b682391b1b5d36e6348fa9a0b6655dc562fd0a33fc8ce9

Observation 02cc1978-67db-47a9-bbb7-fcad49a87637 · outbound

This paper cites Shaker and Salman H.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Shaker and Salman H

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.522179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.522179Z digest=sha256:86b2d920b59dc234d78c5c7243305b28793e61424381658fcd840b93004ae579

Observation 70b2912e-eb61-455e-8ada-155eaddd0cbc · outbound

This paper cites NExT-Chat: An.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment NExT-Chat: An

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.176905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.525748Z digest=sha256:e7159779be85c800acac70e0978c2b2fe7d2cff0f6bc1b176dc3f38b08cefc5f

Observation e9d00afa-6941-4047-8a5c-b6247aaab4f7 · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.529634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.529634Z digest=sha256:f5a63ee263cc4a04229444429fcc0d1aeca39828c49860ef8110c0ac91d1c22e

Observation d660a6f6-af95-4a67-b46f-6defdd0d40d6 · outbound

This paper cites The All-Seeing Project.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The All-Seeing Project

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.534030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.534030Z digest=sha256:b6ec1c84b16fb923aeb1dda1ce0809e7c74264f79c1fa5f4757f7d1511b23182

Observation 474a49c6-8f61-4d83-9ac5-b84e8e4d9867 · outbound

This paper cites Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.163045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.537581Z digest=sha256:23dd030808fddd8edbcbed4c91457cf49cb639a59c07beab718db2ef5bc6b4bc

Observation e313d1cf-c9e1-4d48-985b-13a69e1b5aac · outbound

This paper cites PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.541551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.541551Z digest=sha256:2cfe656c2712f717718fd868479c8d307aaafed1f9adfdaf9a3fa6cab3dc75b0

Observation 3f70b69b-cdfc-4972-b6fc-dbdce5df0209 · outbound

This paper cites Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.149501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.545608Z digest=sha256:d0e9e12ee8ca39fad6ee7e77ec3866d35663869af6ded15978cc3371db44b91b

Observation cd13adba-e00f-494f-af79-67378847e00b · outbound

This paper cites DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.549323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.549323Z digest=sha256:e839801e6e2f41a5eb17860b92d9ab01eef0d6444b0b40b0d83c558de72f5cbd

Observation d323adac-6712-4058-92e6-6a814d721a39 · outbound

This paper cites CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual

Reference 85

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.642572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.554233Z digest=sha256:7593b3bec31a61fa86a23465fe0b1b245be909e1ecadecea6b33517f68185a65

Observation c4518eaa-3692-4674-9be4-5fb1b2a65984 · outbound

This paper cites Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer

Reference 86

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.628745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.558383Z digest=sha256:38dd522b1adcac1b3a02e7c32ca77fbef805494ce2caa6d9731d03fafef9ab7c

Observation ed807199-313f-400c-b98a-76c76d9c7e41 · outbound

This paper cites Latino Language and Communicative Behavior , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Latino Language and Communicative Behavior , editor =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.135163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.562649Z digest=sha256:e85dde10e9d2b050281b559cf4ce52126bccf6ee86676cae001f366252696dd7

Observation 3626a8b2-33ad-484e-9e91-26d287c368d9 · outbound

This paper cites Thara and Prabaharan Poornachandran , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Thara and Prabaharan Poornachandran , title =

Reference 88

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T04:57:30.364833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.567258Z digest=sha256:cd50776c1d3b9a89b7611fed5acfe0e7f90ba9d6727625be5bbddd7c69bc8e7c

Observation 0f9fbaa1-9520-482c-b3a1-b903594a379a · outbound

This paper cites Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.121851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.571385Z digest=sha256:a6383de1c3483fc6588a3e58d66c8e21e561be8a2a4a285c74ba5cb5a46f4a87

Observation 0fb93f31-adc7-44a4-9fd3-175496eb9a09 · outbound

This paper cites Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.575127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.575127Z digest=sha256:b87af97c33e53264059a8a2307554491e2b23af3576bfe90929adbe6d6c73f46

Observation aa44995f-2e8f-4629-81e1-0efa883a7051 · outbound

This paper cites 2025 , howpublished =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , howpublished =

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.579562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.579562Z digest=sha256:db343bd12e88ab07c0132305de175e8ca165ae39ba3d5400c7de1de58124cec6

Observation e8ec4d31-14bf-4d50-ae36-e862ddc6b724 · outbound

This paper cites 18th International Conference on Pattern Recognition.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 18th International Conference on Pattern Recognition

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.583850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.583850Z digest=sha256:b9e48059656e278f776e69585d2147d22d7509bce6a9cedbbf07ce226c104895

Pith citing papers

No inbound Pith citation observations are available.