Pith. sign in

Paper Citation Record · LEDGER

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities

As of 6 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:2607.25948.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25948 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:07:44.700601Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved71
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b7af4607-798a-4f1a-a433-761485bddcb2 · outbound

This paper cites Onestory: Coherent multi-shot video generation with adaptive memory.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Onestory: Coherent multi-shot video generation with adaptive memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:42.837630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:42.837630Z digest=sha256:f4334fb606716f2051b37e66bf19cfc8669a209d71fbf52b106ab313eee02fe4

Observation e7da88ea-7b05-4343-bf8c-161a3853c390 · outbound

This paper cites VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:42.950617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:42.950617Z digest=sha256:c82b4b9413a1e30c9a7f944b1c8709733cd6b1ef6a90479acd6da8e25138d7e5

Observation fdb69622-c75c-475e-9456-fb61c02c6594 · outbound

This paper cites F., Mizrahi, D., Garjani, A., Gao, M., Griffiths, D., Hu, J., Dehghan, A., and Zamir, A.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities F., Mizrahi, D., Garjani, A., Gao, M., Griffiths, D., Hu, J., Dehghan, A., and Zamir, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:42.990734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:42.990734Z digest=sha256:2a067f6da81c97dbea6303296b7bdd25145562cf41404a989e06ff69a32e3873

Observation 9a449c5c-e315-458b-ba95-8405a411baad · outbound

This paper cites F., Amirloo, E., El-Nouby, A., Zamir, A., and Dehghan, A.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities F., Amirloo, E., El-Nouby, A., Zamir, A., and Dehghan, A

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.070887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.070887Z digest=sha256:b00c705d292c96a1fcb2b7f16f490b9db6dbf5f5084be0523339de1556f76b7d

Observation d78549ec-17c0-440f-b1b1-5a9dcc1d261e · outbound

This paper cites Qwen technical report.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Qwen technical report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.123780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.123780Z digest=sha256:dc9c8ab84ad610954a9429a35a23e57ccc01f56053b75c0960c5c46e66c379e9

Observation a9dcbb7d-5890-435a-a94f-0d1725df9d7a · outbound

This paper cites an unresolved cited work.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.193803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.193803Z digest=sha256:e3ed7b6c573ca7dee20c923feaec26f96c67653f2c600e5b8377809d59eabca1

Observation 46c8f8a8-a81a-4992-a0e0-2dccdd5ec62d · outbound

This paper cites an unresolved cited work.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.298142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.298142Z digest=sha256:9f24b34ba1e2ff0b807727fc8a487841889c163c29e49d7452163a38667569ab

Observation 57604e31-2c51-4fbb-8dab-01200429c1fd · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.421541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.421541Z digest=sha256:69c36b5fbcd38d131a1450de07289162dc1176e38bb1ae2fc239d5650664ca7b

Observation b869179f-21cf-46ff-9b69-f8984b1ca880 · outbound

This paper cites Hunyuanimage 3.0 technical report.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Hunyuanimage 3.0 technical report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.569475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.569475Z digest=sha256:d93cd0ee5aacb494234339abf17c2ccca9d3456d2acdfb3ad1fa2fe8cae4c07a

Observation fd71cd17-de0f-481a-980b-a7686a66ae88 · outbound

This paper cites Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.710523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.710523Z digest=sha256:ac87ace25ef5769d3244ceac8c5174d7dc5987779c823738ba23d7771aba792d

Observation 57bf5138-9d70-4549-90b1-e5aa2eeebae5 · outbound

This paper cites Janus-pro: Unified multimodal understanding and generation with data and model scaling.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Janus-pro: Unified multimodal understanding and generation with data and model scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.805619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.805619Z digest=sha256:3f3b69d28102dad93391a8c7e49ec588d6d5295057c3e97eea4f25a7dd262c2a

Observation 97cc8852-456d-46ec-aa5e-f57a33517fb1 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:43.940046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:43.940046Z digest=sha256:94708177e54bc72bd9ef31e4c21bce8482a3ba6b8e195f7829e21596f4a8c69f

Observation a9538d68-2aab-468a-bb71-cd26887243eb · outbound

This paper cites Tts-var: A test-time scaling framework for visual auto-regressive generation.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Tts-var: A test-time scaling framework for visual auto-regressive generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.029530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.029530Z digest=sha256:1a164039199a4cc684d0ccb23924854b5603b25097ffdd2eb8b71df40211d967

Observation 484600dc-4e55-429e-b872-7efc03643c56 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Emerging Properties in Unified Multimodal Pretraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.183367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.183367Z digest=sha256:9b2357e48314e0d31d24f2dd1217fa7bee90161818acc2734894569322ed1514

Observation 044a9e8a-552b-4fad-bc18-b6031c72659d · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Scaling rectified flow transformers for high-resolution image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.293487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.293487Z digest=sha256:aad6b1529238bc9945147de5b0c2dc0a24687ba072308ade70ebbe9b449e6f2f

Observation 0c381132-8b6d-4d64-8131-8c96a11d9f61 · outbound

This paper cites Eva-02: A visual representation for neon genesis.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Eva-02: A visual representation for neon genesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.439838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.439838Z digest=sha256:d0a0672107a243deace0be17b727fabc901d772085f710faf4f697f11bc74485

Observation 92f30247-3fd9-45c7-85b0-32791109a20a · outbound

This paper cites Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Geowizard: Unleashing the diffusion priors for 3d geometry estimation from a single image

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.546079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.546079Z digest=sha256:d5a83d4e55f629429d07ddd303e105495c2eed2ef1dec80a910124b7d277b5d0

Observation 4f3cf5b0-126f-4f7e-94ec-eda2d793df96 · outbound

This paper cites Image Generators are Generalist Vision Learners.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Image Generators are Generalist Vision Learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.549021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.549021Z digest=sha256:92b30aadd51423bb955ce3608a0adfcee7b90bedde7d5606c170b73ef880db31

Observation ea97de85-c3c5-4b00-ae5c-ee1c4d4a60ea · outbound

This paper cites (1D) Ordered Tokens Enable Efficient Test-Time Search.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities (1D) Ordered Tokens Enable Efficient Test-Time Search

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.552269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.552269Z digest=sha256:b20eca60c55cff9b69aba5c8120b3f931c69b4cf23775f1e362fd528e0030d42

Observation caa1b8b8-fe2f-4f42-b983-4e16982c3614 · outbound

This paper cites V., Joulin, A., and Misra, I.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities V., Joulin, A., and Misra, I

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.555149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.555149Z digest=sha256:c062098b166db846a09cb3cde5e9869c883b8b10907605780c46967fc4243591

Observation 8a946ee6-ce03-45a4-82a9-e233a074d4d6 · outbound

This paper cites A., Hu, V.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities A., Hu, V

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.557763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.557763Z digest=sha256:7f4b95e79391c6d5d0b418e3ac741d4ae8c7bd0ce756aa72c381e412894b7b7a

Observation ac2acb9a-f852-4aa6-aa30-1db3ad0e1c23 · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.560922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.560922Z digest=sha256:aa47127d5f372eaf65e84322a541bc841afaaea8fc813df705abcb6b65cd747f

Observation e5c93c0c-38a5-4019-9c8a-35288b16bff7 · outbound

This paper cites Infgen: A resolution-agnostic paradigm for scalable image synthesis.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Infgen: A resolution-agnostic paradigm for scalable image synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.563832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.563832Z digest=sha256:79267c86b422ffda6fb16616bdc1ee3380fc410a1a702fcaf2ed96ccbfd3b62b

Observation 1ab17dc3-94ce-4fac-aa6e-824330cb07eb · outbound

This paper cites Lotus: Diffusion-based visual foundation model for high-quality dense prediction.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Lotus: Diffusion-based visual foundation model for high-quality dense prediction

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.566412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.566412Z digest=sha256:130907b9f4158a0729b2efd4c935381d6d7e95dd9af2e6db72d86a92a130894a

Observation 8917cfc2-05c4-44a9-abfb-18fdbb7f99b6 · outbound

This paper cites P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.569312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.569312Z digest=sha256:13be287eae122f315cfaedb1362d559b112fcb1bcca7ee15f699ee72ada91dbb

Observation 11d2c4f3-34ec-4af4-98bf-a12c513862d1 · outbound

This paper cites F., Yeo, T., Atanov, A., and Zamir, A.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities F., Yeo, T., Atanov, A., and Zamir, A

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.571952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.571952Z digest=sha256:3cf8ce59d467777575edfe4f8ad96913ba6a59d2f62b5db1d34f821c2284f167

Observation e63f2dc1-38e4-4f6c-b67c-4f3af3d3374c · outbound

This paper cites C., and Schindler, K.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities C., and Schindler, K

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.574636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.574636Z digest=sha256:2f7d5c90e41c86d5ebf25009fad29f098be7a4e037dd36a73cfd99c0073521e8

Observation e7ff95fa-c581-4ef3-9356-3e37dbab68d8 · outbound

This paper cites Marigold: Affordable adaptation of diffusion-based image generators for image analysis.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Marigold: Affordable adaptation of diffusion-based image generators for image analysis

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.577470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.577470Z digest=sha256:3d811231945dff5433ceec66addeb78220958d11ab5310a8785772d90ee41107

Observation 44986c82-afe9-488e-9028-bbbc567b9340 · outbound

This paper cites C., Lo, W.-Y., Doll\'ar, P., and Girshick, R.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities C., Lo, W.-Y., Doll\'ar, P., and Girshick, R

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.580767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.580767Z digest=sha256:1597dc1a47385e284edca2b6c09c2bd158641ae64a23e0aabba52181d5b21b59

Observation ce79e231-ea6a-453d-806b-d42d9c932483 · outbound

This paper cites H., Pham, T., Lee, S., Clark, C., Kembhavi, A., Mandt, S., Krishna, R., and Lu, J.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities H., Pham, T., Lee, S., Clark, C., Kembhavi, A., Mandt, S., Krishna, R., and Lu, J

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.584083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.584083Z digest=sha256:741d455a81f630bf00c7881393db2078fe29abeddf420406e1c4907aa7eafaad

Observation 979015d4-c47f-4f5e-b32a-066f1f02aec0 · outbound

This paper cites Llava-onevision: Easy visual task transfer.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Llava-onevision: Easy visual task transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.586778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.586778Z digest=sha256:efa9b1b6b9b82b339e8d2f6b245083076bb0de634fef6b9c3aab3b47b4e7261e

Observation cc3256f6-9602-497b-9d6b-2310fe33d191 · outbound

This paper cites Human Motion Instruction Tuning.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Human Motion Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.589373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.589373Z digest=sha256:2fbe3a3e86e9f46782941243d2a6ee134ab6753ac77c24fce5eb10dd58b37ad5

Observation 1bb69985-5ba9-4f44-9e9f-2f72bfe34b86 · outbound

This paper cites Multiple human motion understanding.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Multiple human motion understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.592156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.592156Z digest=sha256:d2fc2c78a8f2113822013b20dbf12113536333c9a0da976d236fae6c77da0983

Observation f6112e8d-96e3-4ab2-912d-0c6945fc47b6 · outbound

This paper cites Exploring plain vision transformer backbones for object detection.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Exploring plain vision transformer backbones for object detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.595148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.595148Z digest=sha256:632fada824c06d27eaddcb70d19ebba88a7bf5f96525e2530745f295a5b410ef

Observation 0ea61d52-01cf-4487-8107-678ab0f43331 · outbound

This paper cites Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.597954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.597954Z digest=sha256:f2c3c95dc46110aaa5834e212c843c735562c85100efa008e22e37446aa378e9

Observation 95db9159-1b90-4bf7-aff2-269d09e674bc · outbound

This paper cites T., Ben-Hamu, H., Nickel, M., and Le, M.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities T., Ben-Hamu, H., Nickel, M., and Le, M

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.600593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.600593Z digest=sha256:4846244e2c2952c527003aa0852b9ac3f9da741f37e7f5ad62d750bfe64b4a29

Observation 060d3eb6-92ab-4c06-8d76-2fa62e8d3cf3 · outbound

This paper cites an unresolved cited work.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.603534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.603534Z digest=sha256:40236992d5e483e6ad7290c94da0e547d0cb9d890677001b86b86a279f28e5c0

Observation a8d033a5-548d-4048-8bc9-3b95b10626ff · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.605996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.605996Z digest=sha256:20a726669ec2abef5188170f9da678d682b4cdb0eb0c93a5b2862fd1a48b06cb

Observation 34e7b16e-aca6-4f35-9f56-7d7be3dce7a3 · outbound

This paper cites S., Yang, S., Wang, Y., Yang, J., and Cheng, M.-M.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities S., Yang, S., Wang, Y., Yang, J., and Cheng, M.-M

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.608739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.608739Z digest=sha256:7291b2bcec35129d185677d7f980d0e0d19d79cf0a07f8b5718a5ea95e3c858e

Observation 8dd8132b-de87-422f-b50e-ab66fdc5afb5 · outbound

This paper cites Unified-io: A unified model for vision, language, and multi-modal tasks.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Unified-io: A unified model for vision, language, and multi-modal tasks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.611365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.611365Z digest=sha256:c0e87f0ca3296eb8c13c25f2c23b47608d1137abbb559e2fa64012bafcea686c

Observation 469f901b-2d17-4e46-beeb-9986f0ad6037 · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.614565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.614565Z digest=sha256:49239571cacd7cf997128bf8646b560335de1618d09e725e202f783460815078

Observation cc13e216-2ea0-4b37-9789-03efe6ccf0b1 · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.617136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.617136Z digest=sha256:fa054667e42933f5a41966f36e8f8f4a9bafaaebfc43bda48c4b507d7a837f8c

Observation eb873ab9-d4f1-4d5f-a445-405f0ad7e565 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.620666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.620666Z digest=sha256:d9b3ac13bf69cf91cdb3463dd64e42a521788ffea80dee1793404f1e967543ed

Observation c1533c3a-54da-4978-9379-c5a0999bf124 · outbound

This paper cites 4m: Massively multimodal masked modeling.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities 4m: Massively multimodal masked modeling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.623368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.623368Z digest=sha256:f0f65a904bc592b041b947ca084d0831fea0fed82384151666dbdbfc0784098c

Observation 9647dec3-8c43-4dfe-acd5-58ce86a33b38 · outbound

This paper cites C., Scalia, G., and Eraslan, G.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities C., Scalia, G., and Eraslan, G

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.626025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.626025Z digest=sha256:af49e550880bb1860e91881814bee648fe2a680e5d422c5129c5119084ec568f

Observation bee50578-0606-4328-b77e-05eb79acfaf9 · outbound

This paper cites Dinov2: Learning robust visual features without supervision.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Dinov2: Learning robust visual features without supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.629111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.629111Z digest=sha256:7fb21331a5d41658cd7f8861fff7ebacea8bdc678e8f270e91179abf0d314026

Observation 0794a551-45f8-4606-a1f8-af2919677e0f · outbound

This paper cites Aion-1: Omnimodal foundation model for astronomical sciences.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Aion-1: Omnimodal foundation model for astronomical sciences

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.631670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.631670Z digest=sha256:b3d26ae2c8e2831666c2a179168169d5dac761c1d6e3ce6135571d60ee6cd4ab

Observation c9bbe1af-7a38-48ce-8694-0d21755a7e3a · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Kosmos-2: Grounding multimodal large language models to the world

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.634306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.634306Z digest=sha256:6299de37425f73b6d869d813b13385444af47fdee08b68fc9f98f8f5e3680d99

Observation c4b25431-7559-4683-93e9-3f3402e1d031 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.637103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.637103Z digest=sha256:bd2d372350faf21fbd852b750b352fa7a5e0d766cd2320ee153bf348c298084a

Observation e20b6ef7-64f0-4595-bb30-89d2b3923bdc · outbound

This paper cites F., and Zamir, A.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities F., and Zamir, A

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.639843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.639843Z digest=sha256:d9f4e272e7795e65f81094098e60090ea43584dbdcc018d75ce7e560d0853e48

Observation b1647729-333d-4f06-a5b1-be01b9d3a2d3 · outbound

This paper cites M., Xing, E., Yang, M.-H., and Khan, F.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities M., Xing, E., Yang, M.-H., and Khan, F

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.642417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.642417Z digest=sha256:52f9cad5cce5f5a2159ff916bd06f8d142bf646bb42ac5a31e424451440ac5c5

Observation 880ba2f5-b222-4e83-acf9-aee1a6cc14e7 · outbound

This paper cites Grounded sam: Assembling open-world models for diverse visual tasks.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Grounded sam: Assembling open-world models for diverse visual tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.645119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.645119Z digest=sha256:bb9c324cfe1b81d24279e5abb533555d517807183cd7242c84c53e5d75cc91aa

Observation 1478dea4-1425-4e93-88d6-705cc436ecc5 · outbound

This paper cites Prom3e: Probabilistic masked multimodal embedding model for ecology.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Prom3e: Probabilistic masked multimodal embedding model for ecology

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.647782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.647782Z digest=sha256:d453403741d71a13e5c7766e24628483317b8f503d42a80dd86577d1922729ea

Observation ac39a35a-4a55-422e-9216-ad3a863fd25e · outbound

This paper cites A General Framework for Inference-time Scaling and Steering of Diffusion Models.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities A General Framework for Inference-time Scaling and Steering of Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.650363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.650363Z digest=sha256:3502316c37a6c7dc520dd807b45e6082c91513706b5471162f4746df64220af1

Observation 2d7e7bd2-ad2d-4f4a-b135-a3568746643b · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.653249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.653249Z digest=sha256:5b7123a0c5fdcaf3c09392bc88a271ba259badc8df10a102d77781617c852885

Observation 9195b562-99db-4c78-a94f-fed84208ecac · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.656625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.656625Z digest=sha256:4791b6a2e69619234d2a2b4f130d5367757ae58243e15415007faed6826e29c7

Observation 9698ebde-4dc9-42dd-ba28-bea7d9a92224 · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Chameleon: Mixed-modal early-fusion foundation models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.659459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.659459Z digest=sha256:d429cb5111d8508a9fde30f84ac9f1d787912ae4773d08b19a5c5eae563974e7

Observation 6ca3e55d-0779-448e-8486-4fe10a976446 · outbound

This paper cites M., Hauth, A., Millican, K., et al.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities M., Hauth, A., Millican, K., et al

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.662232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.662232Z digest=sha256:488aa21081ae05ffafaab80e4b8de1ebbc4ccfee9fb32ed7caf605440122d267

Observation 29e518d6-f0d9-4c7c-bee8-09f1be1f2fc1 · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.664753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.664753Z digest=sha256:6290ddea0d03ef85076bab6a5ed5b6f19e0437bd4cf47145d6f9b3a6015901d5

Observation fde2313f-77c4-4f3f-a550-f60b3721a6b9 · outbound

This paper cites Llama 2: Open foundation and fine-tuned chat models.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Llama 2: Open foundation and fine-tuned chat models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.667673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.667673Z digest=sha256:4e3e393e6038e47ea2fc0c8ab46a626e184e34228825c29cabb4ed23655af49a

Observation 2657e103-86f4-42af-aa21-92b1bed1d70d · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.670179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.670179Z digest=sha256:6b20bee3646893b9e27649e932c6515450453386cb4deb38f6770f010cc50478

Observation 8de3e3dd-05d3-426c-8048-2cad7032dd68 · outbound

This paper cites Emu3: Next-token prediction is all you need.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Emu3: Next-token prediction is all you need

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.673237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.673237Z digest=sha256:bc450dd80e931325053408d744c864d41d4100860a4b7cc5ebbc0b62ecf66112

Observation 26baf85d-08aa-4ae1-b2c4-84a14ef23af2 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.675843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.675843Z digest=sha256:1d21a44c442ce049deac9445944c3f12f022bdf35e12a6b735d071ca1fdd4848

Observation 923d5874-0203-4783-b7c6-ec6fb541cb08 · outbound

This paper cites Next-gpt: Any-to-any multimodal llm.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Next-gpt: Any-to-any multimodal llm

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.678831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.678831Z digest=sha256:4313525b56ff80a5a3e124b6be3aca990c1ee510f0e147b99133f158e22a234b

Observation 0057a862-fe17-4014-aa9f-b6fbc922aaa5 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.682002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.682002Z digest=sha256:56e67a14300d25e91ef4009c606abf9cf81f040f2fd0726b3d41020fbc5d12ba

Observation 1249a70b-02c1-44e4-89a5-113a54cdf00d · outbound

This paper cites J., Wang, W., Lin, K.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities J., Wang, W., Lin, K

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.685026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.685026Z digest=sha256:40fb6769a5c1bf1bfcb9a179835d91e7681104d57c0a18582529d8ac421fed05

Observation a8e2b309-34a5-40cc-aa63-50de98138b5f · outbound

This paper cites Depth anything v2.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Depth anything v2

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.688020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.688020Z digest=sha256:19c29fef2880624347b9651979e09b659d2210a6e95609a4d4e72f3938de1399

Observation c7270472-a27a-40fc-88d8-d1abc4abd4b7 · outbound

This paper cites Anygpt: Unified multimodal llm with discrete sequence modeling.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Anygpt: Unified multimodal llm with discrete sequence modeling

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.691254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.691254Z digest=sha256:10c5165957e4a92e10936fb4d9ec0006378cc5b1e8ea6d549b453801a0d3d5ba

Observation 26cba2ae-6af0-431b-b1f6-cf3c4e82515b · outbound

This paper cites Inference-time scaling of diffusion models through classical search.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Inference-time scaling of diffusion models through classical search

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.693989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.693989Z digest=sha256:aa5372cad293e0420e11ac8095f34cb19200ef73fb7c83bebbe04e52e1e7b07c

Observation 7e75de20-d6fc-47f2-86bc-d4b90408546b · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.697581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.697581Z digest=sha256:9ba69cf4e25dc3aec7462b4cdfd4678257580e87a05ae2fb34341309264d4120

Observation 6d48867f-a489-449a-9428-abae43d5dd2e · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-01T01:07:44.700601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:07:44.700601Z digest=sha256:dfa7096f0927b77c20140b90aaae485477dede6212aea5e37508fc4940bd61a2

Pith citing papers

No inbound Pith citation observations are available.