Pith. sign in

Paper Citation Record · LEDGER

NeoBabel: A Multilingual Open Tower for Visual Generation

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.06137.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06137 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:30.256170Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact4
  • verified fuzzy1
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb33f64b-e506-4043-9aba-00eb31d940d3 · outbound

This paper cites Maya: An Instruction Finetuned Multilingual Multimodal Model.

NeoBabel: A Multilingual Open Tower for Visual Generation Maya: An Instruction Finetuned Multilingual Multimodal Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.391619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.391619Z digest=sha256:c7f600350970076966dcfc15c2a8d16b9a39723b31866ab9c5d514410589c95e

Observation 631a2c31-c178-42c3-bb22-427e5ee31236 · outbound

This paper cites Qwen2.5-VL Technical Report.

NeoBabel: A Multilingual Open Tower for Visual Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.702878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.702878Z digest=sha256:474bf5a38f839972392262584c4d69ed0db5beca7628a51af4681690fb439062

Observation 9946d252-2a16-4bd1-a9e0-50e93c74d681 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

NeoBabel: A Multilingual Open Tower for Visual Generation Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.468458Z digest=sha256:238fda240a76bb18cf147bf8a6d66b9f68b0ea183d45d6dd77ab935abb7021c5

Observation b134bd5a-39fd-43cc-9b02-961aba8666fb · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

NeoBabel: A Multilingual Open Tower for Visual Generation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.683136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.683136Z digest=sha256:d7dc7f90c65cc074ecdc86d1b8fdc979579f7f5145856b256b75130a0704daf8

Observation 254445f6-6f51-42b2-a357-beb673a5d945 · outbound

This paper cites Unified Autoregressive Visual Generation and Understanding with Continuous Tokens.

NeoBabel: A Multilingual Open Tower for Visual Generation Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.807532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.807532Z digest=sha256:dffcf2de2d91cf0a205c207333b69ae3c164c808adf111c7d47beeed49e3df63

Observation d72cfab3-18e6-48e9-91b5-a235dabc64fd · outbound

This paper cites Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You.

NeoBabel: A Multilingual Open Tower for Visual Generation Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.989066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.989066Z digest=sha256:fc3b5749e072f9fce503ddba25a904b5018078fb48b5954f4cadc486ab394faf

Observation 30340b37-b221-4601-8d44-65aa3e124624 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.073467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.073467Z digest=sha256:d7db9535b7ffb3ba3288a6a169233d6b3de99abf237fed1bc82635430d4bbe30

Observation d2d36bcb-25a6-48a7-9c07-3895644aefac · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

NeoBabel: A Multilingual Open Tower for Visual Generation Gemma 2: Improving Open Language Models at a Practical Size

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.198563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.198563Z digest=sha256:e498adc06da95b942763b733f1bb0431ba2217e936ed15ebd574e5dc3754ef02

Observation 31ce32f2-439b-4933-86bf-cfb3c1f8858f · outbound

This paper cites Query-Key Normalization for Transformers.

NeoBabel: A Multilingual Open Tower for Visual Generation Query-Key Normalization for Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.347457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.347457Z digest=sha256:86e07b27b2a5459e6d9c2d01c504d184af01a7b80ab49c50adb70304d222f819

Observation 4f5c16ae-896a-4719-b33c-75995bd231c8 · outbound

This paper cites Classifier-Free Diffusion Guidance.

NeoBabel: A Multilingual Open Tower for Visual Generation Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.441402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.441402Z digest=sha256:c9830d986e78538709cb18f7be7c513f22e18f5c786aae98a89e9650a855559a

Observation c15b6851-c074-4f8b-a2f5-436b3e2d4194 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

NeoBabel: A Multilingual Open Tower for Visual Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.601678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.601678Z digest=sha256:b6f27a8255b025303d56b01b3d39080edc13c88151d856d6ace89421c4f70f2d

Observation d218fce7-2b80-471a-a9a3-58b2ca2c4d28 · outbound

This paper cites Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?.

NeoBabel: A Multilingual Open Tower for Visual Generation Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:31.199854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:15:27.773933Z digest=sha256:2ea8eb8fcb085132ceffc98a524997ff2e114ab803458df870fb26923055ffb7

Observation 1b1501ce-e4d0-4e53-970f-12d20944505e · outbound

This paper cites UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding.

NeoBabel: A Multilingual Open Tower for Visual Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.932896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.932896Z digest=sha256:ab267fd72a8275a18343b0c3775bcec6128632bf62cee7d506dcca470ada727e

Observation 6b08cd28-f79d-4bcf-a796-148f67398d29 · outbound

This paper cites FastText.zip: Compressing text classification models.

NeoBabel: A Multilingual Open Tower for Visual Generation FastText.zip: Compressing text classification models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.033284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.033284Z digest=sha256:1aca4041b9708da20cd226210db2a845ec16d9d210f1addd8d8f68647b9d7e6d

Observation 3fe09f47-9a25-4651-a236-9982132ba8f6 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.124600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.124600Z digest=sha256:db6d8a1000751d04d82aa7c5702ddca319ae53185ea0ea3a5063889421cfd48e

Observation b8005ad2-6a13-410c-84b9-50cc38d8ebab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

NeoBabel: A Multilingual Open Tower for Visual Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.211962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.211962Z digest=sha256:f2787243ca6c506b23c9020be34f3ea9f4eed583e7c57cfa3141cecd6773a94f

Observation 25c1158c-f34c-409b-92e5-072b748233c3 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

NeoBabel: A Multilingual Open Tower for Visual Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.270949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.270949Z digest=sha256:5c1c31c8f6370b415d5d774f0ec322218355801421459bc3162e4270dc1df030

Observation a9136bbe-be09-4210-b494-91e210e38659 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,.

NeoBabel: A Multilingual Open Tower for Visual Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.353210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.353210Z digest=sha256:8297a05a4018b8aaa932a7aabe9c61400f9b21e55ed4be9748ed84e999f65273

Observation ecf123eb-7021-4755-aed5-e17f9c720df3 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

NeoBabel: A Multilingual Open Tower for Visual Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.431651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.431651Z digest=sha256:05820182d0dc1af3ab86acbff8361eaa89d9f34488b35a96baae665c730563d9

Observation e539eae7-006a-408a-b50f-4dc91a831b44 · outbound

This paper cites Transfer between Modalities with MetaQueries.

NeoBabel: A Multilingual Open Tower for Visual Generation Transfer between Modalities with MetaQueries

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.550123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.550123Z digest=sha256:1be58464a9ab6ab9d2726a6d1bd90674e703e31edae34e0cae371bb14ccf33a3

Observation 84849d24-f5bc-423b-8b0c-b90328786bee · outbound

This paper cites RandAR: Decoder-only Autoregressive Visual Generation in Random Orders.

NeoBabel: A Multilingual Open Tower for Visual Generation RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.641209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.641209Z digest=sha256:e4cf7b6e2109e987c582a0d784141556e19687aa07f8a06e9ad334c0ce09c9af

Observation 87a9d507-2af6-4aa2-8ef5-fccc3ff0a297 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

NeoBabel: A Multilingual Open Tower for Visual Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.700319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.700319Z digest=sha256:c066440c670161fbac7451f003b37fbe5629db9dbcec2a3671e7088f891828a4

Observation ef5c483c-58d7-44d2-a74d-067543e4b338 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

NeoBabel: A Multilingual Open Tower for Visual Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.767123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.767123Z digest=sha256:0a516ab605518f8263470617673fa972625f75d721ec3b8b9097c92008cd508b

Observation f6e251c7-9a5f-43f2-861a-702c69504870 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

NeoBabel: A Multilingual Open Tower for Visual Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.869995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.869995Z digest=sha256:39d77fce6e5a41e4f65dc674a1131fc5c186f8aeb734fb9e9d2beb227965c331

Observation c4138b83-579b-4ade-b54e-9356f01811d9 · outbound

This paper cites Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation.

NeoBabel: A Multilingual Open Tower for Visual Generation Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:30.797941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:15:28.957234Z digest=sha256:a993aa5c128caec563b4060dde738d577e736f1bd88b92feb416b3f2a370d512

Observation 8a5f9eb3-2c81-4037-8a93-fffaf842bcaa · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

NeoBabel: A Multilingual Open Tower for Visual Generation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.024903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.024903Z digest=sha256:373eb34549361482e960151fb20b5aabf5a1d9eba256f1ef59d377944adbafd0

Observation d4245c9d-9cff-43e8-a9f2-9c0a2cb82fd8 · outbound

This paper cites A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics.

NeoBabel: A Multilingual Open Tower for Visual Generation A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.104588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.104588Z digest=sha256:fd7dcb7e17b021f062e0ee65b57e075286c2419bb68f65f8dd15f4eaf97f4643

Observation 750250b7-ee98-4417-a503-cce1a9612577 · outbound

This paper cites Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.

NeoBabel: A Multilingual Open Tower for Visual Generation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.193312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.193312Z digest=sha256:c382cb8999a27760dc57308bed758e535d758aab6fcf69adafa0c218623725cd

Observation 5ef95c25-ff74-4ac7-9f6d-d791075b3e6f · outbound

This paper cites DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies.

NeoBabel: A Multilingual Open Tower for Visual Generation DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.272159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.272159Z digest=sha256:832234ea6bcf203cdfb907f3eb0133c906fe61a0e4e39d986a4c80037a1d18b2

Observation c52fc6c6-c3b3-4d9a-9b8a-4a1a76eff23b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.335024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.335024Z digest=sha256:ae5a3190597a45681e0b3f3a90e122189f6544add2fde60529494f9f65229850

Observation 0ccf7602-5185-4563-bf83-13478538a945 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

NeoBabel: A Multilingual Open Tower for Visual Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.406317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.406317Z digest=sha256:403d6092daafb359200892182197d04b297df12bd52528470f4940795f5c50e9

Observation 84989b3e-ab84-48c4-b495-ada5ae199b79 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

NeoBabel: A Multilingual Open Tower for Visual Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.476204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.476204Z digest=sha256:6865a7643fa35e22dbaf1527a3bb8a92ff590b05365a7a095f4e421bcd9158b0

Observation bcb76d96-db31-412a-9d90-ada7b793bf4f · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

NeoBabel: A Multilingual Open Tower for Visual Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.556430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.556430Z digest=sha256:0a6b8bb5f8756f5e6d39175d14900decd44f6b2e4f7727eecbdbbf8527dbc084

Observation 823772c9-11de-4d68-86e5-bef1879fb05d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

NeoBabel: A Multilingual Open Tower for Visual Generation Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.647335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.647335Z digest=sha256:692b856391866608d9938fe98ba73c35fce738907b96f408538e83e5ba2eff66

Observation b6d61fb3-5fe1-4f19-beca-242e2c1675b5 · outbound

This paper cites Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation.

NeoBabel: A Multilingual Open Tower for Visual Generation Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:30.527154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:15:29.734374Z digest=sha256:8805f857945f572b823ee45c7efb4a9612a6859cfe4793f797689c41aed3bce9

Observation 9950c082-bd59-48a2-b4f5-f46c7fa7a7a6 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

NeoBabel: A Multilingual Open Tower for Visual Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.821240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.821240Z digest=sha256:8073de925f341f07e8d35f76f62a7130e5214afd5bd6ae2a397ccb03a970cb3c

Observation 1d9384ee-5dea-4cc9-9f23-c46383ba6e3c · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

NeoBabel: A Multilingual Open Tower for Visual Generation SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.901066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.901066Z digest=sha256:bb1769bd17ec13f50727ae6ce6dc90ba917f6a1cc6e0b2e1b4e30420873f0a90

Observation a32ea1f1-dd39-4029-90ad-51333bebfebb · outbound

This paper cites Qwen2.5 Technical Report.

NeoBabel: A Multilingual Open Tower for Visual Generation Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.991196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.991196Z digest=sha256:11982c38cd732827cf0bfa4fd13101f7428675accac9501bd23053649beec14c

Observation d24a4c6c-7837-4d13-84c2-e59c8f7621c1 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:30.054891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:30.054891Z digest=sha256:00d67e01fc416d70d110cfff7cc79083d8579fc0105ffb9c03675038a91f7d0e

Observation edc4c914-eb9e-4cdc-a452-ca4cdb0daf9f · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:30.153376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:30.153376Z digest=sha256:ea8ee5b83721858e092869c9c45e7bb5ea76c28e682404c83547a4396b3e74f4

Observation 19bb7719-1b68-4860-ba77-6496cb8cccf0 · outbound

This paper cites Table 7 outlines the key hyperparameters used across the three pretraining stages and two instruction tuning stages ofNeoBabel.

NeoBabel: A Multilingual Open Tower for Visual Generation Table 7 outlines the key hyperparameters used across the three pretraining stages and two instruction tuning stages ofNeoBabel

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:31.763902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:15:30.256170Z digest=sha256:ed2193b5a74de72fb2948df97b05f2889bca38a252a6f76b79515bb7b2c46119

Observation fa08158c-384b-4894-a205-04d72179bc66 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

NeoBabel: A Multilingual Open Tower for Visual Generation No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.147148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.147148Z digest=sha256:644d8028acdb08837a7be9e0f2b8e7b8ac526b59e5ed1aee6f2d4690c2dc8b89

Observation 434fee0b-cea0-43fa-991e-e6853c4146ec · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

NeoBabel: A Multilingual Open Tower for Visual Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.005732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.005732Z digest=sha256:4ecffaac152e9822e36c5a40c3f28b80bc6bcb4186974210c26a0707e97b74cd

Observation 0e4f07a4-8b8b-48ae-b291-2ac684ddf55b · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

NeoBabel: A Multilingual Open Tower for Visual Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.901325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.901325Z digest=sha256:683747cf740c1c23b56433492f315a5ac54f5c2d49a914171301132bde9046ab

Observation c37b8bb7-1171-42e1-97a6-dfa957596b5c · outbound

This paper cites Aya Vision: Advancing the Frontier of Multilingual Multimodality.

NeoBabel: A Multilingual Open Tower for Visual Generation Aya Vision: Advancing the Frontier of Multilingual Multimodality

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.315533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.315533Z digest=sha256:5e610f7e9d3bf6ef6853208cca0f149034cfd25f1e16cf42268061276b2237dd

Observation 47cea15a-199a-409b-b992-c139363e0bdf · outbound

This paper cites The AI Gap: How Socioeconomic Status Affects Language Technology Interactions.

NeoBabel: A Multilingual Open Tower for Visual Generation The AI Gap: How Socioeconomic Status Affects Language Technology Interactions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.790040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.790040Z digest=sha256:3e7932205e7c085606b2d77f513f29c249dff6051327b2fde8b9cdddf0647b3b

Observation 6843ef6c-36d7-4520-9596-96ea59f178fc · outbound

This paper cites Behind Maya: Building a Multilingual Vision Language Model.

NeoBabel: A Multilingual Open Tower for Visual Generation Behind Maya: Building a Multilingual Vision Language Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.470790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.470790Z digest=sha256:514e0148d543b795b568a4d60110d3a0b11d75445ecb4d3adbb9ae666013412f

Observation f2b6346c-e4f2-4a19-a54a-ad7d8a06a3b4 · outbound

This paper cites The translation barrier hypothesis: Multilingual generation with large language models suffers from implicit translation failure.arXiv preprint arXiv:2506.22724,.

NeoBabel: A Multilingual Open Tower for Visual Generation The translation barrier hypothesis: Multilingual generation with large language models suffers from implicit translation failure.arXiv preprint arXiv:2506.22724,

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:15:31.575544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:15:25.568660Z digest=sha256:d3cbdcfb402149116c5acb940a31a98058052877c80a0e5911d25f50dbba68e5

Pith citing papers

No inbound Pith citation observations are available.