Pith. sign in

Paper Citation Record · LEDGER

NeoBabel: A Multilingual Open Tower for Visual Generation

As of 7 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2507.06137.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06137 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:15:30.256170Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact4
  • verified fuzzy1
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cb33f64b-e506-4043-9aba-00eb31d940d3 · outbound

This paper cites Maya: An Instruction Finetuned Multilingual Multimodal Model.

NeoBabel: A Multilingual Open Tower for Visual Generation Maya: An Instruction Finetuned Multilingual Multimodal Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.391619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.391619Z digest=sha256:e7afc73a2bb02f48a5a1eda383dba9ef065337e79d0b2a8be6948d7792ac5c31

Observation 631a2c31-c178-42c3-bb22-427e5ee31236 · outbound

This paper cites Qwen2.5-VL Technical Report.

NeoBabel: A Multilingual Open Tower for Visual Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.702878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.702878Z digest=sha256:48d52fa7fe2f0ae3ee7dfe094ab22098d4bdc825cec595f032d89e3a4ebd5168

Observation 9946d252-2a16-4bd1-a9e0-50e93c74d681 · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

NeoBabel: A Multilingual Open Tower for Visual Generation Unveiling Encoder-Free Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.468458Z digest=sha256:7bc2be7fa66adfdd6920e76614ca574dbdd0b62c07d61b77017ec69d0cb72c39

Observation b134bd5a-39fd-43cc-9b02-961aba8666fb · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

NeoBabel: A Multilingual Open Tower for Visual Generation EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.683136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.683136Z digest=sha256:1551659cfe311fa034de8cdec8d9ea71fc9fd3fc2068e8c927d2307cdc40cb0b

Observation 254445f6-6f51-42b2-a357-beb673a5d945 · outbound

This paper cites Unified Autoregressive Visual Generation and Understanding with Continuous Tokens.

NeoBabel: A Multilingual Open Tower for Visual Generation Unified Autoregressive Visual Generation and Understanding with Continuous Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.807532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.807532Z digest=sha256:cee23a62049f4e6237d48ea0476973e66ce921ac5ecdade2a1d279edea5747fb

Observation d72cfab3-18e6-48e9-91b5-a235dabc64fd · outbound

This paper cites Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You.

NeoBabel: A Multilingual Open Tower for Visual Generation Multilingual Text-to-Image Generation Magnifies Gender Stereotypes and Prompt Engineering May Not Help You

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.989066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.989066Z digest=sha256:f42f56b892ed84c85d581335a5b0a32358f0869024019be0e32bcf917128e7ee

Observation 30340b37-b221-4601-8d44-65aa3e124624 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.073467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.073467Z digest=sha256:403bea7ab60a92e0370f929c5edd291121b7287accab4d1ab32a5483fd850a43

Observation d2d36bcb-25a6-48a7-9c07-3895644aefac · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

NeoBabel: A Multilingual Open Tower for Visual Generation Gemma 2: Improving Open Language Models at a Practical Size

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.198563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.198563Z digest=sha256:28c1e3a4e4f6a89a252c7f940f5db826efecdc6eb26567a746f708d8691ad79b

Observation 31ce32f2-439b-4933-86bf-cfb3c1f8858f · outbound

This paper cites Query-Key Normalization for Transformers.

NeoBabel: A Multilingual Open Tower for Visual Generation Query-Key Normalization for Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.347457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.347457Z digest=sha256:fce4505c2a6a4e1e827bf38cdca9d0004f4e0ac19160cb135da53651d30a2b84

Observation 4f5c16ae-896a-4719-b33c-75995bd231c8 · outbound

This paper cites Classifier-Free Diffusion Guidance.

NeoBabel: A Multilingual Open Tower for Visual Generation Classifier-Free Diffusion Guidance

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.441402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.441402Z digest=sha256:24fc3df2e64c5befbd15956441b71f5ebddd72bfce6842180b9de66ff12985a5

Observation c15b6851-c074-4f8b-a2f5-436b3e2d4194 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

NeoBabel: A Multilingual Open Tower for Visual Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.601678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.601678Z digest=sha256:8e655f4e74ece9fe324dc5fb284520cfa1c4e34f39182165124c86fde7e60f41

Observation d218fce7-2b80-471a-a9a3-58b2ca2c4d28 · outbound

This paper cites Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?.

NeoBabel: A Multilingual Open Tower for Visual Generation Can Machine Translation Bridge Multilingual Pretraining and Cross-lingual Transfer Learning?

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:31.199854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:15:27.773933Z digest=sha256:28fad10082354239065ebf8e3603819eda2d93802335ba7aff86fdf24007e92e

Observation 1b1501ce-e4d0-4e53-970f-12d20944505e · outbound

This paper cites UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding.

NeoBabel: A Multilingual Open Tower for Visual Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:27.932896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:27.932896Z digest=sha256:e77858fe0db5b820278294e84c31a68512279f746a925933c1a8404137c31846

Observation 6b08cd28-f79d-4bcf-a796-148f67398d29 · outbound

This paper cites FastText.zip: Compressing text classification models.

NeoBabel: A Multilingual Open Tower for Visual Generation FastText.zip: Compressing text classification models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.033284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.033284Z digest=sha256:9c7a6d2cd55703e0ab95a19c75753590b4ab2dd0745cda145f49646490932f85

Observation 3fe09f47-9a25-4651-a236-9982132ba8f6 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.124600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.124600Z digest=sha256:ce234a96d4918cb162e8e01759fd0f927dea2a99883f5db770cc5a3460dc252c

Observation b8005ad2-6a13-410c-84b9-50cc38d8ebab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

NeoBabel: A Multilingual Open Tower for Visual Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.211962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.211962Z digest=sha256:96b17b5bdb467c4945b75c6e926c4f2b2c41ea88458cbd3ee9bfbcc9630e92a3

Observation 25c1158c-f34c-409b-92e5-072b748233c3 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

NeoBabel: A Multilingual Open Tower for Visual Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.270949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.270949Z digest=sha256:ff5fb81f4c36b8c5c7363deed850ae9850e9efb8c717fbe038444d02eda863c1

Observation a9136bbe-be09-4210-b494-91e210e38659 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,.

NeoBabel: A Multilingual Open Tower for Visual Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.353210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.353210Z digest=sha256:8d64dbca96def193ac5134479ce35a5df8629a2f44007daea9cedf86ddd29160

Observation ecf123eb-7021-4755-aed5-e17f9c720df3 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

NeoBabel: A Multilingual Open Tower for Visual Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.431651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.431651Z digest=sha256:12112dc56e80a5e5a9f91cbe728758a6d402d2ada39df960db68a53981efceb9

Observation e539eae7-006a-408a-b50f-4dc91a831b44 · outbound

This paper cites Transfer between Modalities with MetaQueries.

NeoBabel: A Multilingual Open Tower for Visual Generation Transfer between Modalities with MetaQueries

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.550123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.550123Z digest=sha256:e0ac78e275387cb2a9c8a7a986f4362563b1a673217b16693a9b789aa4bd61e3

Observation 84849d24-f5bc-423b-8b0c-b90328786bee · outbound

This paper cites RandAR: Decoder-only Autoregressive Visual Generation in Random Orders.

NeoBabel: A Multilingual Open Tower for Visual Generation RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.641209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.641209Z digest=sha256:327f7daea1102764f096eb07e7fdd414f406207e4eb58d59a54781a9cc3941c7

Observation 87a9d507-2af6-4aa2-8ef5-fccc3ff0a297 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

NeoBabel: A Multilingual Open Tower for Visual Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.700319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.700319Z digest=sha256:fcefe2c0fa5fcb6f26b3da25c386898628e9517f545fc00f6b40cc34726b8bef

Observation ef5c483c-58d7-44d2-a74d-067543e4b338 · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

NeoBabel: A Multilingual Open Tower for Visual Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.767123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.767123Z digest=sha256:722ecad134ef54cc3bace3f81e550ea0afd21c6ccc14533fa746cac1b14eec43

Observation f6e251c7-9a5f-43f2-861a-702c69504870 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

NeoBabel: A Multilingual Open Tower for Visual Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:28.869995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:28.869995Z digest=sha256:61cd6e53c8b92d1550472f0698afde6171ce57ae4b6d6b1b639f7d94c0688120

Observation c4138b83-579b-4ade-b54e-9356f01811d9 · outbound

This paper cites Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation.

NeoBabel: A Multilingual Open Tower for Visual Generation Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:30.797941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:15:28.957234Z digest=sha256:9a43a97a4b115d0928c8168264cf1e2932b4105fcea6d1c7d50747352d6378f8

Observation 8a5f9eb3-2c81-4037-8a93-fffaf842bcaa · outbound

This paper cites Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model.

NeoBabel: A Multilingual Open Tower for Visual Generation Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.024903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.024903Z digest=sha256:9b1a2b66ebec7176b02097171a8bd8054f94c548cd1c26d491ad5a5e5042eb80

Observation d4245c9d-9cff-43e8-a9f2-9c0a2cb82fd8 · outbound

This paper cites A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics.

NeoBabel: A Multilingual Open Tower for Visual Generation A Post-trainer's Guide to Multilingual Training Data: Uncovering Cross-lingual Transfer Dynamics

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.104588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.104588Z digest=sha256:fb7af13c8bc673dd121392ef88695035cbe4acdd203d71bd24b666c4da2b0d9e

Observation 750250b7-ee98-4417-a503-cce1a9612577 · outbound

This paper cites Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation.

NeoBabel: A Multilingual Open Tower for Visual Generation Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.193312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.193312Z digest=sha256:15be499ad36142cb063e6d382921a738cad1aeb094976bbef95a1412124258ab

Observation 5ef95c25-ff74-4ac7-9f6d-d791075b3e6f · outbound

This paper cites DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies.

NeoBabel: A Multilingual Open Tower for Visual Generation DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.272159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.272159Z digest=sha256:022f739e19782142674a789bdec620fe17193180801ee09895f53fa83a8c5c04

Observation c52fc6c6-c3b3-4d9a-9b8a-4a1a76eff23b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.335024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.335024Z digest=sha256:906f4648cebd3af1fcaf491bbcdfb5e78fc65912fe6cfef74040db8ba1a266cb

Observation 0ccf7602-5185-4563-bf83-13478538a945 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

NeoBabel: A Multilingual Open Tower for Visual Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.406317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.406317Z digest=sha256:30b2909f32960ef948ef069a07cff7eb359b64eb9f1a0de086646904fcd77cf1

Observation 84989b3e-ab84-48c4-b495-ada5ae199b79 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

NeoBabel: A Multilingual Open Tower for Visual Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.476204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.476204Z digest=sha256:a783732a090c4513f7672e6944473f252efffa6dbe9cd05de990e6fc10c8b10c

Observation bcb76d96-db31-412a-9d90-ada7b793bf4f · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

NeoBabel: A Multilingual Open Tower for Visual Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.556430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.556430Z digest=sha256:9308f16aba92d5262ed84ac3f1d90ead388bc157aa0496a34cdb956fb26c4131

Observation 823772c9-11de-4d68-86e5-bef1879fb05d · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

NeoBabel: A Multilingual Open Tower for Visual Generation Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.647335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.647335Z digest=sha256:9740faf99a009c756fa9ac5582f5ed8b9e27117960a132eec71f89e67cf6be55

Observation b6d61fb3-5fe1-4f19-beca-242e2c1675b5 · outbound

This paper cites Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation.

NeoBabel: A Multilingual Open Tower for Visual Generation Lost in Translationese? Reducing Translation Effect Using Abstract Meaning Representation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:15:30.527154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:15:29.734374Z digest=sha256:b427fe223efc3da7f716664847aa0c0fd1a79869753dc23c546c947c8771aeb5

Observation 9950c082-bd59-48a2-b4f5-f46c7fa7a7a6 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

NeoBabel: A Multilingual Open Tower for Visual Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.821240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.821240Z digest=sha256:43448deca4c1f083423af5daa457225bd1726c35c0413662a07350656847465f

Observation 1d9384ee-5dea-4cc9-9f23-c46383ba6e3c · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

NeoBabel: A Multilingual Open Tower for Visual Generation SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.901066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.901066Z digest=sha256:2f406186c3443be8f92915a92d03b29c89d867990bbab5d86eed20259154a73d

Observation a32ea1f1-dd39-4029-90ad-51333bebfebb · outbound

This paper cites Qwen2.5 Technical Report.

NeoBabel: A Multilingual Open Tower for Visual Generation Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:29.991196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:29.991196Z digest=sha256:8d2b38566b3df4acdb52e1a2debd75f02c3f29913706cea8a83ada0d6851da33

Observation d24a4c6c-7837-4d13-84c2-e59c8f7621c1 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:30.054891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:30.054891Z digest=sha256:66d3a4816d3ec11055999ef51f6b3b5a788f4e39893848fcdb3eb2f0bbb406ec

Observation edc4c914-eb9e-4cdc-a452-ca4cdb0daf9f · outbound

This paper cites Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation.

NeoBabel: A Multilingual Open Tower for Visual Generation Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:30.153376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:30.153376Z digest=sha256:fa34b5b5a13ba850b903161e9aa3f45774b7b08cadc3f0feec6a15cc0c35bbc1

Observation 19bb7719-1b68-4860-ba77-6496cb8cccf0 · outbound

This paper cites Table 7 outlines the key hyperparameters used across the three pretraining stages and two instruction tuning stages ofNeoBabel.

NeoBabel: A Multilingual Open Tower for Visual Generation Table 7 outlines the key hyperparameters used across the three pretraining stages and two instruction tuning stages ofNeoBabel

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:15:31.763902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:15:30.256170Z digest=sha256:2ff7deb7d23811571d4922588cd76734aa993eb78d1855a393ca878c5384d021

Observation fa08158c-384b-4894-a205-04d72179bc66 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

NeoBabel: A Multilingual Open Tower for Visual Generation No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.147148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.147148Z digest=sha256:7645b332a6cec2106d9270ef10470bb6d034945797a3c15f6ef72c0a964bfd8d

Observation 434fee0b-cea0-43fa-991e-e6853c4146ec · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

NeoBabel: A Multilingual Open Tower for Visual Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.005732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.005732Z digest=sha256:837757436fbc013a11dd093e2f3d75bc2dd3ae2704ecb795f765a95d42bef2aa

Observation 0e4f07a4-8b8b-48ae-b291-2ac684ddf55b · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

NeoBabel: A Multilingual Open Tower for Visual Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.901325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.901325Z digest=sha256:6d755a1da2fdf2b8d49b922f04a152b06bb88e9de7de6a2bdad5822a7581d87f

Observation c37b8bb7-1171-42e1-97a6-dfa957596b5c · outbound

This paper cites Aya Vision: Advancing the Frontier of Multilingual Multimodality.

NeoBabel: A Multilingual Open Tower for Visual Generation Aya Vision: Advancing the Frontier of Multilingual Multimodality

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:26.315533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:26.315533Z digest=sha256:c118a1d863f246f85c9aae4d2ea9b052f96342033aee797c6e7e23c199cee187

Observation 47cea15a-199a-409b-b992-c139363e0bdf · outbound

This paper cites The AI Gap: How Socioeconomic Status Affects Language Technology Interactions.

NeoBabel: A Multilingual Open Tower for Visual Generation The AI Gap: How Socioeconomic Status Affects Language Technology Interactions

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.790040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.790040Z digest=sha256:5328d44677f48d5f96107de1363628424ef39292aaad095ae528dde7ef3ef615

Observation 6843ef6c-36d7-4520-9596-96ea59f178fc · outbound

This paper cites Behind Maya: Building a Multilingual Vision Language Model.

NeoBabel: A Multilingual Open Tower for Visual Generation Behind Maya: Building a Multilingual Vision Language Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:15:25.470790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:15:25.470790Z digest=sha256:cffa1245bd73cfa44165eb63eac1674fe1efbf6436e43da02aa4448f920812a7

Observation f2b6346c-e4f2-4a19-a54a-ad7d8a06a3b4 · outbound

This paper cites The translation barrier hypothesis: Multilingual generation with large language models suffers from implicit translation failure.arXiv preprint arXiv:2506.22724,.

NeoBabel: A Multilingual Open Tower for Visual Generation The translation barrier hypothesis: Multilingual generation with large language models suffers from implicit translation failure.arXiv preprint arXiv:2506.22724,

Reference 2025

Resolution
verified exact
raw_fallback, observed 2026-08-06T19:15:31.575544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T19:15:25.568660Z digest=sha256:266c0f3b601da3ec29ef8a7f95124116b73a2df7417db645fe4dbf08ee0f7da4

Pith citing papers

No inbound Pith citation observations are available.