Pith. sign in

Paper Citation Record · LEDGER

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution

As of 6 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2606.27313.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.27313 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:29:01.706006Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact17
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ae17375d-8cfe-46bb-91e5-ab4ce7e5c972 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.806093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:f4111328b3f26fb5e5f06d7fb686f8a2954389410600aad1a523989ce0a3bbb2

Observation 7625bbb0-e978-495d-a6e3-5b411253699b · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.828216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:66e24933e3f1e830bb33bd529d310faee092a4903c2cecf3e83b04d2535bf80c

Observation 2e72d204-3baf-4a88-b48e-6e343d5a7140 · outbound

This paper cites Data Filtering Networks.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Data Filtering Networks

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.778449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:d009d7b48b57e2a8add6e4227f9f68372c74f8a757e1e232d1f32e1d81b111c3

Observation 38873eda-e451-4b91-a063-828699d464ed · outbound

This paper cites Seed1.5-VL Technical Report.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.817141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:6a6bcef262f7e8279fbd2b1ba0fb4ab4ea8928b412c6bd3e117be33c45a0a139

Observation 0e8c2db2-213c-4ac7-bf67-c2e813ea4a1c · outbound

This paper cites Adam: A Method for Stochastic Optimization.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Adam: A Method for Stochastic Optimization

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:35:39.798333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:fad82ac53be18548914d242e3155bfc2ea38f12e40395aec00c090372c03f915

Observation 1bcc357d-d70a-4af0-b359-a918c7e62b88 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:35:39.802336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:f168e1230e44a326bef8020377394e7c5258c3091e789291e4d0b77f5a0b7c0c

Observation 1fa8c5de-9e72-4954-ac08-d538123a1901 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.749751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:432715f1e993a291b3ee6feb700eb260efe7d5bbdbc1ef4212829187c6b646ad

Observation 4e0b7bf6-83b7-486c-abbb-6aad52799b9a · outbound

This paper cites Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Open-MAGVIT2: An Open-Source Project Toward Democratizing Auto-regressive Visual Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.766107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:83b58e5ee3ff99dd49d988395b0a7edc80ebed895ab725ac4798cc734ede2126

Observation dfe4d32a-b1e7-4033-905c-3b8c1291d452 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.810511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:39083bc6880c9c85b601bd43c9f79f8236474e81731df36adfc9299ee3f6c3bf

Observation 354cf9f5-4d8a-4eb2-9241-e8d4a08cab49 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.787611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:eaa8b361094173ceb5de7a65035fb9e7a76dd20ecb3bfc6d5819da32ae6f7199

Observation f9c1cfc3-f1f9-4d7f-b681-9cdc08b910ad · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution DINOv2: Learning Robust Visual Features without Supervision

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:35:39.770072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:e734a9e75ec2d8624b892a2e20257baa42098a1c6bff1c5a895fe8ca18510048

Observation 65cdf4b7-5c5e-497f-a99b-c2daa071f8fe · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.790242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:99d2e87118b05b7666078938ad55521f76bd2ddc2cdcf338dd5c5bf157df9973

Observation c47199f5-6eca-44f0-bc2e-9a0b1c4dfdeb · outbound

This paper cites Kimi-VL Technical Report.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Kimi-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.814018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:b2cfd96523cfd615f4fb2b084d5dedfb00a22edd22756466e7f78ff1d2784623

Observation 176b9ed9-c8b7-4caf-99d7-63051936edfa · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Longcat-next: Lexicalizing modalities as discrete tokens

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.773896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:cb1598284b540f3e711e03112304e1c124c143c19790f0d43d1239c143e1fedf

Observation 3f953edb-c4c1-48be-a023-42c8ed4873be · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.824779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:c7e0f266bca9a8340eae04860723b2f118618520b917c521578c662fdf5e103d

Observation bcf0f9c9-d93c-4c70-9d33-6ee50d755fe5 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Wan: Open and Advanced Large-Scale Video Generative Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.795063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:31f000589164163c7a2d806c67edb912892a07f401d35b0b670a36ff148049b4

Observation 7379842e-0657-4fa3-b8f7-e82f09528701 · outbound

This paper cites Qwen-Image Technical Report.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Qwen-Image Technical Report

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.799075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:94dd78a86f849193e349c988075740c585b813104a0c75ac2295b25ae97799b5

Observation a3f170df-9fc5-46a5-b0a6-1a16aca0cb78 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.768511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:798871e49e70dadb70af46af61ade6fc3888e4e1ceb5278a026922f0a103a411

Observation 3259b063-af9a-419e-bfda-0c4ea5bb1824 · outbound

This paper cites Sail-vl2 technical report.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Sail-vl2 technical report

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:35:39.756114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:a148178ff886a4f369d4edb8e3645823d3759eb9a20acef8a733e01a7927cb9d

Observation 3dfe8624-b133-4ee9-add3-f2b993d7aff9 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:39.821652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:a201a93d0829894ffeeb7a81b8943a4d08daff89dad1c0e6725c9be9f3d596c4

Observation 26ec2148-0860-4d53-a399-51b11bf23add · outbound

This paper cites an unresolved cited work.

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-07-06T18:02:46.620046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-01T06:29:01.706006Z digest=sha256:3847fe133580a0db568ad7fafc5083d237188ad800e4041c73021cc8346eab75

Pith citing papers

No inbound Pith citation observations are available.