Pith. sign in

Paper Citation Record · LEDGER

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment

As of 21 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.14682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14682 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:25:29.882104Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2036b856-62f4-4b1c-b5d2-dbe80dd4b4e7 · outbound

This paper cites Qwen3-VL Technical Report.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:24.924803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:24.924803Z digest=sha256:8668597a450c2641a4ceea1be6785f37dc268d78750786b6cbe7d03cba66a77f

Observation 4717a300-0a56-4541-b646-73bd9e992ac3 · outbound

This paper cites F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment F., Tito, R., Mafla, A., Gomez, L., Rusinol, M., Valveny, E., Jawahar, C., and Karatzas, D

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:25.085350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:25.085350Z digest=sha256:13d30563f4e38ff9f3afb77e71396ca9bf64cb0bd8ba1569acdf22cb1ba6f21a

Observation dc8c3ff4-281b-47bf-9fad-0f310a01fa16 · outbound

This paper cites an unresolved cited work.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:25.697642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:25.697642Z digest=sha256:3cad344f254be0262f4901f8628c63e130b687879e6fca26091f2bcd545aabd0

Observation c4adbef3-0997-4118-80b9-3fc3c9d87338 · outbound

This paper cites Boundingdocs: a unified dataset for document question answering with spatial annotations: S.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Boundingdocs: a unified dataset for document question answering with spatial annotations: S

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:25.976663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:25.976663Z digest=sha256:c20c721084639a8607b296306b79f2429cae90060dd1cd8e3141711a0051a8df

Observation c79a4934-3666-4184-a743-2ee99e67c107 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.105226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.105226Z digest=sha256:d171693bbdaec5bb13985d13cef32ec1c9cb526d2dd508fdcdd92f8c7641191e

Observation eaedfb15-ca91-479b-86df-4c3129ee7264 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.185055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.185055Z digest=sha256:f4ef6c4c3d3d8af576512f0472a33c6a918abb77e6b8b0b4a311c9df2c046c0c

Observation b2a1eabf-7860-4adb-9005-28797899353c · outbound

This paper cites B., and Zhang, K.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment B., and Zhang, K

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.305638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.305638Z digest=sha256:d57f59ff6878e2c22c6d5eb5a58ae46424f06c91b643fad6461819552297797c

Observation e01d50c8-b1d0-4f01-89d7-f736e4cbfdae · outbound

This paper cites 4v (ision) system card https://cdn.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment 4v (ision) system card https://cdn

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.440656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.440656Z digest=sha256:53cabbb36172bb470455504b67487efeaf9301c13c443032c67ae9d349e544ea

Observation 462196c9-676c-458e-bba1-3622dd6bb37c · outbound

This paper cites Docile benchmark for document information localization and extraction.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Docile benchmark for document information localization and extraction

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.589095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.589095Z digest=sha256:d11c113a3d642f1255bb81a7c8439e222ee50be0b7a94b84d75c08cff755107e

Observation f6494e40-5c51-44a8-8481-2c36e3883896 · outbound

This paper cites Unsloth: Fast fine-tuning and training of llms.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Unsloth: Fast fine-tuning and training of llms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.706633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.706633Z digest=sha256:9be20eed023dd4ca78afbd6fc29dec162eccf57a14af958526c10849a11c7cde

Observation f53706a9-e410-4714-acf3-b45d8c1c20d0 · outbound

This paper cites A., Jung, K., J \"a lk \"o , J., D’Andecy, V.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment A., Jung, K., J \"a lk \"o , J., D’Andecy, V

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.737077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.737077Z digest=sha256:688b157b5c531f742a537ed8ce44d2b87f425fbe6396562428dd78797659f2a8

Observation 73bcadbe-187f-4e07-97f5-e3da8ea9d093 · outbound

This paper cites Drishtikon: Multi-granular visual grounding for text-rich document images.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Drishtikon: Multi-granular visual grounding for text-rich document images

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.773547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.773547Z digest=sha256:a88582f07eb613485cb6a4893220a2f8247548b44d5c5c8a172bb2d5cbde774b

Observation 2131c131-dfc1-4d60-a3d9-f6d9647d601b · outbound

This paper cites Towards visual grounding: A survey.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Towards visual grounding: A survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.806212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.806212Z digest=sha256:c20b36656f4271913b22158d4caee82fb13500b5b1e2f656920e941d19b8f6f5

Observation 5c91b62c-ff95-4277-ba1e-b3c1c48aecd0 · outbound

This paper cites Docthinker: Explainable multimodal large language models with rule-based reinforcement learning for document understanding.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Docthinker: Explainable multimodal large language models with rule-based reinforcement learning for document understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.922461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.922461Z digest=sha256:f193ea2078ce1317329cd02b39a3164b04c852921e9bc09325cd096501cb3985

Observation 343f646b-b3b4-4c6f-9796-5d88adeb675c · outbound

This paper cites Dogr: Towards versatile visual document grounding and referring.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Dogr: Towards versatile visual document grounding and referring

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:26.974062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:26.974062Z digest=sha256:33089356f4c5e339bd502455e2497aace28ae23a32411acceeae2c7e020670ae

Observation 881efef6-66ec-4da9-baf2-f95d14aa9d36 · outbound

This paper cites X., Wu, H., Wang, W., Feng, F., Wang, C., Luan, H., and Chua, T.-S.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment X., Wu, H., Wang, W., Feng, F., Wang, C., Luan, H., and Chua, T.-S

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.020594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.020594Z digest=sha256:b5ce8892c2ed04e9bdc28cba29d7c3233a5d517fc1c5e9d61a906b66539e26f4

Observation d406b3d1-687d-44c1-b6c5-f9157c4032d9 · outbound

This paper cites A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.067817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.067817Z digest=sha256:34669af82d6d05ce2216d83179625492a2be07b89edeec9c130e63419f5a8d89

Observation faaf72c8-c70b-48df-8122-aa6204c9120c · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.102713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.102713Z digest=sha256:c2d33f6a555b5c5ea9d03744450fc883094326807ee341e706f35de45fce6ccc

Observation d2601945-8aa6-4a3e-bfda-40ed08127078 · outbound

This paper cites arXiv preprint arXiv:2504.04974 , year=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv preprint arXiv:2504.04974 , year=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.131427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.131427Z digest=sha256:01034c4515fc601db8cdb2dbb69fd6b9dad365d30ea71f90988f1429c009fbd8

Observation a2371b19-3b37-43fd-a7f9-b8aa6524e05d · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.252393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.252393Z digest=sha256:379a8c404141d2b33a30169131120a2f7b5fa4d561287b3da9ec6463b34483ab

Observation b8b2dc5d-28dd-4f54-b0c5-edaeb0f13d97 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.350922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.350922Z digest=sha256:1540f688187b8efef845313e7176471531aa888cd761aad4e2a9fa2779a14a0a

Observation b4ce0d30-568d-4922-b27b-bdb39d6b2803 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.404873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.404873Z digest=sha256:389fadc6233a2013d9ae06380bc1c468676935ea1c31cda1f9f600e648ba93a7

Observation 49bfeff3-2245-4a6e-8937-3016b5f99916 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.506499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.506499Z digest=sha256:1b4bb81aeaecfdaa18a8cf5f4f3bfb91d4ff0608e6fe15a08bd19fc41b328120

Observation 9e956eb2-2da9-4612-9e43-fbb20a34db1a · outbound

This paper cites RL Fine-Tuning Heals OOD Forgetting in SFT.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment RL Fine-Tuning Heals OOD Forgetting in SFT

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.583062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.583062Z digest=sha256:ab4465bd5f3b643c714fb9bd014cb8b2d754ff512698fd6035ee33f1888bf87f

Observation fec2b2e4-f917-40a7-81d4-de4e79f26ab0 · outbound

This paper cites Qwen2.5-VL Technical Report.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Qwen2.5-VL Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.678066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.678066Z digest=sha256:4e64bc878f3e75c2b7dbec1c90212ae58cb772520bf5e9a0eeebff596a7826f5

Observation 780c33be-6348-47c9-bfb3-3d95d8b7cfdf · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.781991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.781991Z digest=sha256:ec60097ee4be5147ea7465f9da4c2c7f4afacc2d8021a3ac8b70d4c61787b1f5

Observation d300b33e-3241-4d92-b30f-7e733b518ebe · outbound

This paper cites Proceedings of the 30th ACM international conference on multimedia , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Proceedings of the 30th ACM international conference on multimedia , pages=

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:27.897612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:27.897612Z digest=sha256:6df6a209342c1afabe953cd0c56c4c512680dca5b79527d3164559e89d3b67c0

Observation 041c6618-0378-4711-b63f-3a4efcce3dc4 · outbound

This paper cites PaddleOCR 3.0 Technical Report.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment PaddleOCR 3.0 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.035910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.035910Z digest=sha256:cf37471a392dac8fecdaea663e070fdc9bdd9e14b087286f05857fb6049331c6

Observation 2115405b-ec84-4213-b783-fee12364c78c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.204527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.204527Z digest=sha256:57e8b90cb5cff375cc8d6613ed2f90fb57433977475f83e49cce1f785f368ee9

Observation 522b424f-c63c-430a-b981-315f58a06ab1 · outbound

This paper cites an unresolved cited work.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.272655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.272655Z digest=sha256:e3ef7a6abb74d1b2d2df19cf96504c14a4c2efd41cb0127d0958f17b6ba812df

Observation 13c8431a-09e6-43fe-bf40-f1ecb82f34a6 · outbound

This paper cites DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.323502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.323502Z digest=sha256:5c38bbac859ede59338410c2a63a7dd1ebdb14e7ae1392c83c9c30238cdd997d

Observation 9432e48e-8716-4962-8802-5ef19a825d05 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.375861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.375861Z digest=sha256:eb180ceb63c1f8d12abc44a35d00c8a3c64c9982fec203900f5ba2a16b8a5acb

Observation d9b59ccb-2086-4ed3-9858-5919c10c2246 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.475131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.475131Z digest=sha256:e419e199e668a3332c42097c733866b4a7031fc21660f651400ce5b5e9fd1b5d

Observation af234a1e-faf5-40b5-8d18-930e6cf4014f · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment KOSMOS-2.5: A Multimodal Literate Model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.575255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.575255Z digest=sha256:f868505fd298b20eb72fc94cde896f5ced8052652070a869a86ac28c617204e1

Observation 9adb5f5e-4e82-433d-95be-3339e344f9c3 · outbound

This paper cites Farrar, Straus and Giroux , year=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Farrar, Straus and Giroux , year=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.646952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.646952Z digest=sha256:7454757af377bc0d319a55f80d7331c6415e6aef6e5395b36217363df85c31a1

Observation 54621a2c-7471-45b4-94fc-9b72a41f2ae0 · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.694126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.694126Z digest=sha256:1c3162fc6bc33dbe6ca6433afefe472fff3f3868db52fee2305220dbf402d1cd

Observation 7419a4c0-5df5-4b11-9c48-0d4ee9eb397d · outbound

This paper cites The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment The Thirty-ninth Annual Conference on Neural Information Processing Systems , year=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.757507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.757507Z digest=sha256:470bbd59bab5cea300ba13d9f1ff28d162a25bad4b0b172ac8d90efdb775761b

Observation 12030f40-4f56-4021-904a-ae70d92b54d8 · outbound

This paper cites arXiv preprint arXiv:2503.20752 , year=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv preprint arXiv:2503.20752 , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.830889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.830889Z digest=sha256:b41ca6eac5b248efd859d5e4d3827b9ce9ae54bf4294977976e5a4a619ff9270

Observation 79ad8275-70e7-4b9b-9103-2f8c9b7edef4 · outbound

This paper cites 2021 , eprint=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment 2021 , eprint=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.922743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.922743Z digest=sha256:abc296e87b7c0ee9498a772d29b44c353ff8c99bef8d3b6879c5df5351d460dd

Observation 7ad47b11-3858-4566-bdcb-edff050db110 · outbound

This paper cites International Conference on Document Analysis and Recognition , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment International Conference on Document Analysis and Recognition , pages=

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.996885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.996885Z digest=sha256:50cd8c48099e8422c68b01e592d283a22048c46c47c1972fed136b7e7dd333f9

Observation 5593d532-4328-4f7c-a157-7e8cd098284f · outbound

This paper cites International Conference on Document Analysis and Recognition , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment International Conference on Document Analysis and Recognition , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.136088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.136088Z digest=sha256:89de5aaf815995046c5a6cf875c7f224ae806f7e927311accd86852b038745b2

Observation f683b3cf-f5ff-4205-bda7-56c9c97b909c · outbound

This paper cites Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.294248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.294248Z digest=sha256:76b639117e40d609a463cc1dcd31e7bdeb3789faa0dfe8f0d76fafb038167c84

Observation fa4f7381-55aa-418b-870c-131d313286db · outbound

This paper cites Giovannini et al.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Giovannini et al

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.426738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.426738Z digest=sha256:d70ecb2d76068eb18e95a0493d43201c0d195868958daba44511f0ed50755f54

Observation 701567ac-527d-4b96-8970-565cee502a47 · outbound

This paper cites arXiv e-prints , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv e-prints , pages=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.514942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.514942Z digest=sha256:b53d0db9bad4f92c301f7e08b030014465267891c156608287cd08b5c3614d5b

Observation 7a79af6e-fca0-46a9-bd40-448eb7c6c724 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.557821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.557821Z digest=sha256:8318939d2a869c6caffc743897d76105976a7e4c8bb11e82e4003dac788561fc

Observation 96b560bd-7b97-4032-a3a5-ba05062031f3 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.590993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.590993Z digest=sha256:db6d0c828819a859fdf7106e080074bad7ef2af2657ce33131ab231a5bc7e3cc

Observation 9d1271e7-f5f7-4949-82fc-f1c27418339a · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.636892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.636892Z digest=sha256:2fed2371b0a18aaf4ed7223bbd7980b80a170f6828655ecb445430c4d0eed6c8

Observation 1be60655-032b-4dae-8849-3444dc55399e · outbound

This paper cites 2021 , eprint=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment 2021 , eprint=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.667195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.667195Z digest=sha256:cd6c8f5e4f61352993c94a34883a2ad767b7d0897cdf594d5c54f7dbdb4d8651

Observation 0652bd7b-1633-4164-8c7a-dd63c29e3102 · outbound

This paper cites GitHub repository , howpublished =.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment GitHub repository , howpublished =

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.699471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.699471Z digest=sha256:ec0189bdadb9af0888a9f93fa87e7eba7d9695617a07e34d2341acf2eb49a8a5

Observation 090466a4-ad6c-47ec-995c-9690a034761f · outbound

This paper cites Proceedings of the IEEE/CVF international conference on computer vision , pages=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Proceedings of the IEEE/CVF international conference on computer vision , pages=

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.732418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.732418Z digest=sha256:8bb36b61a5a12c3f54d5782c8ccf0dba65238346b1e6e1ddb229c2e77181f5c5

Observation 159a1476-8bce-4e1a-a9b3-9158c145de9e · outbound

This paper cites Towards Visual Grounding: A Survey , year=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment Towards Visual Grounding: A Survey , year=

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.781391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.781391Z digest=sha256:0da9040a8db84030c253587cd694a050d737fb5ae06113bc7ab8d0888f1aaad5

Observation e4edaa9f-d5c2-4046-b961-43df417ede1d · outbound

This paper cites arXiv preprint arXiv:2509.10345 , year=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment arXiv preprint arXiv:2509.10345 , year=

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.816900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.816900Z digest=sha256:9351963f7352487352d56176975b0ccdf77b275de5103c4bb1f1c2e02ae56046

Observation 27bf238b-3274-4a5f-a3da-0091ec70a33b · outbound

This paper cites MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding and Grounding.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding and Grounding

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.847526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.847526Z digest=sha256:5254c6b212a99105481601bc0e33522251f487ffd1d3d4f5186a3c4ce3df08a4

Observation f994e11e-2b88-4883-af63-08b7a51a4cb2 · outbound

This paper cites 2025 , eprint=.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment 2025 , eprint=

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:29.882104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:29.882104Z digest=sha256:973461da65f8c8eb4d9542982c077947afa7b1c1eb05d5d0bd8cb00d9e1c24ff

Pith citing papers

No inbound Pith citation observations are available.