Pith. sign in

Paper Citation Record · LEDGER

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

As of 17 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2605.19528.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19528 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T06:31:04.432486Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:04:38.497190Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact18
  • verified fuzzy33
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b018fbbf-9a92-485f-b01d-38de39e39884 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.355288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:5f8824ec51afd633606bda9ef693c35b214706ccf054396dfd5026de83622396

Observation dbf85dab-361b-4472-a52f-3fa479283465 · outbound

This paper cites Claude Opus 4.7.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Claude Opus 4.7

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.359051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:41fe25a13c88273692b6f30039e13f2a8de9dbdc7778d74edd87e35ab0bcd01a

Observation f58ab116-a982-48a2-867b-bdf495e5577d · outbound

This paper cites Claude Sonnet 4.6.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Claude Sonnet 4.6

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.374685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:f07767c13ba1941fabb60354c4495e772376d23a13b3f95513bce1b68e73ecb7

Observation 8a789041-8fd0-463a-87a5-48ce9809560b · outbound

This paper cites Qwen Technical Report.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Qwen Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.731646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:5f6c03cf04d9e5a43c036478c810d98e5330884246ba0dbaf41d00f5fcff2259

Observation 4231b4db-0387-4bf2-b5ca-bcafb36b1360 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Scaling spatial intelligence with multimodal foundation models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:33:05.738868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:79097b2c54bad6828b65f4eb38c3cb904fdac080b1ee7705e6955aac8e906581

Observation a0931d2b-9e4a-4d55-90b2-fdb856a52e83 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs SAM 3: Segment Anything with Concepts

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.735281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:1cc4126f4d980af76f81ea39b9cea7e40b1967dab289c0c208698c05f9342b11

Observation c3d88411-23e0-4ded-9057-fc5727cf7f59 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.397771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:e69591d635c734c2add9bce62074dd99a3ea520c56e680d175c0a9ef88e3694c

Observation d516a63c-5ac9-43a1-8973-23a5153b4a34 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.399560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:a5fdfd234b799be2f01077698c783324d4dc4dbf6a94f4db06d32c67b49d80c3

Observation 4df36fd8-41c1-4182-90b6-793b32e47183 · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Spatialrgpt: Grounded spatial reasoning in vision-language models.Advances in Neural Information Processing Systems, 37:135062–135093

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.415791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:f1b17c6023fff9ac4f24c627a4924a6935994d627d81e32f553a5dadbfa1e08e

Observation 8beecc7a-8a68-4681-838b-7399bedac2ec · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.744553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:a215e034ad4e272c6ebdf9981880ee8abbc77554121ebfdfee0624e9e5376760

Observation 3f26e81f-a39b-49c7-8ca6-929a00667dcf · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.403334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:deb66c4c9bb1f189a843411bab9fd57a157bcb45ef35a9e1ba3ee07d994f3282

Observation 41ef58a2-7630-4b27-8d9f-455db7947792 · outbound

This paper cites DeepSeek-V4.https://www.deepseek.com/.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs DeepSeek-V4.https://www.deepseek.com/

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.380340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:81e722acba5fde02f3eb24c83cae7880d1f8f1b7d6188b39f1a599553c20076b

Observation 9a01ac1f-6c16-4883-8bf5-7a95a003cd18 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.687991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:d3c98d41afd0635dae0d4e5d391883ad406d53a18ca3492e684d8a991c940e8b

Observation d61abc35-7a43-4701-9c77-1672ea97cc24 · outbound

This paper cites A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs A Survey of Large Language Model-Powered Spatial Intelligence Across Scales: Advances in Embodied Agents, Smart Cities, and Earth Science

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:33:05.703873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:bfa4649f7517e14000fa47230257a41ff147c5ba46028bba294fea010767e0ea

Observation 6741944c-ac64-4752-98e7-08b6811f916c · outbound

This paper cites Refocus: Visual editing as a chain of thought for structured image understanding.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Refocus: Visual editing as a chain of thought for structured image understanding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.367200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:178b5f5d7659b81e47dccef99922ef13cdf3175a9125ffb5b340eba4e595dd20

Observation 51260352-c3c5-41a2-8eab-777dcc0ae125 · outbound

This paper cites Pal: Program-aided language models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Pal: Program-aided language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.376825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:94c7ee0cf9992f74e4a2a1c489ce1b37a185c49fa6c40eb3f676434a892ba7a9

Observation 16fa4cc2-0d9b-423d-93bb-13a411b84980 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.395857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:9d1dd2c9e8ac77ed1d8995f4467b3dfb5b883af019878111b6c0a9c5813f3b0a

Observation 464f2f10-5ede-4dac-b5fe-94dc1c7e1222 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.Advances in Neural Information Processing Systems, 37:113991–114017.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Chat-scene: Bridging 3d scene and large language models with object identifiers.Advances in Neural Information Processing Systems, 37:113991–114017

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.411025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:ff22d0131742d4769359340ffdf34e1e6c3d3365664bd300fb7e3983af4f2e68

Observation fd083deb-647f-4eee-92ea-b034dd64cc69 · outbound

This paper cites Vision-r1: Incentivizing reasoning capability in multimodal large language models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Vision-r1: Incentivizing reasoning capability in multimodal large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.365364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:9a8cf0d78b6acf12f363de2afa66bddc18471b7de813a13885511649e25b9e6c

Observation 0642c3c6-f259-40b0-8616-5db05cd96702 · outbound

This paper cites GPT-4o System Card.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs GPT-4o System Card

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.712873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:9fc3ac03e0c7b90fb558bbbd12886fc976330e5400a7fab464c6f56f77846f67

Observation f509e4db-a57f-4a80-91bb-6d4cfcef46d7 · outbound

This paper cites Neuro-symbolic data generation for math reasoning.Advances in Neural Information Processing Systems, 37:23488–23515.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Neuro-symbolic data generation for math reasoning.Advances in Neural Information Processing Systems, 37:23488–23515

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.363390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:aecc4f139b8a96b7a30a5b7e95346fffbaeab4fb2ac164eb012d38d5f1122c7e

Observation 3ebddc71-cdaa-4af5-9437-b71b396eb854 · outbound

This paper cites Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Visual instruction tuning.Advances in Neural Information Processing Systems, 36:34892–34916

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.361487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:ade9c1c2f983183a4bf6795d43305d2d5d5dd2e58a5155dc391868e2618ada12

Observation 07fca96f-014a-437c-84b6-ae87327736e5 · outbound

This paper cites Ssr: Enhancing depth perception in vision-language models via rationale-guided spatial reasoning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Ssr: Enhancing depth perception in vision-language models via rationale-guided spatial reasoning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.419559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:43e287d7ed3d1e5c9ace463014ae615be8dc3295c0b4527d727a04d4795c08f8

Observation aede7af8-9d83-46ce-bad5-6683b9f34a5f · outbound

This paper cites Kimi K2.6.https://www.kimi.com/blog/kimi-k2-6.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Kimi K2.6.https://www.kimi.com/blog/kimi-k2-6

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.357170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:7d8acefd75ab9f80af57c22857880cb54321313ee8ae27b7f1d9ae350301f066

Observation e5cdae1f-dd1b-4ab7-8b57-a9df05839849 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs WebGPT: Browser-assisted question-answering with human feedback

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.719493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:8f56e02a68309185524081c9cca0584fc48d6dd1d6e6b0ed2176b697326f9b53

Observation 2b31955a-f8a2-47ec-93fd-479d5d088f03 · outbound

This paper cites an unresolved cited work.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-05-20T06:33:06.371191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:4704ac89f50287164a1c96e9727be90ca158f2873704b0f3e17647be243239e1

Observation 717405a7-da50-424f-9667-7f76cc25c438 · outbound

This paper cites Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Logic-lm: Empowering large language models with symbolic solvers for faithful logical reasoning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.373033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:ce62e0c9de4d4a859a1aa429145a26d1a098cc2a9d5e868e062ab9bcd19e5fec

Observation 98d0cef2-4776-484f-b8a2-85fccda41788 · outbound

This paper cites Unidepthv2: Universal monocular metric depth estimation made simpler.IEEE Transactions on Pattern Analysis and Machine Intelligence.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Unidepthv2: Universal monocular metric depth estimation made simpler.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.378698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:a10e9f8f2a8467628cc438c3a32616ec06fa2308896103cf3ea34812f7332a2a

Observation 83069b3c-c50e-4ca7-8f80-c90187e3cb5b · outbound

This paper cites Qwen3.6-Flash.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Qwen3.6-Flash

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.405392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:c75426afe481253538735f0bd8464d5f0eaca1f67be30d655b77113d0351fafa

Observation 989ca4d5-d57c-459d-8170-0e9177981012 · outbound

This paper cites Grounded Reinforcement Learning for Visual Reasoning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Grounded Reinforcement Learning for Visual Reasoning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.700678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:d8522b5214b65b7b950b8bc24bf29f7cbe7690861f84a478948b50f7b6a2df26

Observation 7d642168-4f02-494c-a833-a484ff0b14c5 · outbound

This paper cites Seed-2.0-Pro.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Seed-2.0-Pro

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.407252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:e63ac8d474429ef0b32ba44fe7bd2b4184d68a24729c9c49d112af9d7cc0c37b

Observation 0b9e7535-b4a5-432d-b23c-bbc92082924a · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.Advances in Neural Information Processing Systems, 36:38154–38180.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.Advances in Neural Information Processing Systems, 36:38154–38180

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.413355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:18f4b07d77c11ecf002e26dd0abc9b637e58e76618bac968b15ba160c5e1a63c

Observation 48c9885d-82e3-4d9b-8623-238a7677be5b · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.716648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:a9fb8a82b967a674c67c5d41246fc60adde7aa741beacf80c445a82b62c16b6d

Observation 33c71e8f-7213-4dbf-ac02-e8eddbf7b81b · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.694300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:980590cc4f76caf2cdfbbd3b402153672543860c3eb83e4dde63569f40fb53aa

Observation 17590dce-3ff6-4194-a235-81d871965e58 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Vipergpt: Visual inference via python execution for reasoning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.388244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:ad7b09ef2a02e37a1555b1c987cc74ddb5ec79e42d0fdc193fe3b7481bf2c16f

Observation efe77d19-7766-48c6-858b-690ce6dc8aa5 · outbound

This paper cites LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs LAST: Leveraging Tools as Hints to Enhance Spatial Reasoning for Multimodal Large Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.710072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:c05682d4d461e85c271bfe1f6fd2adfa7c715aad393941dfe5190f560abc536b

Observation 12b23ed4-f774-4e6d-b7c9-8ff1d397f9e6 · outbound

This paper cites Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning.Advances in Neural Information Processing Systems.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Pixel reasoner: Incentivizing pixel-space reasoning with curiosity-driven reinforcement learning.Advances in Neural Information Processing Systems

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.401468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:c460b9f80d856c0e795f3a293d4816108e60a4435d882332d810ed040949e698

Observation 65968de8-2794-4ad3-925d-6d7079370425 · outbound

This paper cites Last: Learning to think in space and time for generalist vision-language models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Last: Learning to think in space and time for generalist vision-language models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:33:05.725596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:5c84bcfdbb1adcfe6f780d628d2a6cf72f96d2ea88bd145eaa71d8ed7251018f

Observation 58c9da7f-4ac8-46cd-a0fb-e43ee5de1961 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:33:05.691342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:322752de48ae8e5bcf2f6eabb85bfdcdece3970289d4bab069ca6e92db7da83c

Observation 9cbe3419-19ab-4ab9-a1b2-2551a758edba · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.707113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:207d12ca9ac3c500dab864633f6f25909cde233d3aeca875de3f56d3266bbf1a

Observation 359663de-b04e-48d5-b469-a1c93fa6e64d · outbound

This paper cites Spatial-mllm: Boosting mllm capabilities in visual-based spatial intelligence.Advances in Neural Information Processing Systems.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Spatial-mllm: Boosting mllm capabilities in visual-based spatial intelligence.Advances in Neural Information Processing Systems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.384334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:f8c644f6f9102fe984344f42fada2eedaca795eb3ac82d0fc8879d0a77448460

Observation 84b7596f-66be-4b32-8fcd-933260d174e6 · outbound

This paper cites Qwen3 Technical Report.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Qwen3 Technical Report

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.728593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:11ba854b7dda4bfaded798d4680acd0a30b6741bc2860a5ae6ffc8e2de1cb49d

Observation 9667e7f5-7210-430a-8bee-ac98dc17ba79 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.417707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:a1f985e9eeb6f631de0b77b2da870579e6e077b10656987306ffe6df051f8cb1

Observation 78c4f4fa-be29-405c-aef0-63b94960db16 · outbound

This paper cites arXiv:2509.18905 (2025) 6, 9, 17.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs arXiv:2509.18905 (2025) 6, 9, 17

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T06:33:05.741795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:c7624529bf03a116a4cd0f18af869c9a98f8972b07b4f0122fe8be50a19e0420

Observation a5ed18cd-046a-494d-8671-70d03bf3b99c · outbound

This paper cites On the generalization capacities of mllms for spatial intelligence.International Conference on Learning Representations.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs On the generalization capacities of mllms for spatial intelligence.International Conference on Learning Representations

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.409140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:678b583fc395328239301dcd8ecac00c38ac28daa3289ebda15cf6dd49ad25cf

Observation 24e6a5ef-c728-4afd-8b6d-fb7e93100324 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.Advances in Neural Information Processing Systems.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs From flatland to space: Teaching vision-language models to perceive and reason in 3d.Advances in Neural Information Processing Systems

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.369425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:5d365c709da134c8990d80cc88f75080a08b643a960470f9c3f4b6b72630a683

Observation 3fc31a3f-9c9c-41da-8174-258d9a3343d5 · outbound

This paper cites Thyme: Think beyond images.International Conference on Learning Representations.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Thyme: Think beyond images.International Conference on Learning Representations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.382236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:550cbb86f18f69ce11444d9f913a08525e03a346efa8abb019b4eb8a2212f137

Observation 674fa5ca-92d5-4974-b4eb-1c5e01aec9e1 · outbound

This paper cites Think3d: Thinking with space for spatial reasoning.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Think3d: Thinking with space for spatial reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:33:05.697703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:3d591c1f18a4f79c45db3e4e215da8c4576f6a483c58df121cbb856f2adfa033

Observation 75b8135d-50bf-4af4-af68-eec2164ce0a2 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.Advances in Neural Information Processing Systems.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.Advances in Neural Information Processing Systems

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.390191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:7c26761d946c43c1a9bac992081290931bf905401c1c0ddadf38b1a4ea3f875f

Observation 4061159a-c108-4a67-831a-a54657811161 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.386406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:509848cb6a6d5d3b9b2366a537dbeffe89bd4e6dd2cd0b6a757b9167dda5af43

Observation 4bd15e3c-f212-4fe4-b767-5bd536ac8bf9 · outbound

This paper cites thinking with images.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs thinking with images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.391842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:63c9a0640177df8818631bdbfef8aca50d28865369b082535894861fbcab1a33

Observation 6da640a0-7c45-405c-8d7c-037d84899326 · outbound

This paper cites Roborefer: Towards spatial referring with reasoning in vision- language models for robotics.Advances in Neural Information Processing Systems.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Roborefer: Towards spatial referring with reasoning in vision- language models for robotics.Advances in Neural Information Processing Systems

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T06:33:06.393740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:61af86c773d51d0dc7c18d3501a3553a3033123586be459b6054c598a0f91ad7

Observation 3b664314-82e5-4238-843b-8a7b39bcb08b · outbound

This paper cites Reinforced Visual Perception with Tools.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs Reinforced Visual Perception with Tools

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:33:05.722823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:2cec7ea205c0f548892ac3dd6c4b40941967eedecc46a295985ac816c346bb17

Pith citing papers

Observation 7af2ca59-638c-4623-b582-eed7b3e31756 · inbound

3D-Aware VLMs with Implicit and Explicit Geometries cites this paper.

3D-Aware VLMs with Implicit and Explicit Geometries Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T07:04:38.497190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:04:38.497190Z digest=sha256:7ab51f77ccf61522d91b65a4a5d1bb1856c0e07fa361d111c05f24db5ce13c4a