Pith. sign in

Paper Citation Record · LEDGER

Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 47 inbound Pith citation observations for arXiv:2209.14610.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2209.14610 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 47 of 47 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:24:08.397621Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T17:41:17.457873Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ed20a350-dee8-46d9-af3b-1970e308611a · inbound

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks cites this paper.

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T16:48:28.089252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T16:48:27.918334Z digest=sha256:675a305d865a02a36a0de2fb942cae196ae56c2f6198e91dc3ebbe689a3f5081

Observation 78cad326-9ff8-4673-99e7-855a21908352 · inbound

Multimodal Chain-of-Thought Reasoning in Language Models cites this paper.

Multimodal Chain-of-Thought Reasoning in Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:12:27.523797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T18:12:27.396607Z digest=sha256:a23f698d2adf367e81283f000b818caeda321d87c641946b231e9b48b45692d5

Observation 1030b891-771b-4678-b98c-b8c96b3c4b84 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:06.336163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:ba6180d6c94caa6d106cc6f52e3449ccdcc664ab5cfc23fd0eff3b3baaf62185

Observation 91e0f16f-d3b7-4386-98bc-8395f3a1f7f8 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T20:58:59.136825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:99fee3e298224ee8b8f5d781cf26f7d43f9f98ab5888c26803db800b8c668e88

Observation 728c5717-d3d1-43d0-b4dd-8ebdd0893478 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T10:46:28.848496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:540b053e2d2af9e4338b997023e585f03b1b6ca52b638f834961e519c4521442

Observation 19f63a1e-4256-4725-8fb2-444834dd1666 · inbound

Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment cites this paper.

Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T17:45:06.323743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:45:06.323743Z digest=sha256:5a98d02a33e46ac432754ac9dc5b4761a0203f40c9b1c20ca9646d09255dec32

Observation 1967235b-fc79-49f4-bb7a-73e95f8b9a92 · inbound

LLaSA: Large Language and Structured Data Assistant cites this paper.

LLaSA: Large Language and Structured Data Assistant Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:24:49.855253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:24:49.855253Z digest=sha256:fd1b55a2c5d36ff0c1b6ddf32149318148f56fe22f21adb04e5e75a27ce8d4b5

Observation e228ef7e-5dc0-4e3c-b8d8-009d12d57122 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 166

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:57.777515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:3dcf8fd82a2f789793df6b48dbd6ce1eb485267e2e69726e9335d0699346c838

Observation 7eaa812c-b129-4d95-9dda-f3f94b1b37af · inbound

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models cites this paper.

PVC: Progressive Visual Token Compression for Unified Image and Video Processing in Large Vision-Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:42.538842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:42.538842Z digest=sha256:da1d0343c13cb1f2c448b9236c292e9a428b0078585a36d0e97374e79fa3d887

Observation f716cc29-6023-4004-ac70-5c53ffd6d5bd · inbound

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding cites this paper.

HoVLE: Unleashing the Power of Monolithic Vision-Language Models with Holistic Vision-Language Embedding Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T10:49:08.214747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:49:08.214747Z digest=sha256:e9f501d10dd2a380aa9a50fe7b3f9a11c7f62da9a8c3a8b10bf7bb631126f134

Observation 602a996e-fd75-43f3-ae7e-7f3003a15456 · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.071219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.071219Z digest=sha256:b6f14f11b5eeb8d050864bba400b011aa375fba4277b52e636b4d3b29e68374d

Observation b0a9aab9-bf2a-4992-bba8-1c11427c7e97 · inbound

CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning cites this paper.

CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T17:27:43.099255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:27:43.099255Z digest=sha256:6f64f6d64b65a578055dbff716add7d018a2840e482caeb90a33029db950f6b7

Observation 219b968c-698f-47ae-aa01-efe93dbd34e2 · inbound

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model cites this paper.

InternLM-XComposer2.5-Reward: A Simple Yet Effective Multi-Modal Reward Model Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T17:18:40.590732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:18:40.590732Z digest=sha256:89a33b51b2660bfa43956030da3ed5ab478b22efb82fbe25abb4221ae9c99f1e

Observation 691127aa-01a8-4bd5-975d-89fc8eb190a5 · inbound

What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects cites this paper.

What Really Matters for Table LLMs? A Meta-Evaluation of Model and Data Effects Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:56:30.878124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:56:30.878124Z digest=sha256:74f7467f5d11cf70a6d7f06be911a50ce7ff97786767777c97cd5aa1cf49de14

Observation 261cf6af-5fae-4493-a957-9e769eec7bf1 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.155023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.155023Z digest=sha256:8bd59d886190ccb83d390feab61b84df017270a04eaa637c495b5782071a4a08

Observation a1d1e027-957c-42bf-94e0-f94604204e11 · inbound

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types cites this paper.

TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:36.638524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:06:36.638524Z digest=sha256:435bbb406de998f19ff5ed7a702668c206b3ea1f572a82b57ad95257852713ef

Observation c46353c7-91bd-440e-9d6e-04019cfad104 · inbound

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL cites this paper.

LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:15:46.298494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T15:15:46.255296Z digest=sha256:3ee75f4cc98c966a4c54fd01ea3f22d005f59c7037668c00ae7bb2c85f90909e

Observation f6784cd3-7111-4c36-9b89-9954220371a4 · inbound

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding cites this paper.

FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T19:52:01.892384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-22T19:49:00.961388Z digest=sha256:7aa5611cf6036638845fc41d6a7f215dba8f42cc38c515ad5643a9870a8629ea

Observation e44e2bb2-2481-4e40-a0b9-7802a9de8efe · inbound

Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering cites this paper.

Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.067596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:25.067596Z digest=sha256:1c6ac9ee838156cdd4390bc5d80ed6b28375b1c552ca63f1a58d867ef34553c7

Observation 563f274e-f6b0-4a07-967c-6cc9a724897b · inbound

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning cites this paper.

Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:04.652599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:04.652599Z digest=sha256:efbbb45d0280c312a5670aa4562285ebad3c0e8a714c0f48881c7e24805835c3

Observation 779f9fba-6b85-4f38-a0b7-12883815d016 · inbound

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding cites this paper.

MMTABREAL: Real-World Benchmark for Multimodal Table Understanding Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:19.359442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:19.359442Z digest=sha256:eff15165fdcb7651649970ab10773e49ebaaaa799bffe2eefd397f3f0be437c2

Observation c06c8f85-6815-4e2f-a344-637fcd7bc9a1 · inbound

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start cites this paper.

Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:12:54.957538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:12:54.957538Z digest=sha256:d7ec21735bdb9b7783d6e406ba97dc51ac32140805ee5c11090d22af124fcd6d

Observation 3f10de56-a0b1-4c17-ae11-78cfbcbc2cd5 · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:00.575535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:54:00.575535Z digest=sha256:797694282c46194f6f76af5e8de5ad3ebe0f7d0038e150bb8fa3011f74c2bc2f

Observation 4f30c604-bac2-41b0-9774-f0e2832b62cd · inbound

CoMemo: LVLMs Need Image Context with Image Memory cites this paper.

CoMemo: LVLMs Need Image Context with Image Memory Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:30.357233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:02:30.357233Z digest=sha256:4b6b63d95b8e063926b4549427df9460c005af7ecf4295e02b4b471e8b970413

Observation d29d96b1-1b6e-4443-a13f-25d9904f8dcb · inbound

MarginSel : Max-Margin Demonstration Selection for LLMs cites this paper.

MarginSel : Max-Margin Demonstration Selection for LLMs Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:12.917956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:12.917956Z digest=sha256:b28971f8fff2b9b0db7d70313aa5126e0b7135ba70e2c9e55336d16983cb3043

Observation 1c1c7cfb-783d-4f56-a472-3b1a2373e1b6 · inbound

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design cites this paper.

FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:45.415945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:45.415945Z digest=sha256:407999739c9914660e62b9fb4238dda46dd5e64f2a0206a6b11be7f407661dd9

Observation 1ca9bdcc-28dc-4284-918d-cdca4ea567c6 · inbound

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning cites this paper.

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:36.733960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:36.733960Z digest=sha256:3f2bb14481acab62e7e700d6037d46b09b90fd76d36aa8a2cded3f850f1b2fd5

Observation f8b413f7-7d5a-45e3-8655-2a7e0c828785 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:03.783517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:03.783517Z digest=sha256:73471252d4a150c6c76bec408fcef5425dbad211baaa1778a031c50260698223

Observation 319aec87-0af3-4f86-bfe8-be9ab2818488 · inbound

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning cites this paper.

MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:18:36.719126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:18:36.719126Z digest=sha256:c549126adf6a72fbef4abfba3d41e6e6e00abdd8082e481909cc28e37576c0c8

Observation 6ae840bb-6856-43d5-b93b-3e4b89998ffd · inbound

A Compute-Matched Re-Evaluation of TroVE on MATH cites this paper.

A Compute-Matched Re-Evaluation of TroVE on MATH Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:05:19.206563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:05:19.206563Z digest=sha256:cae3354a2a43e47f85f69ec194e1e1208fc87b82fca8b2424a2cfec652b9d0e7

Observation 96109899-f9ff-431b-b58f-a0ed11b218ed · inbound

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation cites this paper.

Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T17:24:08.397621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:24:08.397621Z digest=sha256:4e7ef94d13c8341a36c9158d1107489a29c3c042c5e3ffd4013d3bbc793a81f2

Observation d6f8beef-a452-4e5f-8a3f-d6c6b9153366 · inbound

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning cites this paper.

TableMind: An Autonomous Programmatic Agent for Tool-Augmented Table Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:56:52.574532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:56:52.574532Z digest=sha256:b2a7fc5481aa45b93c83f51230564304e9042bd9e24874ff4e0f786510bc8ecb

Observation 13894834-534b-453e-8427-7c16af940138 · inbound

Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation cites this paper.

Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T09:21:10.687359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T09:17:00.389716Z digest=sha256:503d8305dda460cb9d8ec1a6f0371b6f49bc42914194b1014bd1f0cf9f482f30

Observation 5e68657e-cbf9-4d1b-8731-2699fe041a04 · inbound

From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs cites this paper.

From Implicit to Explicit: Token-Efficient Logical Supervision for Mathematical Reasoning in LLMs Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:11:08.204801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T17:09:58.963561Z digest=sha256:5f2d27c2326db1e9c65f51eeb2217cd2759faa71fcb16955d394c7c54c8b105d

Observation 0f778b08-1513-4b56-9814-1927e1be69d3 · inbound

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning cites this paper.

From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T06:52:53.945286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T06:52:53.945286Z digest=sha256:96c15d1ed3fafeb43c0cb8d37a6a30ed83dfb1413e9ee7ebf0899690a7cf6aa7

Observation 0fbc24ed-2b5b-41c0-9fa1-275302ea34a7 · inbound

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment cites this paper.

PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T19:57:32.329557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:57:32.329557Z digest=sha256:dfb2cc87be84ddc3ead0eb779d3cc339ae31951390ab3c680af33a2682d9a640

Observation 9fd8fad2-2019-49a6-9e9d-17cf251d7597 · inbound

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning cites this paper.

Self-Consistency from Only Two Samples: CoT-PoT Ensembling for Efficient LLM Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T17:41:17.460571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-05T17:40:18.030591Z digest=sha256:5771be256352950b45059d8a398ef88b8cbc227db2d5232b88aa3281a7b5285d

Observation e873eadc-c41f-4092-8528-5a7a74e87e95 · inbound

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization cites this paper.

V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:01:04.237102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-09T23:48:32.613988Z digest=sha256:3e72dd946988862e75e7f5e9c33fafb9096bdba50a780e33316d72f3497e1c12

Observation 4cb5e9e7-0f80-4054-90b7-8b11f5b14ca7 · inbound

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild cites this paper.

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:41:24.776125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-09T19:28:54.531339Z digest=sha256:e7db1364a020b46d2538644f1c75008e1f16d4f348d00b99c25d0ae1cbdb2dbe

Observation 60c70ffd-6c9f-4b66-9502-b67ca9d6ef01 · inbound

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild cites this paper.

WildTableBench: Benchmarking Multimodal Foundation Models on Table Understanding In the Wild Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:10:24.092184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-25T06:07:28.392100Z digest=sha256:f956351848636c9de2b168aaff3e343f63cfadf4c4f754cc1b66c037c67764ff

Observation b26d3fdd-c3ec-4262-a7b3-20a15995e983 · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:23.486623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:c8b5bf92ec0269ec5da38a2436fdda9dad69fb3a67b072285df9eea5fe861cc2

Observation 9d902a10-81d4-448a-89b2-3ffd33bc34d9 · inbound

Zamba2-VL Technical Report cites this paper.

Zamba2-VL Technical Report Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 123

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T19:26:00.662647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T22:34:20.970856Z digest=sha256:863cf54edfdeae53f8e56ddab7fc46731b46133cf3658c1836f984019f5d9e92

Observation 03258f5a-a419-4f52-962e-1e01d6164398 · inbound

On the Generalization Gap in Self-Evolving Language Model Reasoning cites this paper.

On the Generalization Gap in Self-Evolving Language Model Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:24.934086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T17:18:37.671369Z digest=sha256:dab067633c7d0ad42d3940c19ce31c84a5dbd6f6c445f1ba29f349c554e3032c

Observation 383af810-4417-49a0-baf6-34ad11371b6c · inbound

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks cites this paper.

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T04:06:35.186296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-28T09:27:30.923556Z digest=sha256:b76a9122e4542d02ae0246077d22de8b1b9de1eb9a3871be242e724b93fb11b2

Observation 2d89b93a-b6d1-4900-a27c-3c96111fac85 · inbound

From Context-Aware to Conflict-Aware: Generalizing Contrastive Decoding for Knowledge Conflict in LLMs cites this paper.

From Context-Aware to Conflict-Aware: Generalizing Contrastive Decoding for Knowledge Conflict in LLMs Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.341911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:38:31.421486Z digest=sha256:116a56fed223c24c3dce5339c133c84343a6bce7b5c7cdc0e8bbd29cac55c2bd

Observation 9cff557a-1171-4ee2-80d1-1f1ba5084f87 · inbound

MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning cites this paper.

MoCA-Agent: A Market-of-Claims Code Agent for Financial and Numerical Reasoning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:57:56.203277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T10:17:42.831128Z digest=sha256:8b14ad6bf6004c30b7e6a3737e44f90eb02239d502c1894aef7cb0312d67fa61

Observation bf856bad-8392-4af0-9dc4-fc151234e38d · inbound

Adapting Generalist Robot Policies with Semantic Reinforcement Learning cites this paper.

Adapting Generalist Robot Policies with Semantic Reinforcement Learning Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:45:42.737459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-07-01T05:09:29.625066Z digest=sha256:11cd6361c053a51fbe60543449291cb545e625151359f02c9fab375bb78011f2