Pith. sign in

Paper Citation Record · LEDGER

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance

As of 14 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 2 inbound Pith citation observations for arXiv:2510.08048.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.08048 v4

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:53:12.499757Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-15T19:27:17.802597Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-10T12:15:01.137692Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved62
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4776558-b14b-4346-a29c-313ad2cf4861 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:06.813916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:06.813916Z digest=sha256:9b180e99af2b456aa70e556d6d5cbcb6aa4dd0f2f0ed23c73f72d956fa68e150

Observation 17158c38-5797-4a8f-b40e-c7524912d7e9 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:06.876600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:06.876600Z digest=sha256:466cdd4ba193d22a330d4ee9a6cca9d0602fe1674bb4cdcf5bd88207b82f1266

Observation 14f14b6d-bbf2-4d64-ab39-0715baae8e00 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:06.913693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:06.913693Z digest=sha256:d5883fe5ea3a47c9a56f4e1f779e0a630e77fbbd1562f81f3e714fb27d92f234

Observation 6ea9a0cc-9ee5-4cb4-bf21-5dc92a896cab · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:06.966825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:06.966825Z digest=sha256:ed517fc08da97b6de552e227eb91ec7c2538407f59e3a883566dd7660e7afb0e

Observation 19ee9732-5b7e-412e-bce6-4021e1b15ebf · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.048784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.048784Z digest=sha256:86de2073084e6e6f49df4e9dcd026b1b7d2a6a61e48351cc87940a8e442f2353

Observation 8f0edeb4-5518-4f7c-ad17-b3878085ebe7 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.123147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.123147Z digest=sha256:440b32fda9685db7a7762045e8e2399b93b24ce9eb63da360892e6ae40f90de7

Observation d4e1c867-3bf8-4f43-ab92-765e2712d7a1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.232318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.232318Z digest=sha256:010ffe4f48a5a640449cddcfb04b78e73629254d28ec9455b2d338be016db44d

Observation dbadd8c7-f774-4d60-bddd-95d24b3e29bb · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.376532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.376532Z digest=sha256:78572cc57f8f5de02c4fa3d5d5bb9496dfa63c0aea0934b5f83e962b189df076

Observation 66057b99-1e90-4006-9ef7-f250c3b6ae66 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-04T10:53:07.521983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.521983Z digest=sha256:923efa542b0ed690974e93c77fc4472838e00abdc3bfdc2a088f674074c4cc89

Observation c486bdb2-97f6-43d8-9f28-b573a0b765de · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.624556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.624556Z digest=sha256:6ca2f63f4b4ad2778f3dd0a1e4d5645154b764785ad863c46fb29e456d85c419

Observation 74609662-e9ec-4c31-9b32-d63db4327337 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.720624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.720624Z digest=sha256:8c24d6d835eb533990fccc7d661afb08fa87a68e63850e39969fd24cbba043f5

Observation d63d37f7-3e97-45d9-8c41-6a220bdba68b · outbound

This paper cites The Llama 3 Herd of Models.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.864133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.864133Z digest=sha256:f5ffa88328fc124500f074680139fd1b2b8a752acd66b003228b74b8f3f5bedb

Observation b40857a2-b4f1-4ddf-88c5-fe132cf4ca19 · outbound

This paper cites A Survey on LLM-as-a-Judge.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance A Survey on LLM-as-a-Judge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.943286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.943286Z digest=sha256:f36c3674feb5bafd789b600110d3c54f58ec4a5601c8c2b5f14e1095c14ed134

Observation 7c481f30-7042-4931-9867-16411b6c3d47 · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Teaching Large Language Models to Reason with Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.022212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.022212Z digest=sha256:4c2042db9f6f5555f5929984ddf157e5bd49c787b28ca462c93387abb2a2212e

Observation f0a69c2f-ff18-4b32-8a53-0770e230e8fe · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.114591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.114591Z digest=sha256:7c12743cfbfac6c1f92a081c055801714991d2845af33867a058a0f700afec35

Observation d40c7261-d733-4f21-b951-b6c3b7de3e0c · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.181854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.181854Z digest=sha256:aca0306f52caef0cd3a5d982c7123c77f24a03adc78e6b888ca011a690a3e728

Observation 180f8dfb-2cea-4771-a861-afbe7e63938b · outbound

This paper cites Qwen2.5-Coder Technical Report.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Qwen2.5-Coder Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.220029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.220029Z digest=sha256:4a42464e711fe9af300fcc2f95fc87985a47adea795e66fac53b0adc5efbd5c0

Observation 6175ae2b-d845-479c-b20a-350fee2eb1da · outbound

This paper cites A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance A Study of Plasticity Loss in On-Policy Deep Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.287562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.287562Z digest=sha256:22c46563c87ee6d00100a92c887360a7f8ee4e446f146a225c1db9b2a416b99d

Observation 4f94b7ac-1717-4bf1-99c2-0a696b6575f8 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.355132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.355132Z digest=sha256:89566d414f109dc664a91020814727a65874071f5e7227696a30370776f6b3d9

Observation 188aae36-e751-45d8-9c83-e3f9a25cde77 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.424629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.424629Z digest=sha256:5316148e13bc2264fd280653da257d481da1e29f41fd60a202325d449bfe533a

Observation 6a4c5a46-318a-46b7-b102-08b4c8236c83 · outbound

This paper cites Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.542610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.542610Z digest=sha256:865143c520e98f6d0d31d80b2b33e95328f253c2de51c952e3326c454d802178

Observation a32544b8-85f4-4acc-adbb-b4baf62aa0ad · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.603311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.603311Z digest=sha256:89cb66da2196504fd47343adea665644fdb2bb9f9593bb946ed9e7d96005f4bc

Observation 11411b44-0b68-42f9-aeb6-eab51a369da2 · outbound

This paper cites Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Adaptive Guidance Accelerates Reinforcement Learning of Reasoning Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.661292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.661292Z digest=sha256:3aff7685c267bb6b1df537e519fe723ad1f64793cd0abb13ba22584ab62a2538

Observation 9d902ca6-1795-492b-99ce-79f74f15bf14 · outbound

This paper cites Passage Re-ranking with BERT.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Passage Re-ranking with BERT

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.718998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.718998Z digest=sha256:2e853e3d6473d2b89ab140e0e58982c4148e903ff840063bd1fd7b61f8762530

Observation e9650e6d-1eac-4f64-9e69-2b5cda75e202 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.788211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.788211Z digest=sha256:3c3448b6ebcd3d12eb482276400fd17679e95933cf586e0f54d9acce9a88eb2c

Observation 292be729-1e32-40da-9838-6f0a080823f7 · outbound

This paper cites OpenAI o1 System Card.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance OpenAI o1 System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.866608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.866608Z digest=sha256:a73ac704056745ac3aa3a0a1ff770a26acf210416f6b1bf0af4e888f9541741f

Observation 07f689de-5442-481b-b2ce-fab67f6074f1 · outbound

This paper cites LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance LLM-Driven Intrinsic Motivation for Sparse Reward Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.945022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.945022Z digest=sha256:6b740fa3d695eb3a893106b859f6e18fe5ff560bc143c0bce363542f36109a46

Observation 16b7bb27-178a-4e87-bbb9-8aca8f4efa76 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.015333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.015333Z digest=sha256:a9da95543d75a0c78d6fd4c20aa0faf7145058b0c5adc6c2ef9080faa56a015c

Observation 2055a67a-652b-4683-a375-b6c463c0399d · outbound

This paper cites Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Learning to Plan & Reason for Evaluation with Thinking-LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.093199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.093199Z digest=sha256:3673a0388e717ae58d121f704ac33c60ae1cbc974e0bfdfeab3bab3676d677c8

Observation 6047bd7e-680c-478d-9606-c39ef50e21c0 · outbound

This paper cites Seed-Coder: Let the Code Model Curate Data for Itself.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Seed-Coder: Let the Code Model Curate Data for Itself

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.168062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.168062Z digest=sha256:ace860bfe6e744d83e0f1409442f32903657712e26ef7f22e0ac05305971ea8b

Observation e9a35561-696a-49d8-953a-0e4456d84efd · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.237921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.237921Z digest=sha256:4f434026f8bc5f6774d010b11491bfa11fc29a2a7e61ed1398dd086dc55ab60f

Observation 190731fb-787c-4547-aea5-b2772cdd64d4 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.317845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.317845Z digest=sha256:f7a2423e8aa9348e894d9ed71f45c8bfeb62124c035d84f3390e1a80e28c6288

Observation 7c50dd99-db9d-4b5c-a77a-c5f41bbaf12d · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.467665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.467665Z digest=sha256:c6aa4969c89fd5fabb5de68465cd6b59ab9abe87a0c1fad468de349f0620a94a

Observation 3b27a105-e121-4ea0-b3ca-cdedb13ce357 · outbound

This paper cites Svore and Christopher J.C.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Svore and Christopher J.C

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.563069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.563069Z digest=sha256:025a18c7525298168bdf23ef6e6cd92155ef911c2dd202404e889640f1bd66f3

Observation 4bd72136-490c-47fc-8d15-5f0bcba62315 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.610992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.610992Z digest=sha256:70a5f12c9822de35a0589e282378a10e15cbcda8b77077f72dfd93f3e4672417

Observation f535e962-0cea-4939-880a-c2b5915eb6ad · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.647203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.647203Z digest=sha256:dbd8de74d95caff5a3d89a687546e6b0179a9ecb7804cb5f9dc0639d43de0e6c

Observation 1908b08b-7257-48ef-86a5-d1e67b8ac8e7 · outbound

This paper cites Attention Is All You Need.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Attention Is All You Need

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.715085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.715085Z digest=sha256:6c00afc91b84a36af3fd761f58c90f9d3dde7f9cb6e01fc6242fbb972a76a06a

Observation ff131638-77fc-41eb-bd47-63dfda5b131e · outbound

This paper cites Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.781696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.781696Z digest=sha256:735bfb7f6fc18103780d70b1a9811b97f16734d6bba47ef807d0366a2b1c178b

Observation 03bf37e0-7be3-4c0b-94c2-d2a2857d85c4 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.842462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.842462Z digest=sha256:1ac2a58f1587bc2db34cf91cf6bf613c9dfe1c7a3f7195b4db185a98344e0a16

Observation fede5782-71b7-4a70-bb3d-1a35c1d150d9 · outbound

This paper cites Qwen3 Technical Report.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.989206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.989206Z digest=sha256:b3d2debfb2f9075f9e9f96d054b97f1148eb2d87d8339472928de22094d7ea3d

Observation 02c16df3-b86f-4b5d-a211-95fd29048e87 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.105189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.105189Z digest=sha256:492644b0946c4a3a893fef1e3fa48b3c22b2a1dcca530e977e91d3f1ec1feb34

Observation e9e39826-6a62-4149-a3fa-a5f00ddd7bec · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.259947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.259947Z digest=sha256:c4952a72eac6a7d79186601dc538ebaeaf98ab85ad1361a0dcd41f04b50e39d6

Observation b05400e5-2a05-404f-845b-3a0f5569c5a6 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.374837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.374837Z digest=sha256:2a560328449c7ae7a168616898818a87cc89eca722ef5e868b71397c0a1e8ce7

Observation 2cab5881-6639-4941-b133-9639d1c52f32 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.552125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.552125Z digest=sha256:dad243e8866635f429ef09fb5eb8fc8e57caa4f719ed071c2c68c302f01666f1

Observation 73cd5e14-6d18-4fe8-a88f-51686fe1fe90 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.151336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.151336Z digest=sha256:0eac0e852a7100b90c76540d8390ee25bd3fcf0a17f810ac794bc6afccf6c9f3

Observation 03f4f0ee-7b24-455c-aa94-43efc46b328b · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.776382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.776382Z digest=sha256:f3f6adec35066890dc0ea566905e936cbdaffd08e555c944dab236d066c695a0

Observation 27f1b512-88a9-4cbb-929a-98e9068fa0ce · outbound

This paper cites Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Critique-GRPO: Advancing LLM Reasoning with Natural Language and Numerical Feedback

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.877236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.877236Z digest=sha256:b61cb2fff193c8309c134c6166208669b05e6702aa6fcb8fcc513135a046b1e4

Observation 8602e035-1add-4b8e-b107-eb26c08f2193 · outbound

This paper cites LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance LLaMA-Berry: Pairwise Optimization for O1-like Olympiad-Level Mathematical Reasoning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.440907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.440907Z digest=sha256:9332e22e7df17b9c702e97c8bde652bbe651623e6ee004d13aa814f8a9edd994

Observation 63ed6416-f211-42d3-8ae8-6792164f8643 · outbound

This paper cites StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:10.631113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:10.631113Z digest=sha256:1de565fdd2b670434b669d262f73bf90a6ac3936ac5f04cb65c6f55aa27282e0

Observation 3960b980-434d-44fa-ad6a-b55140f46ade · outbound

This paper cites Undyed Cashmere Coat contains cashmere.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Undyed Cashmere Coat contains cashmere

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:11.229914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:11.229914Z digest=sha256:5ef014bbd60887ef330568767ad754f514e82a326d81f8127fe936bbd2911c1c

Observation 3b9933cb-e7dc-4c12-8f03-8891239a1170 · outbound

This paper cites Relevance label is 4-Excellent.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Relevance label is 4-Excellent

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:11.294586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:11.294586Z digest=sha256:b3b1079b811e2b4cafa298f0a3bd168d5d2ae0ab75cbddedb0cd3b167fa72e02

Observation 66931e4b-d4ef-4ecf-9f64-7b3201dd6bc6 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 59

Resolution
parse uncertain
no resolver link, observed 2026-08-04T10:53:11.468139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:11.468139Z digest=sha256:a91b2d9dd855ab08e1c59cdadbaab9aae882151bfbd6ba7c63af3c9eab71945e

Observation c008981b-b227-4a8f-9726-370fc4d806d5 · outbound

This paper cites Undyed Cashmere Coat contains cashmere, but the content is below 50%.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Undyed Cashmere Coat contains cashmere, but the content is below 50%

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:11.603766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:11.603766Z digest=sha256:75ea4572f83d7e8801306315170203201e0d64ad648c0568588128bf0f2ab11d

Observation ec4d19da-f532-4ab0-9cf6-b20cdc334408 · outbound

This paper cites chiffon dress.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance chiffon dress

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:11.682245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:11.682245Z digest=sha256:8ea31d182362364a8e2e0bda672a8ce1ab0971e1a9511a1d24925b0c2389eb0b

Observation 4e7cbcae-0a8d-4da7-984c-fd477b74c166 · outbound

This paper cites The dress contains chiffon, but the content is less than 50%.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance The dress contains chiffon, but the content is less than 50%

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.055879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.055879Z digest=sha256:e7a668f3a76912a75dc8a4ed61db10569ce65d97a93376a283aff512055dc1d9

Observation 88fa4c85-b240-4bee-a8e2-4f25f8b4bd2e · outbound

This paper cites Relevance label is 2-Mismatch.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Relevance label is 2-Mismatch

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.137741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.137741Z digest=sha256:13872b1afb28c4208b56ac34dd01f2d14627205faf78098502735ddb99862fb9

Observation 389a4d66-81e1-42dd-8fbd-a2b41bcf6121 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.209015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.209015Z digest=sha256:fab68a6e7d0f08bcc24a3ed432ea322e6606a272e184117920641baa053228cb

Observation ffdf0d9c-5fa7-4d01-84c4-3858b2e39e21 · outbound

This paper cites an unresolved cited work.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Unresolved cited work

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.280665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.280665Z digest=sha256:eb51e84ee5c62d80ed8125f693d77b0460184bda09b9b41ba90ff9b80d981620

Observation 8a0c056b-2e2c-4a50-9ffd-d517cf316c4d · outbound

This paper cites The conclusion is Excellent.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance The conclusion is Excellent

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.347899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.347899Z digest=sha256:a220dd41d958b1ae96d8c2b22c895b886d6d96fe29ee040117e5a1b7834bfd85

Observation ce67b22d-60d8-4dc8-8c6f-2c7132176819 · outbound

This paper cites The dress contains chiffon.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance The dress contains chiffon

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.424739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.424739Z digest=sha256:64c26c61c3a5447b244374f91f14ae56e91c50f540f1b659eb1d04eb8dc5c504

Observation f1253ef2-405a-426a-9426-d1c2c67fc721 · outbound

This paper cites Relevance label is 4-Excellent.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Relevance label is 4-Excellent

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:12.499757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:12.499757Z digest=sha256:14b915fb62642d886faa0b67cd997d91999a2421e097705f7a2d615ca740a8e2

Observation 05404bdc-9d67-4074-b0fe-25f965c19a8b · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Go-Explore: a New Approach for Hard-Exploration Problems

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:07.774986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:07.774986Z digest=sha256:2dddf33f96a4293f29d6094cb52002b033f32671c54c794f8becca007fbdc9b0

Observation c015ebd5-0bb5-4d99-86fc-21eb80bd959a · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:08.480368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:08.480368Z digest=sha256:f053de0ff608a4fe011cbc31ff497f58bfac4b54a57121866d1d1ebdf45f2d92

Observation 73dfea2d-def0-4b49-b3ad-a517d841eae6 · outbound

This paper cites Defining and Characterizing Reward Hacking.

TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance Defining and Characterizing Reward Hacking

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:53:09.398856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:53:09.398856Z digest=sha256:ab06aa06fad9d7abb22ad8e1402b1c29dcf17db2f7aebc2263a2cf0b99284bb1

Pith citing papers

Observation 1f86f5e7-f68a-4dca-baf8-2f190263c6a5 · inbound

Synthetic Data Powers Product Retrieval for Long-tail Knowledge-Intensive Queries in E-commerce Search cites this paper.

Synthetic Data Powers Product Retrieval for Long-tail Knowledge-Intensive Queries in E-commerce Search TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:12.927851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T19:27:17.802597Z digest=sha256:719550dc5bbd7c9d7a00ba24b267f844a8c1c535a8f70f2dd67b54955a8f9935

Observation 0087070d-441f-4858-89c9-78506279a086 · inbound

K-CARE: Knowledge-driven Symmetrical Contextual Anchoring and Analogical Prototype Reasoning for E-commerce Relevance cites this paper.

K-CARE: Knowledge-driven Symmetrical Contextual Anchoring and Analogical Prototype Reasoning for E-commerce Relevance TaoSR-AGRL: Adaptive Guided Reinforcement Learning Framework for E-commerce Search Relevance

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:12.927851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T15:28:44.678081Z digest=sha256:96564c45db519af4a972eb269b329b68af072806c6b3fd969a44187bb30098dc