{"work":{"id":"38c07273-5071-4bbb-b2af-067af11becc7","openalex_id":null,"doi":null,"arxiv_id":null,"raw_key":"raw:3bb7980b3cfdb40588267a5f","title":"online\" 'onlinestring :=","authors":null,"authors_text":"ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school","year":null,"venue":null,"abstract":null,"external_url":null,"cited_by_count":null,"metadata_source":"raw_reference","metadata_fetched_at":"2026-07-09T10:56:11.938856+00:00","pith_arxiv_id":null,"created_at":"2026-05-13T17:18:01.826474+00:00","updated_at":"2026-07-09T10:56:11.938856+00:00","title_quality_ok":false,"display_title":"online\" 'onlinestring :=","render_title":"online\" 'onlinestring :="},"hub":{"state":{"work_id":"38c07273-5071-4bbb-b2af-067af11becc7","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":171,"external_cited_by_count":null,"distinct_field_count":11,"first_pith_cited_at":"2024-04-02T17:49:40+00:00","last_pith_cited_at":"2026-05-24T19:27:20+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-20T22:59:22.491314+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":6},{"context_role":"dataset","n":2},{"context_role":"baseline","n":1},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":6},{"context_polarity":"use_dataset","n":2},{"context_polarity":"baseline","n":1},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"online\" 'onlinestring :=","claims":[{"claim_text":"2Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, P. R. China, 3Zhejiang University, Hangzhou, China, 4Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, guoyanl@zju.edu.cn Abstract Recent reinforcement learning (RL) ap- proaches have advanced radiology report gen- eration (RRG), yet two core limitations per- sist: (1) report-level rewards offer limited evidence-grounded guidance for clinical faith- fulness; and (2) current metho","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"within a single expression. To this end, we recruit human annotators to rate the figurative expressions from the MetFuse dataset. Each annotator was provided with a sample and asked to rate the fig- urative texts on a scale of 1 to 5, with the literal sentences as references. We used four criteria from Chakrabarty et al. (2021), in addition to one of ours: (1)Fluency(\"How fluent, grammatical, well formed and easy to understand are the generated utterances?\"), (2)Meaning(\"Are the input and the ou","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"vendors-we guard against this persuasion bias while simultaneously enabling the benchmark to simulate the full spectrum of real-world purchase intentions. We model scenario difficulty with five calibrated customer profiles (easy,medium,hard, very hard,adversarial). Each profile k defines two interpretable controls (Table 2): (i) prior buy propensity pk ∈[0,1] (0.8 foreasyto 0.05 for Figure 2: Overview of the SalesLLM Benchmark Script Generation Pipeline. The pipeline follows a hierarchical proce","claim_type":"dataset","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Multi-hop retrieval.LetC={d 1, . . . , d|C|}be a fixed passage corpus. Given a natural-language queryq, the task is to retrieve a ranked list ˆP⊆ C such that ˆPcovers as many passages in the gold support setG={g 1, g2} ⊆ Cas possible within the top-Kresults. We follow Guti 'errez et al. (2025) and Wang and Han (2025) in using R@5 =E q \" |G ∩ ˆP5| |G| # (1) as the evaluation metric, where ˆP5 is the top-5 re- trieved passages. Equation (1) rewards retrieving allgold passages and penalizes partial","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"anchors, the system traverses the cross-layer links Ecross to access specific archived Event Progres- sion Graphs stored in the archive set Sarch. These archived graphs constitute the raw source of the topic nodes. Formally, we define the candidate set C by aggregating all event nodes from the archival graphs linked to any node in the anchor set: C= [ u∈Vanchor n v∈ V(G ′) G′ ∈ Sarch, (u,G ′)∈ E cross o , (10) where V(G ′) denotes the node set of an archived graph G′. This strategy ensures that ","claim_type":"baseline","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"grows linearly with the number of steps, optimizing πθ for complex tasks becomes increasingly chal- lenging due to the accumulation of history. 2.2 Agentic Reinforcement Learning Agentic RL typically adopts policy-gradient meth- ods to optimize the agent policy πθ. We formulate the agentic RL training objective as: max πθ Ex∼D,H∼πθ(·|x) [rϕ(x,H)]−βD KL [πθ(·)∥πref(·)] (1) where πθ represents the policy LLM, πref is the reference LLM, rϕ and DKL denotes the reward function and KL divergence respe","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks online\" 'onlinestring := because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (6 contexts).","role_counts":[{"n":6,"context_role":"background"},{"n":2,"context_role":"dataset"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-20T06:31:48.962570+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"5e3d68df-d072-4ce9-bb1b-53ac1ed495cc","orcid":null,"display_name":"ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school"}]},"error":null,"updated_at":"2026-05-20T06:31:49.516894+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-17T10:39:51.003293+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"write newline","work_id":"b6ace3ad-bd44-4a95-86ee-c7b73bd4cda2","shared_citers":46},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":14},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":7},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":6},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":5},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":5},{"title":"Qwen2 Technical Report","work_id":"a1857881-ab9b-4b80-9b5f-9ae4b5c2566d","shared_citers":5},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":5},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":5},{"title":"Bleu: a method for automatic evaluation of machine translation","work_id":"7a0e7f56-7c92-470e-b7bf-2e9974bf9a93","shared_citers":4},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":4},{"title":null,"work_id":"d715ff4d-f31d-49cf-bf54-8217373c5c82","shared_citers":4},{"title":"BERTScore: Evaluating Text Generation with BERT","work_id":"9eaaaac1-0a96-4f5f-9b13-30c46e9e1346","shared_citers":3},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":3},{"title":"doi: 10.18653/v1/N18-1101","work_id":"9737cdf0-fd48-4485-ab04-c8ec9386e782","shared_citers":3},{"title":"Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D., and M \\`a rquez, L","work_id":"11bfc949-547c-40f3-a86d-953eb9b2154c","shared_citers":3},{"title":"HuggingFace's Transformers: State-of-the-art Natural Language Processing","work_id":"9d86da8d-01d3-41af-a0d2-ee14897927a9","shared_citers":3}],"time_series":[{"n":11,"year":2025},{"n":36,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-17T10:39:51.030543+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-17T10:39:48.350149+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"online\" 'onlinestring :=","claims":[{"claim_text":"2Key Laboratory of Smart Manufacturing in Energy Chemical Process, Ministry of Education, P. R. China, 3Zhejiang University, Hangzhou, China, 4Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security, guoyanl@zju.edu.cn Abstract Recent reinforcement learning (RL) ap- proaches have advanced radiology report gen- eration (RRG), yet two core limitations per- sist: (1) report-level rewards offer limited evidence-grounded guidance for clinical faith- fulness; and (2) current metho","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"within a single expression. To this end, we recruit human annotators to rate the figurative expressions from the MetFuse dataset. Each annotator was provided with a sample and asked to rate the fig- urative texts on a scale of 1 to 5, with the literal sentences as references. We used four criteria from Chakrabarty et al. (2021), in addition to one of ours: (1)Fluency(\"How fluent, grammatical, well formed and easy to understand are the generated utterances?\"), (2)Meaning(\"Are the input and the ou","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"vendors-we guard against this persuasion bias while simultaneously enabling the benchmark to simulate the full spectrum of real-world purchase intentions. We model scenario difficulty with five calibrated customer profiles (easy,medium,hard, very hard,adversarial). Each profile k defines two interpretable controls (Table 2): (i) prior buy propensity pk ∈[0,1] (0.8 foreasyto 0.05 for Figure 2: Overview of the SalesLLM Benchmark Script Generation Pipeline. The pipeline follows a hierarchical proce","claim_type":"dataset","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Multi-hop retrieval.LetC={d 1, . . . , d|C|}be a fixed passage corpus. Given a natural-language queryq, the task is to retrieve a ranked list ˆP⊆ C such that ˆPcovers as many passages in the gold support setG={g 1, g2} ⊆ Cas possible within the top-Kresults. We follow Guti 'errez et al. (2025) and Wang and Han (2025) in using R@5 =E q \" |G ∩ ˆP5| |G| # (1) as the evaluation metric, where ˆP5 is the top-5 re- trieved passages. Equation (1) rewards retrieving allgold passages and penalizes partial","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"anchors, the system traverses the cross-layer links Ecross to access specific archived Event Progres- sion Graphs stored in the archive set Sarch. These archived graphs constitute the raw source of the topic nodes. Formally, we define the candidate set C by aggregating all event nodes from the archival graphs linked to any node in the anchor set: C= [ u∈Vanchor n v∈ V(G ′) G′ ∈ Sarch, (u,G ′)∈ E cross o , (10) where V(G ′) denotes the node set of an archived graph G′. This strategy ensures that ","claim_type":"baseline","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"grows linearly with the number of steps, optimizing πθ for complex tasks becomes increasingly chal- lenging due to the accumulation of history. 2.2 Agentic Reinforcement Learning Agentic RL typically adopts policy-gradient meth- ods to optimize the agent policy πθ. We formulate the agentic RL training objective as: max πθ Ex∼D,H∼πθ(·|x) [rϕ(x,H)]−βD KL [πθ(·)∥πref(·)] (1) where πθ represents the policy LLM, πref is the reference LLM, rϕ and DKL denotes the reward function and KL divergence respe","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks online\" 'onlinestring := because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (6 contexts).","role_counts":[{"n":6,"context_role":"background"},{"n":2,"context_role":"dataset"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-20T06:31:48.960017+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"online\" 'onlinestring :=","claims":[],"why_cited":"Pith tracks online\" 'onlinestring := because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-17T10:39:51.007083+00:00"}},"summary":{"title":"online\" 'onlinestring :=","claims":[],"why_cited":"Pith tracks online\" 'onlinestring := because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"write newline","work_id":"b6ace3ad-bd44-4a95-86ee-c7b73bd4cda2","shared_citers":46},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":14},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":7},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":7},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":6},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":6},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":5},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":5},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":5},{"title":"Qwen2 Technical Report","work_id":"a1857881-ab9b-4b80-9b5f-9ae4b5c2566d","shared_citers":5},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":5},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":5},{"title":"Bleu: a method for automatic evaluation of machine translation","work_id":"7a0e7f56-7c92-470e-b7bf-2e9974bf9a93","shared_citers":4},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":4},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":4},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":4},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":4},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":4},{"title":null,"work_id":"d715ff4d-f31d-49cf-bf54-8217373c5c82","shared_citers":4},{"title":"BERTScore: Evaluating Text Generation with BERT","work_id":"9eaaaac1-0a96-4f5f-9b13-30c46e9e1346","shared_citers":3},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":3},{"title":"doi: 10.18653/v1/N18-1101","work_id":"9737cdf0-fd48-4485-ab04-c8ec9386e782","shared_citers":3},{"title":"Hellaswag: Can a machine really finish your sentence? In Korhonen, A., Traum, D., and M \\`a rquez, L","work_id":"11bfc949-547c-40f3-a86d-953eb9b2154c","shared_citers":3},{"title":"HuggingFace's Transformers: State-of-the-art Natural Language Processing","work_id":"9d86da8d-01d3-41af-a0d2-ee14897927a9","shared_citers":3}],"time_series":[{"n":11,"year":2025},{"n":36,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"5e3d68df-d072-4ce9-bb1b-53ac1ed495cc","orcid":null,"display_name":"ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school","source":"manual","import_confidence":0.72}]}}