Pith. sign in

hub

write newline

80 Pith papers cite this work. Polarity classification is still indexing.

80 Pith papers citing it

hub tools

citation-role summary

background 1

citation-polarity summary

claims ledger

  • background sess the impact of this balance, we conduct an ablation study by varying ζ and observing its effect on Best-of-N (BoN) se- lection performance. 6.2 Experimental Setup. We perform BoN selection on our 1,000-sample CFLUE test set, using the same fine-tuned Qwen2.5-7B-Instruct model as the generator. For N values of 2, 4, 8, and 16, we vary ζ across the range [0.0, 2.0] and plot the resulting accuracy. 6.3 Results and Analysis. As shown in Figure 4, the model's performance is sensitive to the value

co-cited works

roles

background 1

polarities

background 1

representative citing papers

Dynamic Tool Dependency Retrieval for Lightweight Function Calling

cs.LG · 2025-12-18 · unverdicted · novelty 7.0

DTDR dynamically retrieves relevant tools by modeling dependencies from demonstrations and conditioning on the evolving agent plan, improving function calling success rates by 23-104% over static retrievers across benchmarks.

Incremental Data-Driven Policy Synthesis via Game Abstractions

cs.GT · 2025-11-14 · unverdicted · novelty 7.0

An incremental rank-lifting algorithm updates winning regions and policies in data-driven stochastic game abstractions by exploiting monotonic growth of under-approximations and shrinkage of over-approximations.

TRAM: Test-Time Risk Adaptation with Mixture of Agents

cs.LG · 2024-08-16 · unverdicted · novelty 7.0

TRAM is a test-time mixture method that scores and composes risk-neutral source policies using reward and occupancy-based risk to achieve new reward-risk tradeoffs without parameter updates.

From Next Token Prediction to (STRIPS) World Models

cs.AI · 2025-09-16 · reject · novelty 6.0

A specialized transformer trained only on action traces can recover an exact propositional STRIPS planning model, provided the number of world atoms is known in advance.

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal

cs.CL · 2025-09-07 · unverdicted · novelty 6.0

Sparse autoencoders plus greedy filtering and factorization-machine interaction modeling identify minimal sets of features in Gemma-2-2B-IT and LLaMA-3.1-8B-IT whose ablation produces jailbreaks by flipping refusal to compliance.

citing papers explorer

Showing 50 of 80 citing papers.