Pith. sign in

Canonical reference

Title resolution pending

Canonical reference. 74% of citing Pith papers cite this work as background.

126 Pith papers citing it
818 external citations · Crossref
Background 74% of classified citations

citation-role summary

background 17 baseline 1 method 1

citation-polarity summary

representative citing papers

LogicHunter: Testing LLM Agent Frameworks with an Agentic Oracle

cs.SE · 2026-07-07 · conditional · novelty 7.0

LogicHunter combines specification-driven test generation with a ReAct-based agentic oracle to discover 40 previously unknown bugs in LangChain, LlamaIndex, and CrewAI, achieving 91.17% oracle precision.

Fork-Think with Confidence

cs.LG · 2026-06-30 · unverdicted · novelty 7.0

Fork-think with confidence identifies forking points via model confidence in a single path before sampling continuations, cutting tokens up to 30% and runtime up to 57% on reasoning benchmarks while matching or exceeding parallel thinking performance.

Structured Inference with Large Language Gibbs

cs.LG · 2026-06-17 · unverdicted · novelty 7.0

Large Language Gibbs uses LLM next-token conditionals as MCMC transition operators for iterative resampling of structured variables, aiming to produce a stationary distribution that compromises across all local conditionals.

What Gets Cited: Competitive GEO in AI Answer Engines

cs.AI · 2026-05-25 · unverdicted · novelty 7.0

Across 252,000 paired trials on six LLMs, topical relevance and list position emerged as the strongest drivers of first citation in competitive RAG, with price information and recency providing consistent secondary gains.

ContractBench: Can LLM Agents Preserve Observation Contracts?

cs.SE · 2026-05-17 · conditional · novelty 7.0

ContractBench shows that LLM agents frequently violate observation contracts by using expired artifacts or corrupting their byte integrity, with no model exceeding 80% success and notable scaling irregularities across families.

Zero-Shot Goal Recognition with Large Language Models

cs.AI · 2026-05-14 · unverdicted · novelty 7.0

Frontier LLMs show uneven zero-shot performance on goal recognition in PDDL domains: some scale with accumulating evidence toward landmark-based accuracy while others stay anchored to world-knowledge priors.

Synthesizing Multi-Agent Harnesses for Vulnerability Discovery

cs.CR · 2026-04-22 · unverdicted · novelty 7.0

AgentFlow uses a typed graph DSL covering roles, prompts, tools, topology and protocol plus a runtime-signal feedback loop to optimize multi-agent harnesses, reaching 84.3% on TerminalBench-2 and discovering ten new zero-days in Chrome including two critical sandbox escapes.

HorizonBench: Long-Horizon Personalization with Evolving Preferences

cs.CL · 2026-04-19 · unverdicted · novelty 7.0

HorizonBench generates 6-month conversation histories from structured mental state graphs to test AI models on tracking evolving user preferences, finding that frontier models mostly fail at belief updates and perform near or below chance.

citing papers explorer

Showing 50 of 126 citing papers.