Pith. sign in

hub Canonical reference

PIXIU: A large language model, instruction data and evaluation benchmark for finance

Canonical reference. 80% of citing Pith papers cite this work as background.

22 Pith papers citing it
Background 80% of classified citations

hub tools

citation-role summary

background 4 method 1

citation-polarity summary

representative citing papers

SAGE: A Service Agent Graph-guided Evaluation Benchmark

cs.AI · 2026-04-10 · unverdicted · novelty 7.0

SAGE is a new multi-agent benchmark that formalizes service SOPs as dynamic dialogue graphs to measure LLM agents on logical compliance and path coverage, uncovering an execution gap and empathy resilience across 27 models in 6 scenarios.

DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs

cs.AI · 2026-05-24 · unverdicted · novelty 5.0

DarkForest improves multi-agent LLM reasoning via independent agents, semantic clustering of responses, and calibrated belief estimation with controlled communication, yielding up to 30.7% better benchmark metrics and 6.5x lower token consumption than heavy-communication baselines.

citing papers explorer

Showing 22 of 22 citing papers.