A local Llama 3.2 pre-flight rewriter that translates and compacts non-English coding prompts cuts prompt tokens by 34–47% on a new 200-task benchmark while preserving accuracy across three commercial backends.
BrowseComp: A simple yet challenging benchmark for browsing agents
3 Pith papers cite this work. Polarity classification is still indexing.
years
2026 3representative citing papers
XpertBench provides 1,346 rubric-scored expert tasks showing leading LLMs achieve a maximum ~66% success rate and ~55% mean score across domains.
CORE is a lightweight two-stage prompt compression method for edge-device RAG QA that builds answer and clue sets via NER and semantic matching then refines them to deliver higher accuracy and lower resource costs than baselines.
citing papers explorer
-
Cross-Lingual Token Arbitrage: Optimizing Code Agent Context Windows via Local LLM Preprocessing
A local Llama 3.2 pre-flight rewriter that translates and compacts non-English coding prompts cuts prompt tokens by 34–47% on a new 200-task benchmark while preserving accuracy across three commercial backends.
-
Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation
XpertBench provides 1,346 rubric-scored expert tasks showing leading LLMs achieve a maximum ~66% success rate and ~55% mean score across domains.
-
Less is More: Lightweight Prompt Compression for Question Answering Applications on Edge Devices
CORE is a lightweight two-stage prompt compression method for edge-device RAG QA that builds answer and clue sets via NER and semantic matching then refines them to deliver higher accuracy and lower resource costs than baselines.