{"total":12,"items":[{"citing_arxiv_id":"2605.12000","ref_index":65,"ref_count":2,"confidence":0.55,"is_internal_anchor":false,"paper_title":"Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation","primary_cat":"cs.LG","submitted_at":"2026-05-12T11:49:08+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"MA-BC partitions divergent expert data and pools non-conflicting pairs to achieve faster convergence to Pareto-optimal policies in MOMDPs, with a matching minimax lower bound.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"Multi-agent adversarial inverse reinforcement learning. InInternational conference on machine learning, pages 7194-7201. PMLR, 2019. [64] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric. P Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena, 2023. [65] B. D. Ziebart, A. Maas, J. A. Bagnell, and A. K. Dey. Maximum entropy inverse reinforcement learning. InNational Conference on Artificial Intelligence (AAAI), 2008. 14 Contents of Appendix A Related Work 16 B Omitted Proofs 18 B.1 Proof of Theorem 2.1 . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 B.2 Proof of Theorem 2."},{"citing_arxiv_id":"2605.06914","ref_index":20,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"Regulating Branch Parallelism in LLM Serving","primary_cat":"cs.DC","submitted_at":"2026-05-07T20:23:32+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"TAPER regulates LLM branch parallelism by admitting extra branches opportunistically when predicted externality fits slack, delivering 1.48-1.77x higher goodput than eager or fixed-cap baselines on Qwen3-32B while keeping over 95% SLO attainment.","context_count":1,"top_context_role":"dataset","top_context_polarity":"use_dataset","context_text":"How common are decomposable requests?We characterize three representative datasets (Figure 1). Theproportion of decomposable requests(PDR) measures how often parallelism appears: Math-220K has the highest PDR at 84.2%, followed by the RAG-12K dataset at 67.0% and ShareGPT Vicuna at 43.5%. Theparallel token share(PTS) measures how much of a decomposable response is parallel: ShareGPT [20] (70.5%) and RAG-12K [21] (68.9%) are dominated by parallel tokens when they decompose, while Math-220K [22] (30.6%) produces short, narrow parallel stages. Theaverage branch fanout(ABF) determines the maximum step width the runtime could admit: ShareGPT averages 5.2 branches per parallel stage, RAG-12K 4.2, and Math-220K 2.7. The pattern is not uniform."},{"citing_arxiv_id":"2604.06427","ref_index":31,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning","primary_cat":"cs.LG","submitted_at":"2026-04-07T20:04:14+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LLMs discover latent planning strategies up to five steps during training and execute them up to eight steps at test time, with larger models reaching seven under few-shot prompting, revealing a dissociation between discovery and execution.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2509.21319","ref_index":51,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"RLBFF: Binary Flexible Feedback to bridge between Human Feedback & Verifiable Rewards","primary_cat":"cs.CL","submitted_at":"2025-09-25T16:19:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"RLBFF extracts binary principles from human feedback to train reward models that outperform Bradley-Terry models on RM-Bench and JudgeBench and enable customizable inference-time focus for LLM alignment.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2410.10813","ref_index":107,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory","primary_cat":"cs.CL","submitted_at":"2024-10-14T17:59:44+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LongMemEval benchmarks long-term memory in chat assistants, revealing 30% accuracy drops across sustained interactions and proposing indexing-retrieval-reading optimizations that boost performance.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2408.04840","ref_index":271,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models","primary_cat":"cs.CV","submitted_at":"2024-08-09T03:25:42+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"mPLUG-Owl3 introduces hyper attention blocks to integrate vision and language for long image-sequence understanding and reports SOTA results on single-image, multi-image, and video benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2405.01470","ref_index":22,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"WildChat: 1M ChatGPT Interaction Logs in the Wild","primary_cat":"cs.CL","submitted_at":"2024-05-02T17:00:02+00:00","verdict":"ACCEPT","verdict_confidence":"LOW","novelty_score":8.0,"formal_verification":"none","one_line_summary":"WildChat releases a dataset of 1 million ChatGPT conversations with timestamps, demographics, and headers, claimed to be the most diverse and multilingual such resource available.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2403.04652","ref_index":95,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"Yi: Open Foundation Models by 01.AI","primary_cat":"cs.CL","submitted_at":"2024-03-07T16:52:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Yi models are 6B and 34B open foundation models pretrained on 3.1T curated tokens that achieve strong benchmark results through data quality and targeted extensions like long context and vision alignment.","context_count":1,"top_context_role":"dataset","top_context_polarity":"use_dataset","context_text":"Stage 2: we scale up the image resolution of ViT to 4482, aiming to further boost the model's capability for discerning intricate visual details. The dataset used in this stage includes 20 million image-text pairs derived from LAION-400M. Additionally, we incorporate around 4.8 million image-text pairsn from diverse sources, e.g., CLLaV A [45], LLaV AR [91], Flickr [85], VQAv2 [25], RefCOCO [37], Visual7w [95] and so on. Stage 3: the parameters of the entire model are trained. The primary goal is to enhance the model's proficiency in multimodal chat interactions, thereby endowing it with the ability to seamlessly integrate and interpret visual and linguistic inputs. To this end, the training dataset encompasses a diverse range of sources, totalling approximately 1 million image-text pairs, including GQA [32],"},{"citing_arxiv_id":"2311.12983","ref_index":206,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"GAIA: a benchmark for General AI Assistants","primary_cat":"cs.CL","submitted_at":"2023-11-21T20:34:47+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"GAIA benchmark shows humans at 92% accuracy on simple real-world questions far outperform current AI systems at 15%, proposing this gap as a key milestone for general AI.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2311.07911","ref_index":28,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"Instruction-Following Evaluation for Large Language Models","primary_cat":"cs.CL","submitted_at":"2023-11-14T05:13:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"IFEval is a new benchmark of 25 verifiable instruction types and ~500 prompts for objective, reproducible evaluation of LLMs' instruction-following abilities.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2308.14132","ref_index":38,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"Detecting Language Model Attacks with Perplexity","primary_cat":"cs.CL","submitted_at":"2023-08-27T15:20:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Jailbreak prompts with adversarial suffixes have high GPT-2 perplexity, and a LightGBM model on perplexity and length detects most attacks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2306.14048","ref_index":51,"ref_count":1,"confidence":0.55,"is_internal_anchor":false,"paper_title":"H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models","primary_cat":"cs.LG","submitted_at":"2023-06-24T20:11:14+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"H2O evicts non-heavy-hitter tokens from the KV cache using a dynamic submodular policy, retaining recent and frequent-co-occurrence tokens to reduce memory while preserving accuracy.","context_count":1,"top_context_role":"dataset","top_context_polarity":"use_dataset","context_text":"OPT [39] with model sizes, LLaMA [40], and GPT-NeoX-20B [41]. We sample eight tasks from two popular evaluation frameworks (HELM [16] and lm-eval-harness [15]): COPA [42], MathQA [43], OpenBookQA [44], PiQA [45], RTE [46], Winogrande [47], XSUM [48], CNN/Daily Mail [49]. Also, we evaluate our approach on recent generation benchmarks, AlpaceEval [50] and MT-bench [51], and the details are included in Appendix. We use NVIDIA A100 80GB GPU. Baselines. Since H2O evenly assigns the caching budget to H2 and the most recent KV, except for full KV cache, we consider the \"Local\" strategy as a baseline method. In addition, we also provide two different variants of Sparse Transformers (strided and fixed) as strong baselines."}],"limit":50,"offset":0}