{"work":{"id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","openalex_id":null,"doi":"10.1016/j.aiopen.2022.12","arxiv_id":"2505.09388","raw_key":null,"title":"Qwen3 Technical Report","authors":null,"authors_text":"An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng","year":2025,"venue":"cs.CL","abstract":"In this work, we present Qwen3, the latest version of the Qwen model family. Qwen3 comprises a series of large language models (LLMs) designed to advance performance, efficiency, and multilingual capabilities. The Qwen3 series includes models of both dense and Mixture-of-Expert (MoE) architectures, with parameter scales ranging from 0.6 to 235 billion. A key innovation in Qwen3 is the integration of thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework. This eliminates the need to switch between different models--such as chat-optimized models (e.g., GPT-4o) and dedicated reasoning models (e.g., QwQ-32B)--and enables dynamic mode switching based on user queries or chat templates. Meanwhile, Qwen3 introduces a thinking budget mechanism, allowing users to allocate computational resources adaptively during inference, thereby balancing latency and performance based on task complexity. Moreover, by leveraging the knowledge from the flagship models, we significantly reduce the computational resources required to build smaller-scale models, while ensuring their highly competitive performance. Empirical evaluations demonstrate that Qwen3 achieves state-of-the-art results across diverse benchmarks, including tasks in code generation, mathematical reasoning, agent tasks, etc., competitive against larger MoE models and proprietary models. Compared to its predecessor Qwen2.5, Qwen3 expands multilingual support from 29 to 119 languages and dialects, enhancing global accessibility through improved cross-lingual understanding and generation capabilities. To facilitate reproducibility and community-driven research and development, all Qwen3 models are publicly accessible under Apache 2.0.","external_url":"https://arxiv.org/abs/2505.09388","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-11T03:27:46.490459+00:00","pith_arxiv_id":"2505.09388","created_at":"2026-05-08T17:13:38.674290+00:00","updated_at":"2026-07-11T11:50:26.030339+00:00","title_quality_ok":false,"display_title":"Qwen3 Technical Report","render_title":"Qwen3 Technical Report"},"hub":{"state":{"work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","tier":"mega_hub","tier_reason":"1,000+ Pith inbound or 100,000+ external citations","pith_inbound_count":3168,"external_cited_by_count":null,"distinct_field_count":51,"first_pith_cited_at":"2024-10-22T18:47:46+00:00","last_pith_cited_at":"2026-07-09T16:50:41+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"needed","recognition_status":"needed","updated_at":"2026-08-23T20:29:16.811238+00:00","tier_text":"mega_hub"},"tier":"mega_hub","role_counts":[{"context_role":"background","n":284},{"context_role":"method","n":69},{"context_role":"baseline","n":43},{"context_role":"dataset","n":26},{"context_role":"other","n":14}],"polarity_counts":[{"context_polarity":"background","n":270},{"context_polarity":"use_method","n":70},{"context_polarity":"baseline","n":43},{"context_polarity":"unclear","n":27},{"context_polarity":"use_dataset","n":26}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Qwen3 Technical Report","claims":[{"claim_text":"In this work, we present Qwen3, the latest version of the Qwen model family. Qwen3 comprises a series of large language models (LLMs) designed to advance performance, efficiency, and multilingual capabilities. The Qwen3 series includes models of both dense and Mixture-of-Expert (MoE) architectures, with parameter scales ranging from 0.6 to 235 billion. A key innovation in Qwen3 is the integration of thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework. This eliminates the need to switch between different models--","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T17:23:29.390948+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"f6e1a34c-959d-4c3d-b317-1a65d6fe682c","orcid":null,"display_name":"An Yang"},{"id":"067eafca-97ae-45d8-9656-e86a5a431723","orcid":null,"display_name":"Anfeng Li"},{"id":"3e0735af-daed-4e22-b48a-b7778e5e9a45","orcid":null,"display_name":"Baosong Yang"},{"id":"54001702-5e02-4890-9876-458e5739ed23","orcid":null,"display_name":"Beichen Zhang"},{"id":"785c6603-fbf9-4a35-bfb2-686be03040d6","orcid":null,"display_name":"Binyuan Hui"},{"id":"cc48b5a8-a38d-4d8c-ace7-5add79c07fd0","orcid":null,"display_name":"Bo Zheng"}]},"error":null,"updated_at":"2026-05-13T17:24:03.750030+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-13T17:23:28.176903+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":244},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":236},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":218},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":149},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":130},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":127},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":121},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":120},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":103},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":93},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":87},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":84},{"title":"gpt-oss-120b & gpt-oss-20b Model Card","work_id":"178c1f7e-4f19-4392-a45d-45a6dfa88ead","shared_citers":72},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":71},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":69},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":65},{"title":"DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models","work_id":"07c85cc5-4086-4abc-823b-6d0f4ff784d0","shared_citers":64},{"title":"Kimi K2: Open Agentic Intelligence","work_id":"7f18284c-12d3-4137-bea1-1da97e8cf3c1","shared_citers":64},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":57},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":57},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":56},{"title":"OpenAI GPT-5 System Card","work_id":"ca87689a-0d29-4476-b504-b65dbbb08af4","shared_citers":56},{"title":"Measuring Mathematical Problem Solving With the MATH Dataset","work_id":"50652ac6-fb7c-4675-a2c2-159c241feb17","shared_citers":55},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":51}],"time_series":[{"n":10,"year":2025},{"n":1076,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T01:22:28.613632+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"fixed":1,"items":[{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-13T17:23:27.442755+00:00"},"reader_index":{"job_type":"reader_index","status":"succeeded","result":{"note":"annotated reader requires full-text/OA fetch; shell is wired for mega hubs","status":"reader queued"},"error":null,"updated_at":"2026-05-13T17:23:29.392123+00:00"},"recognition_alignment":{"job_type":"recognition_alignment","status":"succeeded","result":{"modules":["IndisputableMonolith.Cosmology.InflationModelsFromConfigDim","IndisputableMonolith.Education.PedagogyModelsFromConfigDim","IndisputableMonolith.RRF.Models","IndisputableMonolith.Physics.StandardModelGroupStructure","IndisputableMonolith.Physics.StandardModelLagrangianStructure","IndisputableMonolith.Chemistry.VanDerWaals","IndisputableMonolith.Physics.GrandUnificationFromRS","IndisputableMonolith.Physics.DarkMatterCrossSectionBandScoreCard"],"query_chars":1804},"error":null,"updated_at":"2026-05-13T17:23:37.298302+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Qwen3 Technical Report","claims":[{"claim_text":"In this work, we present Qwen3, the latest version of the Qwen model family. Qwen3 comprises a series of large language models (LLMs) designed to advance performance, efficiency, and multilingual capabilities. The Qwen3 series includes models of both dense and Mixture-of-Expert (MoE) architectures, with parameter scales ranging from 0.6 to 235 billion. A key innovation in Qwen3 is the integration of thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework. This eliminates the need to switch between different models--","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T17:23:28.186109+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Qwen3 Technical Report","claims":[{"claim_text":"In this work, we present Qwen3, the latest version of the Qwen model family. Qwen3 comprises a series of large language models (LLMs) designed to advance performance, efficiency, and multilingual capabilities. The Qwen3 series includes models of both dense and Mixture-of-Expert (MoE) architectures, with parameter scales ranging from 0.6 to 235 billion. A key innovation in Qwen3 is the integration of thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework. This eliminates the need to switch between different models--","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-13T17:23:28.183092+00:00"}},"summary":{"title":"Qwen3 Technical Report","claims":[{"claim_text":"In this work, we present Qwen3, the latest version of the Qwen model family. Qwen3 comprises a series of large language models (LLMs) designed to advance performance, efficiency, and multilingual capabilities. The Qwen3 series includes models of both dense and Mixture-of-Expert (MoE) architectures, with parameter scales ranging from 0.6 to 235 billion. A key innovation in Qwen3 is the integration of thinking mode (for complex, multi-step reasoning) and non-thinking mode (for rapid, context-driven responses) into a unified framework. This eliminates the need to switch between different models--","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen3 Technical Report because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":244},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":236},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":218},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":149},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":130},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":127},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":121},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":120},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":103},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":93},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":87},{"title":"GPT-4o System Card","work_id":"f37bf1c7-4964-4e56-9762-d20da8d9009f","shared_citers":84},{"title":"gpt-oss-120b & gpt-oss-20b Model Card","work_id":"178c1f7e-4f19-4392-a45d-45a6dfa88ead","shared_citers":72},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":71},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":69},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":65},{"title":"DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models","work_id":"07c85cc5-4086-4abc-823b-6d0f4ff784d0","shared_citers":64},{"title":"Kimi K2: Open Agentic Intelligence","work_id":"7f18284c-12d3-4137-bea1-1da97e8cf3c1","shared_citers":64},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":57},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":57},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":56},{"title":"OpenAI GPT-5 System Card","work_id":"ca87689a-0d29-4476-b504-b65dbbb08af4","shared_citers":56},{"title":"Measuring Mathematical Problem Solving With the MATH Dataset","work_id":"50652ac6-fb7c-4675-a2c2-159c241feb17","shared_citers":55},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":51}],"time_series":[{"n":10,"year":2025},{"n":1076,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"067eafca-97ae-45d8-9656-e86a5a431723","orcid":null,"display_name":"Anfeng Li","source":"manual","import_confidence":0.72},{"id":"f6e1a34c-959d-4c3d-b317-1a65d6fe682c","orcid":null,"display_name":"An Yang","source":"manual","import_confidence":0.72},{"id":"3e0735af-daed-4e22-b48a-b7778e5e9a45","orcid":null,"display_name":"Baosong Yang","source":"manual","import_confidence":0.72},{"id":"54001702-5e02-4890-9876-458e5739ed23","orcid":null,"display_name":"Beichen Zhang","source":"manual","import_confidence":0.72},{"id":"785c6603-fbf9-4a35-bfb2-686be03040d6","orcid":null,"display_name":"Binyuan Hui","source":"manual","import_confidence":0.72},{"id":"cc48b5a8-a38d-4d8c-ace7-5add79c07fd0","orcid":null,"display_name":"Bo Zheng","source":"manual","import_confidence":0.72}]},"citers":{"total":3168,"items":[{"citing_arxiv_id":"2607.08690","ref_index":32,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"A Practical Investigation of Training-free Relaxed Speculative Decoding","primary_cat":"cs.LG","submitted_at":"2026-07-09T16:50:41+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Relaxed speculative decoding methods that trade exactness for speed only produce useful capability-speed trade-offs when the drafter is a strong standalone language model; lightweight MTP drafters are largely unsuited to relaxation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08535","ref_index":23,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability","primary_cat":"cs.CL","submitted_at":"2026-07-09T14:31:05+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Judge upgrades are not interchangeable: only Qwen3 1.7B→4B yields robust adjacent gains, MiniMax adjacent releases do not, and stronger judges reduce but do not remove bias or correlated jury errors.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08497","ref_index":40,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing","primary_cat":"cs.CV","submitted_at":"2026-07-09T13:55:55+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A modular 8B agent with episodic visual memory and RL-trained retrieval reaches 91.4% cross-turn image recall over 20 turns, outperforming 32B all-context baselines with ~1.8× lower latency.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08423","ref_index":31,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"OmniFood-Bench: Evaluating VLMs for Nutrient Reasoning and Personalized Health Advice","primary_cat":"cs.AI","submitted_at":"2026-07-09T12:46:00+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A new benchmark shows current vision-language models can name foods but fail at estimating portion and nutrient values and often give unsafe dietary advice for chronic-disease patients.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08326","ref_index":11,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Diagnosing and Repairing Persona Collapse in LLM Advice","primary_cat":"cs.CY","submitted_at":"2026-07-09T10:12:18+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"LLMs collapse advice into a single supportive persona; Inverse-Process Distillation restores human-like persona diversity, yet raters still prefer the collapsed default.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08317","ref_index":49,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Blind-Spots-Bench: Evaluating Blind Spots in Multimodal Models","primary_cat":"cs.AI","submitted_at":"2026-07-09T09:56:50+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A 235-item multimodal stress-test shows frontier closed models outpace open-weight peers by ~10% and leaves shared failures on counting, spatial, and character-level tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08268","ref_index":5,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Different Teachers, Different Capabilities: Sub-1B On-Device Distillation for Structured Text Enrichment","primary_cat":"cs.AI","submitted_at":"2026-07-09T09:10:49+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Distilling an 8B reasoning teacher into a 0.6B student recovers most summary quality at ~50× speed, but teacher type—not scale alone—determines which capabilities transfer.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08257","ref_index":26,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"MentalHospital: A Virtual Environment for Evaluating Psychiatric Clinical Encounters","primary_cat":"cs.AI","submitted_at":"2026-07-09T09:03:42+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"EHR-derived standardized patients and dual-track evaluation reveal LLMs trail clinicians by 37.28 points on full psychiatric encounters, with mental-status assessment the main bottleneck.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08221","ref_index":32,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"LUMI: Tokenizer-Agnostic LLM-Based Lossless Image Compression","primary_cat":"cs.CV","submitted_at":"2026-07-09T08:16:14+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A tokenizer-free pixel embedding, position encoding, and 256-way head let frozen LLMs act as portable entropy models for lossless RGB compression across model families.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08208","ref_index":19,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech","primary_cat":"cs.CL","submitted_at":"2026-07-09T08:07:34+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"A Qwen3-ASR-based two-speaker, 21-language transcription system cuts its official error metric from 30.53 to 23.70 on the MLC-SLM 2026 dev set; supervised fine-tuning delivers most of the gain.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08196","ref_index":177,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"A First-Principles Theory of Slow Thinking and Active Perception","primary_cat":"cs.AI","submitted_at":"2026-07-09T07:54:39+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":7.5,"formal_verification":"none","one_line_summary":"Active lifting of data distributions via latent-sequence sampling and max-rate uncertainty reduction formally derives slow-thinking LLMs and places them on representation and sampler hierarchies that can be climbed.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08194","ref_index":49,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Dive Into the Implicit Biases of Low-rank Vision-language Alignment","primary_cat":"cs.CV","submitted_at":"2026-07-09T07:50:49+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Low-rank LLM adaptation during vision-language alignment outperforms full fine-tuning by preserving per-token visual structure and favoring flat, noise-robust subspaces.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08170","ref_index":50,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Understanding Layer Patching in Model Size Interpolation","primary_cat":"cs.LG","submitted_at":"2026-07-09T07:14:12+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Optimal layer-patching order for boomerang distillation is a shortest path on a KL-weighted Boolean lattice; greedy KLPatch and simple sequential orders often yield near-optimal interpolations.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08046","ref_index":16,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"What LLM Forecasters Know but Don't Say: Probing Internal Representations for Calibration and Faithfulness","primary_cat":"cs.CL","submitted_at":"2026-07-09T01:52:09+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Linear probes on intermediate LLM activations produce better-calibrated confidence than verbalized probabilities, detect hidden evidence influence, and reveal that forecasts are largely pre-committed before reasoning begins.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08038","ref_index":24,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis","primary_cat":"cs.AI","submitted_at":"2026-07-09T01:30:24+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A multi-agent LLM framework with safety-oriented reasoning gates improves diagnostic accuracy and must-not-miss condition coverage over standalone LLMs across case-report benchmarks and a blinded physician evaluation of 43 real-world ED notes.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08017","ref_index":3,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework","primary_cat":"cs.CL","submitted_at":"2026-07-09T00:45:26+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Graph Reasoning Coherence Score (GRCS) and Graph Self-Consistency (GSC) quantify and select faithful LLM reasoning better than final-answer majority voting, with the selected medoid acting as a load-bearing path under ablation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08010","ref_index":53,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Tool-Making and Self-Evolving LLM Agents in Low-Latency Systems","primary_cat":"cs.CL","submitted_at":"2026-07-09T00:27:13+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Compiling repeated SOP nodes into environment-grounded, versioned tools cuts production p50 latency by 42% and end-to-end error rate by up to 53% in a 44-node fulfillment-center alarm-triage agent.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.08009","ref_index":194,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs","primary_cat":"cs.CL","submitted_at":"2026-07-09T00:27:07+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.5,"formal_verification":"none","one_line_summary":"On 2,520 programming tasks, matched Qwen general and coder models reliably raise Bloom cognitive demand but fail to lower it, so execution skill does not imply educational control.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07993","ref_index":18,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Hallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator","primary_cat":"cs.CL","submitted_at":"2026-07-08T23:54:36+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Hallucination Self-Play co-evolves a generator and detector from one base LLM via RLAIF and RLVR, lifting a 7B model to match advanced LLMs on RAGTruth faithfulness detection.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07976","ref_index":34,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning","primary_cat":"cs.CL","submitted_at":"2026-07-08T23:00:54+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"TACO soft-suppresses positive GRPO credit on high tail-risk tokens (surprisal above local entropy) and consistently beats GRPO-style baselines on three LLMs and eight reasoning benchmarks while stabilizing long training.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07918","ref_index":51,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Efficient Safety Alignment of Language Models via Latent Personality Traits","primary_cat":"cs.LG","submitted_at":"2026-07-08T21:03:27+00:00","verdict":null,"verdict_confidence":null,"novelty_score":null,"formal_verification":null,"one_line_summary":null,"context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07916","ref_index":74,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Persona Cartography: Charting Language Model Personality Traits in Weight Space","primary_cat":"cs.AI","submitted_at":"2026-07-08T21:00:44+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Composable LoRA adapters can amplify or suppress OCEAN traits in LLMs, combine approximately additively, preserve moderate-scale capability, and move safety-relevant behaviours.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07907","ref_index":2,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks","primary_cat":"cs.LG","submitted_at":"2026-07-08T20:42:46+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A system-first taxonomy and literature synthesis of multimodal unlearning across vision, language, video, and audio, with datasets, benchmarks, metrics, applications, and open challenges.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07847","ref_index":53,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"When Does Continual Learning Require Learning","primary_cat":"cs.LG","submitted_at":"2026-07-08T18:27:40+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Different patterns of environmental change (space vs time) require different LLM update behaviors; no single family of methods—prompts, distillation, RL, or compression—handles all regimes.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07824","ref_index":5,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"From Triggers to Emotions: A CPM-Grounded Appraisal Multi-Agent for Dynamic Emotional Evolution in Persona-Based Dialogue","primary_cat":"cs.MA","submitted_at":"2026-07-08T18:06:03+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A CPM-grounded multi-agent system extracts dialogue triggers, appraises them on relevance/implication/coping/norms, and updates a persona’s latent multi-emotion state more coherently than standard prompting baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07708","ref_index":82,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning","primary_cat":"cs.CL","submitted_at":"2026-07-08T17:59:59+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A single autoregressive foundation model unifies protein, molecule, and crystal structures into a shared token vocabulary and generates inspectable reasoning traces, achieving SOTA on 67 of 86 scientific tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07779","ref_index":273,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier","primary_cat":"cs.CL","submitted_at":"2026-07-08T17:46:36+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"LLM formal provers must shift from competition solvers to research agents that handle open-ended, under-specified frontier mathematics under machine-checked rigor.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07674","ref_index":17,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Max Out GRPO Signal: Adaptive Trace Prefix Control for Hard Reasoning Problems","primary_cat":"cs.LG","submitted_at":"2026-07-08T17:32:58+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AdaPrefix-GRPO treats solution-prefix length as a feedback controller targeting 50% rollout success rate during GRPO training, then anneals to zero prefix, yielding 1.6–2.1× accuracy gains over vanilla GRPO at matched compute on hard math.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07626","ref_index":20,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Future Confidence Distillation in Large Language Models","primary_cat":"cs.CL","submitted_at":"2026-07-08T16:43:11+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Linear probes trained on pre-solution hidden states, supervised by post-solution correctness probe outputs, recover 32–66% of the calibration gap between pre- and post-solution confidence across five open-source LLMs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07593","ref_index":34,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"What Makes a Good Bug Report for an AI Agent?","primary_cat":"cs.SE","submitted_at":"2026-07-08T16:13:15+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"AI repair agents solve bugs more reliably when reports include executable reproduction scripts, file-level localization cues, and clear structure, while longer prose reports and human-oriented steps to reproduce show no benefit or hurt.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07548","ref_index":21,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Think Big, Search Small: Where Capacity Matters in Hierarchical Search Agents?","primary_cat":"cs.CL","submitted_at":"2026-07-08T15:46:48+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Scaling the delegation backbone in hierarchical search agents improves EM by ~11 points while scaling the executor moves EM by only ~2.6 points, and a 1.7B SFT executor matches a frontier sub-agent at 37% fewer tokens.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07508","ref_index":20,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning","primary_cat":"cs.LG","submitted_at":"2026-07-08T15:02:19+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SAO stabilizes asynchronous RL for LLMs by replacing group-wise sampling with single-rollout updates, token-level importance sampling, and targeted value-model training, outperforming GRPO on reasoning and coding benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07435","ref_index":20,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"RLVP: Penalize the Path, Reward the Outcome","primary_cat":"cs.LG","submitted_at":"2026-07-08T14:06:14+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Pairing outcome rewards with verifiable per-action path penalties reduces constraint violations nearly sixfold at equal task success, while a progress potential accelerates learning only where partial progress is reachable.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07409","ref_index":25,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"DeLS-Spec: Decoupled Long-Short Contexts for Parallel Speculative Drafting","primary_cat":"cs.CL","submitted_at":"2026-07-08T13:41:52+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"DeLS-Spec improves block-parallel speculative decoding by fusing DFlash logits with an independently trained lightweight local head and a unigram prior correction.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07403","ref_index":9,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Multi-Agent Robotic Control with Onboard Vision-Language Models","primary_cat":"cs.MA","submitted_at":"2026-07-08T13:37:31+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"A multi-agent architecture using compact VLMs on edge hardware controls a mobile manipulator across five warehouse task categories in hardware-in-the-loop simulation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07388","ref_index":32,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"TF-Engram: A Train-Free Engram with SSD-Backed Memory for Large Language Models","primary_cat":"cs.CL","submitted_at":"2026-07-08T13:19:52+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A train-free phrase memory system for LLMs uses SSD-backed hierarchical storage and early-exit predictive prefetching to improve downstream accuracy without backbone training.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07387","ref_index":15,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"A Large Language Model-Driven Agent-Based Modeling Framework with Multi-Round Communication for Simulating Vaccine Opinion Dynamics","primary_cat":"cs.MA","submitted_at":"2026-07-08T13:19:47+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"An LLM-driven agent-based model with multi-round dialogue reproduces non-linear social influence patterns in vaccination opinion dynamics, with memory increasing resistance and prompt diversity increasing adoption.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07318","ref_index":12,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"R^3: Advertisement Compliance Rectification via Group-Relative Experience Extractor and Curriculum Reinforcement","primary_cat":"cs.CL","submitted_at":"2026-07-08T12:05:41+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"R3 uses curriculum RL with hierarchical rewards and a group-relative experience extractor to automatically rectify non-compliant text in video ads while preserving semantic intent.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07287","ref_index":23,"ref_count":2,"confidence":0.98,"is_internal_anchor":true,"paper_title":"TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation","primary_cat":"cs.RO","submitted_at":"2026-07-08T11:28:13+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.5,"formal_verification":"none","one_line_summary":"A multi-timescale tactile hierarchy with subtask planning, tactile world-model goals, and residual refinement raises real-robot success by about 16–19 points over strong baselines on six contact-rich tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07748","ref_index":22,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Selective Left-Shift: Turning Test-Time Compute and Difficulty-based Curation into Training Data for Low-Resource Code Generation","primary_cat":"cs.LG","submitted_at":"2026-07-08T09:18:56+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Left-shifting iterative compiler/test refinement into verified SFT data, then GRPO on difficulty-curated IO rewards, lifts Qwen3-8B Julia pass@1 past prior SOTA at 1/3 data and 1/6 cost, and bootstraps Ballerina.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07128","ref_index":42,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Distributed Sparse Interventions in Language Models","primary_cat":"cs.LG","submitted_at":"2026-07-08T08:19:11+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Sparse interventions on 8–64 neurons distributed across layers can activate task behavior in instruction-tuned LLMs, outperforming first-order linear steering approaches by modeling nonlinear neuron interactions.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.07740","ref_index":1,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPE","primary_cat":"cs.LG","submitted_at":"2026-07-08T06:23:42+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Jet-Long is a tuning-free bifocal RoPE method that dynamically sets remote group size from sequence length, recovering the base model within the pretrained window and beating prior zero-shot extenders on RULER, HELMET-RAG, and PG-19 with near-FA2 throughput.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06993","ref_index":15,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Large Behavior Model: A Promptable Digital Twin of the Retail Customer","primary_cat":"cs.AI","submitted_at":"2026-07-08T04:31:18+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Grounding an LLM in verbalized transaction histories via Person–Environment prompting, continued pre-training, SFT, and GRPO yields stronger retail decision simulation than frontier models, with partial cross-domain transfer.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06990","ref_index":60,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"A Closed-Loop Multi-Agent Framework for Robust Multi-Robot Manipulation","primary_cat":"cs.RO","submitted_at":"2026-07-08T04:23:41+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A closed-loop multi-agent LLM framework enables heterogeneous robots to collaboratively manipulate objects by decomposing tasks, grounding actions via visual tools, and recovering from execution failures hierarchically.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06987","ref_index":36,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma","primary_cat":"cs.LG","submitted_at":"2026-07-08T04:21:42+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Replacing the importance sampling ratio with a stop-gradient self-anchored ratio for positive advantages yields unclipped, REINFORCE-equivalent gradients that improve exploration without training instability.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06974","ref_index":45,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning","primary_cat":"cs.CL","submitted_at":"2026-07-08T03:51:37+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"MILES dynamically expands step-wise memory with learnable selection heads that rerank candidates and guide reasoning, improving LLM test-time performance under limited supervision.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06855","ref_index":37,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Geometric Self-Distillation for Reasoning Generalization","primary_cat":"cs.LG","submitted_at":"2026-07-07T23:16:19+00:00","verdict":"ACCEPT","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A geometric self-distillation objective using Hellinger loss and Fisher-Rao proximal regularization prevents predictive drift in LLM post-training, improving out-of-distribution reasoning by 5.7-8.6 points.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06827","ref_index":3,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs","primary_cat":"eess.AS","submitted_at":"2026-07-07T21:45:04+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Learned pooling of speech KV caches from an intermediate LLM layer compresses speech to text-level length while matching or exceeding the uncompressed baseline on ASR and entity recognition, with 1.49–2× decoding speedup.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06807","ref_index":33,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems","primary_cat":"cs.CR","submitted_at":"2026-07-07T21:12:27+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Activation-space divergence detects and corrects compromised LLM agents in multi-agent systems without interaction graphs or synchronized rounds, outperforming graph baselines especially under async stealthy attacks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06765","ref_index":31,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"When and How to Ask: Dynamic Preference Elicitation Strategies for Conversational Recommendation","primary_cat":"cs.IR","submitted_at":"2026-07-07T19:54:00+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Optimal preference elicitation in conversational recommenders is stage-dependent (attributes early, items later), and a MoE model trained on a new annotated dataset improves offline recommendation and response quality.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06720","ref_index":35,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning","primary_cat":"cs.AI","submitted_at":"2026-07-07T18:36:04+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"When reflections localize early errors, in-context search solves exp-small pass-rate problems with poly sequential attempts; otherwise it offers no asymptotic gain over parallel sampling, and the update is learnable and RLVR-optimal.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06534","ref_index":47,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models","primary_cat":"cs.CV","submitted_at":"2026-07-07T17:39:41+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Topology-aware attention over hierarchical scene graphs lets a 3D-LLM ground, caption, and answer questions across multi-room homes, with large gains on a new HM3D benchmark.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06482","ref_index":32,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities","primary_cat":"cs.CL","submitted_at":"2026-07-07T16:43:05+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"DataGovBench is a new benchmark using 178 large multi-tabular government datasets showing state-of-the-art LLMs and agents achieve below 40% QA accuracy and below 50% insight scores, far from real-world data analysis demands.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06402","ref_index":67,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"What Images Cannot Say: Language-Guided Olfactory Representation Learning","primary_cat":"cs.CV","submitted_at":"2026-07-07T15:31:55+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SCENT uses VLM-generated scene descriptions as a semantic bridge to align electronic-nose signals with visual and textual embeddings, improving cross-modal smell retrieval and enabling object-context odor disentanglement.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06262","ref_index":1,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Optimal Transport Q-Learning for Flow Policy Steering and Acceleration","primary_cat":"cs.RO","submitted_at":"2026-07-07T13:29:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Advantage-weighted conditional optimal transport flow matching simultaneously steers flow policies toward high-value actions and straightens their integration paths, enabling 2-3 step inference while improving task success.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06202","ref_index":51,"ref_count":2,"confidence":0.98,"is_internal_anchor":true,"paper_title":"UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods","primary_cat":"cs.DC","submitted_at":"2026-07-07T12:25:16+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"UBEP replaces BSP All-to-All for MoE on multi-tier superpods with dependency-driven kernel decomposition, topology-aware token scheduling, and Data-as-Flag atomics, cutting All-to-All latency up to 52.4% and TPOT up to 11.1%.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06157","ref_index":173,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"LLM Agents for Deliberative Collaboration: A Study on Joint Decision Making Under Partial Observability","primary_cat":"cs.CL","submitted_at":"2026-07-07T11:34:10+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A benchmark for LLM agents in partially observable joint decision-making reveals that deliberation challenges current models but can enable reflection and error correction.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06109","ref_index":67,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"RoME: Robust Mixture of Low-Rank Experts against Multiple Adversarial Perturbations","primary_cat":"cs.CV","submitted_at":"2026-07-07T10:20:51+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Low-rank additive experts with dual-scale gating and threat-guided diversification improve multi-perturbation adversarial robustness by routing different threat types through distinct model pathways.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06065","ref_index":41,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SWE-Review: Closing the Loop on Issue Resolution with Agentic Code Review","primary_cat":"cs.SE","submitted_at":"2026-07-07T09:37:45+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"An agentic code reviewer that explores repositories to judge and diagnose AI-generated pull requests improves resolve rates from 27.5% to 56.9% and outperforms single-turn review baselines.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.06014","ref_index":3,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models","primary_cat":"cs.SD","submitted_at":"2026-07-07T08:57:00+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ORCA splits Q-Former queries into orthogonally constrained groups, reversing directional collapse and speaker-indistinguishability in audio-LLM connectors and gaining 26.4 points on SAKURA multi-hop reasoning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05992","ref_index":33,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages","primary_cat":"cs.CL","submitted_at":"2026-07-07T08:25:29+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"PluraMath extends PolyMath with human-validated math problems in 18 mid-to-extreme low-resource languages and benchmarks 27 reasoning LLMs, finding a persistent high- vs low-resource performance gap.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05971","ref_index":16,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking","primary_cat":"cs.MM","submitted_at":"2026-07-07T08:04:56+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"VTMR is a two-stage video-to-music recommender: joint audio-visual-text retrieval of candidates, then temporal-sequence reranking, lifting R@10 to 18.3 and matching commercial preference.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05861","ref_index":18,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization","primary_cat":"cs.CL","submitted_at":"2026-07-07T05:34:25+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"MARGO mitigates thinking-induced hallucination in large reasoning models by using mixed-mode GRPO rollout groups that compare thinking trajectories against same-model non-thinking references.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05849","ref_index":28,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script","primary_cat":"cs.CL","submitted_at":"2026-07-07T05:12:13+00:00","verdict":"CONDITIONAL","verdict_confidence":"UNKNOWN","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A multi-stage pipeline that pivots Traditional Mongolian script through Cyrillic before translation improves MT quality across multiple backbones and target languages, and generates useful synthetic parallel data.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05804","ref_index":29,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training","primary_cat":"cs.AI","submitted_at":"2026-07-07T03:56:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"TurnOPD improves on-policy distillation for long-horizon agents by adaptively budgeting rollout depth and progressively shifting KL loss from token-level to turn-balanced weighting, achieving up to 2.29x faster training with better accuracy on ALFWorld, WebShop, and Multi-Hop Search.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05721","ref_index":35,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SpanUQ: Span-Level Uncertainty Quantification for Large Language Model Generation","primary_cat":"cs.CL","submitted_at":"2026-07-07T01:09:46+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"A DETR-style probe distills multi-sample claim uncertainty into single-pass span detection and continuous Mixture-of-Beta scores, outperforming baselines on a new 293K-span benchmark.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05716","ref_index":22,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models","primary_cat":"cs.CV","submitted_at":"2026-07-07T01:00:51+00:00","verdict":"ACCEPT","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Scene-graph-aligned SFT plus node-as-proxy GRPO rewards let small MLLMs outperform larger baselines on fine-grained visual reasoning tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05708","ref_index":47,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Akashic: A Low-Overhead LLM Inference Service with MemAttention","primary_cat":"cs.AI","submitted_at":"2026-07-07T00:06:22+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.5,"formal_verification":"none","one_line_summary":"Akashic’s MemAttention plus locality-aware placement improves agent task accuracy by up to 10.2 points and throughput by up to 1.21× over prior memory systems across four long-horizon workloads.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05394","ref_index":3,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Weak-to-Strong Generalization via Direct On-Policy Distillation","primary_cat":"cs.LG","submitted_at":"2026-07-06T17:59:58+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Transferring a weak model’s RL-induced log-ratio policy shift on a strong student’s own rollouts raises AIME accuracy more cheaply than imitating the weak teacher or running matched-step RL on the student.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05391","ref_index":25,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"LLM-as-a-Verifier: A General-Purpose Verification Framework","primary_cat":"cs.AI","submitted_at":"2026-07-06T17:59:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Expecting over scoring-token logits yields continuous, scalable verification that improves agent trajectory selection and dense RL rewards across coding, robotics, and medical benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05355","ref_index":41,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Faithfulness to Refusal: A Causal Audit of Neuron Selectors","primary_cat":"cs.CL","submitted_at":"2026-07-06T17:33:36+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A causal audit via neuron-row zeroing shows attribution methods (LRP, IG) faithfully identify dispensable neurons and can install refusal behavior, while rank-stability proxies systematically miss selector failures.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05339","ref_index":8,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"TREK: Distill to Explore, Reinforce to Refine","primary_cat":"cs.LG","submitted_at":"2026-07-06T17:21:16+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"TREK uses verified teacher proposals to expand a student model's exploration support before standard GRPO refinement, improving performance on hard math and agentic tasks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05196","ref_index":188,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Unified Audio Intelligence Without Regressing on Text Intelligence","primary_cat":"cs.CL","submitted_at":"2026-07-06T15:11:57+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Audex unifies audio understanding and generation on a strong text MoE backbone with multi-stage SFT plus text-only Cascade RL, matching open SOTA audio scores while mostly retaining text capability.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.05184","ref_index":13,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Rethinking On-Policy Self-Distillation for Thinking Models","primary_cat":"cs.AI","submitted_at":"2026-07-06T15:01:35+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Privileged-context on-policy self-distillation degrades thinking models' long-budget accuracy by suppressing forking and self-correction behaviors, while helping instruction-tuned models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02502","ref_index":14,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"DemoPSD: Disagreement-Modulated Policy Self-Distillation","primary_cat":"cs.LG","submitted_at":"2026-07-02T17:58:29+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.5,"formal_verification":"none","one_line_summary":"Disagreement-modulated reverse-KL barycenter targets let on-policy self-distillation attenuate privileged leakage while preserving exploration, beating SDPO and GRPO on SciKnowEval and GPQA.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02464","ref_index":42,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Will Scaling Improve Social Simulation with LLMs?","primary_cat":"cs.CL","submitted_at":"2026-07-02T17:30:38+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Using 85 controlled and 35 public LLMs, the authors show social-simulation accuracy generally improves with compute, but some behavioral and low-resource tasks do not scale.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02460","ref_index":8,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation","primary_cat":"cs.LG","submitted_at":"2026-07-02T17:27:24+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Neuron-OPSD uses neuron activations to guide data selection and teacher construction for annotation-free on-policy self-distillation in LLMs, claiming better in-domain results without harming cross-domain performance or calibration.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02423","ref_index":25,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Neuron-Aware Active Few-Shot Learning for LLMs","primary_cat":"cs.LG","submitted_at":"2026-07-02T16:51:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"NeuFS selects active few-shot samples for LLMs by representing samples via neuron activation patterns and applying a dual-criteria strategy of diversity and neuron consensus to identify informative examples.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02402","ref_index":42,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Show Me Examples: Inferring Visual Concepts from Image Sets","primary_cat":"cs.CV","submitted_at":"2026-07-02T16:35:52+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A new task and architecture for inferring an unlabeled visual concept from a small image set and re-instantiating it in a query image, with experiments showing gains over VLMs.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02383","ref_index":105,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"MEDIAREF: A Public Knowledge Store for Media Background Checks","primary_cat":"cs.CL","submitted_at":"2026-07-02T16:20:28+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"MEDIAREF is a 21,921-document knowledge store covering 200 media outlets that enables cheaper, reproducible media background check generation; adding its evidence to LLM prompts modestly improves fact recall without reducing errors.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02374","ref_index":11,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models","primary_cat":"cs.AI","submitted_at":"2026-07-02T16:15:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"User-attribute memory induces measurable medium-to-large reasoning drift in LLMs above pragmatic noise, only partly reduced by GRPO/DPO post-training.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02303","ref_index":16,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"A Hippocampus for Linear Attention: An Exact Memory for What the Recurrent State Forgets","primary_cat":"cs.AI","submitted_at":"2026-07-02T15:19:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"HOLA pairs a compressive delta-rule recurrent state with a residual-selected exact KV cache and decoupled RMSNorm-gamma read, yielding lower perplexity than both standard linear attention and full-attention baselines on Wikitext and LAMBADA plus stronger needle-in-haystack recall.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02262","ref_index":5,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"CheckRLM: Effective Knowledge-Thought Coherence Checking in Retrieval-Augmented Reasoning","primary_cat":"cs.CL","submitted_at":"2026-07-02T14:50:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"CheckRLM extracts factual claims from reasoning chains, detects inconsistencies via RAG, and refines them with low-cost corrections to reduce error accumulation in Reasoning Language Models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02214","ref_index":33,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning","primary_cat":"cs.CL","submitted_at":"2026-07-02T14:22:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"SpeechCombine produces instruction-following SLMs via speech pre-training followed by direct weight combination with the text LLM instruction delta, without any speech instruction tuning.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02119","ref_index":31,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation","primary_cat":"eess.AS","submitted_at":"2026-07-02T12:55:11+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Extends vLLM with delay-pattern de-interleaving, multi-stream sampling, and co-scheduled CFG to achieve 80% of non-CFG throughput for unified audio tasks while open-sourcing the pipeline.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02118","ref_index":27,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Enhancing Fitness Intelligence through Domain-Specific LLM Post-Training","primary_cat":"cs.AI","submitted_at":"2026-07-02T12:53:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"FitOne-8B/32B models improve average scores on ACSM-EP and NSCA-CSCS certification exams by up to 12.73% over base Qwen3 while retaining general capabilities.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02073","ref_index":15,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Evidence-State Rewards for Long-Context Reasoning","primary_cat":"cs.AI","submitted_at":"2026-07-02T12:11:49+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Maven is an RL method using answer-conditioned evidence-state values to assign rewards to add, link, and drop actions on evidence memory, outperforming outcome-only baselines on LongBench v2, LongReason, and RULER.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02052","ref_index":41,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Mitigating Package Hallucinations in Large Language Models via Model Editing","primary_cat":"cs.SE","submitted_at":"2026-07-02T11:27:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"BOUND refines LLMs' package-validity boundary via targeted editing to cut package hallucination rates by 79.9% on edit prompts and 65.4% on unseen prompts in recommendation tasks while generalizing to code generation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.02047","ref_index":21,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets","primary_cat":"cs.CL","submitted_at":"2026-07-02T11:14:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"OpenSafeIntent benchmark shows models fail to calibrate safety across intent shifts in matched dual-use prompts, indicating current evaluations are insufficient.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01977","ref_index":16,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"OntoLearner: A Modular Python Library for Ontology Learning with Large Language Models","primary_cat":"cs.AI","submitted_at":"2026-07-02T10:09:03+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"OntoLearner supplies the first cross-domain ontology collection and benchmarking infrastructure for LLM-driven ontology learning, finding that failure scales with ontological complexity instead of model size.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01964","ref_index":2,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing","primary_cat":"cs.CL","submitted_at":"2026-07-02T09:57:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"Unsupervised LLM input rewriting for dialogue discourse parsing introduces more regressions than repairs and has a practical ceiling on error repair, necessitating rewritability prediction for effective use with frozen parsers.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01953","ref_index":27,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Underspecification does not imply Incoherence: The Risks of Semantic Collapse in Coding Models","primary_cat":"cs.SE","submitted_at":"2026-07-02T09:43:15+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Coding LLMs exhibit detrimental semantic collapse on underspecified prompts by producing consistent but incorrect code rather than incoherent variations, affecting 3-32% of tasks across MBPP, HumanEval, and LiveCodeBench.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01927","ref_index":3,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"TUDUM: A Turkish-Thinking Reasoning Pipeline for Qwen3.5-27B","primary_cat":"cs.CL","submitted_at":"2026-07-02T09:22:36+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":3.0,"formal_verification":"none","one_line_summary":"TUDUM applies LoRA-based SFT on 15,991 Turkish reasoning examples followed by GRPO reinforcement learning on Turkish math problems to a 27B Qwen model, producing shorter Turkish reasoning traces with mixed benchmark results.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01883","ref_index":46,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"PairCoder++: Pair Programming as a Universal Paradigm for Verified Code-Driven Multimodal and Structured-Artifact Generation","primary_cat":"cs.CL","submitted_at":"2026-07-02T08:36:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"PairCoder is a two-agent pair-programming method that leverages toolchain verification oracles to improve LLM generation of verifiable structured artifacts on 17 benchmarks across seven models.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01876","ref_index":32,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"SAB-LVLM: Significance-Aware Binarization for Large Vision-Language Models","primary_cat":"cs.CV","submitted_at":"2026-07-02T08:30:00+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"SAB-LVLM proposes a significance-aware binarization technique for LVLMs that uses modality-guided Hessian-based maps to reweight binarization errors and improve performance under 1-bit constraints.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01855","ref_index":49,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Regression Accumulation in Multi-Turn LLM Programming Conversations","primary_cat":"cs.SE","submitted_at":"2026-07-02T08:15:40+00:00","verdict":"CONDITIONAL","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Regression accumulation affects 40-73% of 8-turn LLM coding tasks on extended HumanEval+/MBPP+ benchmarks, with verification gates improving final-turn pass rates on prior tests.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01831","ref_index":58,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Lynx: Progressive Speculative Quantization for accelerating KV Transfer in Long-Context Inference","primary_cat":"cs.DC","submitted_at":"2026-07-02T07:52:43+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"Lynx partitions KV cache bits into anchor and residual streams for progressive transfer, enabling speculative decoding on partial data followed by verification to match BF16 accuracy at 4-bit-like TTFT.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01830","ref_index":5,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Many Voices, One Reward: Multi-Role Rubric Generation for LLM Judging and Reward Modeling","primary_cat":"cs.LG","submitted_at":"2026-07-02T07:50:38+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"MRRG elicits evaluation criteria from multiple complementary roles to build rubrics that outperform single-role baselines for validating LLM preferences and providing rewards in RLVR.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01814","ref_index":78,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support","primary_cat":"cs.AI","submitted_at":"2026-07-02T07:30:19+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"MMIR-TCM is a multimodal framework using MLLM, memory-SAM, and RAG that claims to outperform GPT-4o and Gemini on TCM tongue diagnosis tasks via a new dataset and custom metric.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2607.01800","ref_index":10,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"Do LLMs Truly Generalize in the Molecular Domain? A Perturbation-Based Analysis","primary_cat":"cs.LG","submitted_at":"2026-07-02T07:16:28+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Molecular LLMs suffer large performance drops from single graph edits; in-context tuning on similar molecules partially widens their reliable region.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":100,"offset":0}}