Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:50.175416Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2607.06411.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:50.175416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 25153bba-7766-48ff-8da2-350b24c0482f · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f097b81-7b4e-4618-9fa7-f37081d6ea79 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-Bench+: Enhanced Coding Benchmark for LLMs
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20828afc-0fec-467e-9933-0bd1187ff184 · outbound
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2414804-da3f-4c9a-a437-59199db19a10 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Prathifkumar, N
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49f5b076-22b7-43ab-a7d4-47a40943f5e4 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67a67033-2161-4465-93c2-71fcfe096b3f · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Badertdinov, A
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca14c0ea-849c-4e29-b863-c20cd752f562 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba49c9f0-5eed-44f6-afb9-69243e087812 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Chervyakov, A
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb4a7d5-bffc-4c28-9bb3-313c05f67bdb · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398a7918-8eb2-47f4-bcc1-6394e362bd83 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffc6dabd-4f26-488a-8066-685959e86027 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1327ca94-c8d9-4bf7-a999-1f1c80b399dd · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 448f7704-aa64-4cd9-b2ed-553bc86daad9 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications GAIA: a benchmark for General AI Assistants
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961877a6-14a1-4ed2-aa27-9d8b3536c034 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 666f9baf-c066-44ac-af9a-50f1e1eb8987 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Evaluating Large Language Models Trained on Code
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf58742-749a-4db9-b16e-c2eb66ffed38 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b95bbdd-6e60-4b0f-9f58-9c40fb52cc0f · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications AI Agents That Matter
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc84386f-7bd2-4ece-b298-087048b31f8a · outbound
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087ac713-6a1f-45db-ba43-78014aef5628 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Measuring AI Ability to Complete Long Software Tasks
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58fcc192-7e31-4da4-944f-af8b18859fb3 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4bb9bc-feab-490c-b45e-35719505f9ec · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 220659b0-5cb8-4906-bd48-13c981ac6111 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 036e3ce2-2e83-4ef4-820a-aec381f84720 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e3db29c-191f-47ba-8277-1711100b9253 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3877d2cc-b2d2-4b87-acf0-9b9500240cf6 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications The Leaderboard Illusion
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fcad652-77e6-4baf-95bc-178d169639c5 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9069d21-4c4a-469a-ad8c-0bd1f78a22aa · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa951d9d-e589-41f3-9cc1-f482a6729cc1 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Execution-Based Evaluation for Open-Domain Code Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05b12bbe-234b-4b37-b526-95b20b9c02b4 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e75f657-3fea-452f-a6db-d3c2d022cfea · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Если FSM-контекст для апдейта недоступен — сценная обвязка спокойно пропускает апдейт дальше, а не падает
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7413c14b-3663-49a1-8334-82ff3405278a · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Для стратегий FSM, которые и так работают в разрезе чата (CHAT и CHAT_TOPIC), контекст должен определяться и без пользователя — по самому ча- ту/каналу
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8adc0628-f9d6-4f33-b356-ab6f51a01607 · outbound
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Существующее поведение для личек и групп ломать нельзя: боты со сценами и без, которые сейчас работают, должны работать ровно как раньше
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.