Pith. sign in

Paper Citation Record · LEDGER

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications

As of 13 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2607.06411.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06411 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:21:50.175416Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 25153bba-7766-48ff-8da2-350b24c0482f · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:49.708376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:49.708376Z digest=sha256:a87df404a396f48e5574d3c40b0e3c2a85ff8710614672284bb0e85af9a20250

Observation 7f097b81-7b4e-4618-9fa7-f37081d6ea79 · outbound

This paper cites SWE-Bench+: Enhanced Coding Benchmark for LLMs.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-Bench+: Enhanced Coding Benchmark for LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:49.775844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:49.775844Z digest=sha256:12fec37e34e98163f6ccb9bb76d3fbf3a69e4a479993cc16749c7a7b69a57886

Observation 20828afc-0fec-467e-9933-0bd1187ff184 · outbound

This paper cites Liang, S.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Liang, S

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:49.976954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:49.976954Z digest=sha256:ff1bb25efbbc67fb3475120c714dcb6724e10ad9b47ea2c908ee577f0b586bae

Observation c2414804-da3f-4c9a-a437-59199db19a10 · outbound

This paper cites Prathifkumar, N.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Prathifkumar, N

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.107696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.107696Z digest=sha256:672b1c336e12b66aa157fcc78e5a75da196e28c8c42bc64730ffca1d2f4ed141

Observation 49f5b076-22b7-43ab-a7d4-47a40943f5e4 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.110905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.110905Z digest=sha256:a85ed8e55ef6bc36c4912054bf93cb93367d724d8fdfad21d1debbd64c9c7777

Observation 67a67033-2161-4465-93c2-71fcfe096b3f · outbound

This paper cites Badertdinov, A.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Badertdinov, A

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.113607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.113607Z digest=sha256:21c05c68b483b6dc4a33567df0e57a9e52ce90a5dddb0af011243da55011c5e1

Observation ca14c0ea-849c-4e29-b863-c20cd752f562 · outbound

This paper cites SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.116198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.116198Z digest=sha256:71385289f10c4f0ecab420036ad184d392a62f99453513322f58c40e9fe6a554

Observation ba49c9f0-5eed-44f6-afb9-69243e087812 · outbound

This paper cites Chervyakov, A.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Chervyakov, A

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.119591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.119591Z digest=sha256:1094cf2646fa8971314156c6907b589b3e3b727d2214fa979472c0a413b96d4e

Observation 1bb4a7d5-bffc-4c28-9bb3-313c05f67bdb · outbound

This paper cites an unresolved cited work.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.122122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.122122Z digest=sha256:4c5d7f830a306f70cbb0292858189ca7c92130ccf7c66feb035925f58821443b

Observation 398a7918-8eb2-47f4-bcc1-6394e362bd83 · outbound

This paper cites Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.124329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.124329Z digest=sha256:e5994fd40d9effa0236845b2343b33321ecae31683b5cd08f3da01118068dff8

Observation ffc6dabd-4f26-488a-8066-685959e86027 · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.127140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.127140Z digest=sha256:36c2e42e5ccd43d18206bb58b0f4c5c4344be9076d9529adf009bec38eafb275

Observation 1327ca94-c8d9-4bf7-a999-1f1c80b399dd · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.129688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.129688Z digest=sha256:9e6ba5ad1acb84df9f36cb9b118e5a7f8071a144f96ad1a0efaa1c26cd37623e

Observation 448f7704-aa64-4cd9-b2ed-553bc86daad9 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications GAIA: a benchmark for General AI Assistants

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.132378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.132378Z digest=sha256:1a6db72203bc7b7042c75d4dae3512518d2f93f3e986a00f6a87bc0fe764cc3c

Observation 961877a6-14a1-4ed2-aa27-9d8b3536c034 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.134928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.134928Z digest=sha256:43daf13b5116de30a0673eb1183d3f1c1bf19e6299ab59fdd30b224818c1c901

Observation 666f9baf-c066-44ac-af9a-50f1e1eb8987 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Evaluating Large Language Models Trained on Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.137021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.137021Z digest=sha256:e42c8218eab4c05badfc05b1331b06b0376012a8b50c110641e580b5f0300c3d

Observation dbf58742-749a-4db9-b16e-c2eb66ffed38 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.139165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.139165Z digest=sha256:aaee9fbe9bef26bfced0ae2f9baa4a1b4fc873403e7dc2d7a2749403c59422ba

Observation 3b95bbdd-6e60-4b0f-9f58-9c40fb52cc0f · outbound

This paper cites AI Agents That Matter.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications AI Agents That Matter

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.141580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.141580Z digest=sha256:56f487ba26764b3e4704f8856fcec01607b66242856ce14e34f03106495e8862

Observation fc84386f-7bd2-4ece-b298-087048b31f8a · outbound

This paper cites Kapoor, B.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Kapoor, B

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.143937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.143937Z digest=sha256:a139e23c32903b5130da8909af6c7c1e358b0efb8ac14bdc2931dc0ff2b7edec

Observation 087ac713-6a1f-45db-ba43-78014aef5628 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Measuring AI Ability to Complete Long Software Tasks

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.145877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.145877Z digest=sha256:77f3f29d508bc2354dc9f7d50556209d4c9178cd47416a7ef46fe89b58ae1612

Observation 58fcc192-7e31-4da4-944f-af8b18859fb3 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.147998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.147998Z digest=sha256:562d408458868d27510957265461e65f5643ce2d41d8345d1a4fa17bfd5e8119

Observation dc4bb9bc-feab-490c-b45e-35719505f9ec · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.150187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.150187Z digest=sha256:4f769426eac29143f37a7b0648bdc8993dfdc8b2fd43892cbf364df77292ae22

Observation 220659b0-5cb8-4906-bd48-13c981ac6111 · outbound

This paper cites Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.152279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.152279Z digest=sha256:5a2ac3f22bd8b38d2a6adc97b33d8f28fc397e853de547995bc6781d9ef60f14

Observation 036e3ce2-2e83-4ef4-820a-aec381f84720 · outbound

This paper cites ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.154596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.154596Z digest=sha256:ff7009c5d414bb66424e6d177fdcb6eb76f7b50549fcb62a00039e8a174d2f1f

Observation 1e3db29c-191f-47ba-8277-1711100b9253 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.156961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.156961Z digest=sha256:a24d6de38bbd0177892ab71f903e08487bf937e90155a94033d5826014ef4b26

Observation 3877d2cc-b2d2-4b87-acf0-9b9500240cf6 · outbound

This paper cites The Leaderboard Illusion.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications The Leaderboard Illusion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.159277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.159277Z digest=sha256:9c84b961cc3b2bce563e34767b7020aac435468ac1c75fbbe55a8b29a3169e6f

Observation 9fcad652-77e6-4baf-95bc-178d169639c5 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.161820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.161820Z digest=sha256:aed03ab7d8258df2bfa16c5d27b2e754529833b187064deaccdc123e424cd056

Observation d9069d21-4c4a-469a-ad8c-0bd1f78a22aa · outbound

This paper cites MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.163874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.163874Z digest=sha256:a6b118b421d8bd73c49298e0bc642087df0b823e31c23b9f3051a9d639cf30aa

Observation aa951d9d-e589-41f3-9cc1-f482a6729cc1 · outbound

This paper cites Execution-Based Evaluation for Open-Domain Code Generation.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Execution-Based Evaluation for Open-Domain Code Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.166018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.166018Z digest=sha256:92133e666171e91cb62e2e9ff442570cb579ad97e03c4ddc1c35c5573c2afbd1

Observation 05b12bbe-234b-4b37-b526-95b20b9c02b4 · outbound

This paper cites HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.168183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.168183Z digest=sha256:2b86469dc06b60a57fcc7cccdf16aa85f56bfc320be1d34c1fb98429e40c1b01

Observation 1e75f657-3fea-452f-a6db-d3c2d022cfea · outbound

This paper cites Если FSM-контекст для апдейта недоступен — сценная обвязка спокойно пропускает апдейт дальше, а не падает.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Если FSM-контекст для апдейта недоступен — сценная обвязка спокойно пропускает апдейт дальше, а не падает

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.171250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.171250Z digest=sha256:b4d113e9d76a4dbbb48709badb6d58583157efe4619fb7359eb7d6a34605c727

Observation 7413c14b-3663-49a1-8334-82ff3405278a · outbound

This paper cites Для стратегий FSM, которые и так работают в разрезе чата (CHAT и CHAT_TOPIC), контекст должен определяться и без пользователя — по самому ча- ту/каналу.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Для стратегий FSM, которые и так работают в разрезе чата (CHAT и CHAT_TOPIC), контекст должен определяться и без пользователя — по самому ча- ту/каналу

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.173335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.173335Z digest=sha256:d2a2ab1edb62c056b66af9f5a6ecdab229afb809d6bf838935af10aed59441be

Observation 8adc0628-f9d6-4f33-b356-ab6f51a01607 · outbound

This paper cites Существующее поведение для личек и групп ломать нельзя: боты со сценами и без, которые сейчас работают, должны работать ровно как раньше.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications Существующее поведение для личек и групп ломать нельзя: боты со сценами и без, которые сейчас работают, должны работать ровно как раньше

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.175416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.175416Z digest=sha256:832d8a3cbbf9c1d820931ee137e790261f9e533ed336a3742560178c4a93fa97

Pith citing papers

No inbound Pith citation observations are available.