Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2604.06683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.06683 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:29:09.202017Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact8
  • verified fuzzy25
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 82741301-51c7-42dc-aa9a-70212ffc191e · outbound

This paper cites Mathqa: Towards interpretable math word problem solving with operation-based formalisms.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Mathqa: Towards interpretable math word problem solving with operation-based formalisms

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.022252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:9bb9bd8e8ab7b4e9b90a9ed63b25800ccbde972a438453e142bde5e5fde81b48

Observation 8b2ecd7d-23f2-4028-b36b-623bd33ec61d · outbound

This paper cites Program Synthesis with Large Language Models.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Program Synthesis with Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:30:53.126418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:9095118fe80978b18e70c68e82ff441481e17a31cde0a6f13ca3a571403f449d

Observation e8f42a16-a396-4752-b0b7-8c2f24543b01 · outbound

This paper cites Software architecture documentation in practice: Documenting architectural layers.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Software architecture documentation in practice: Documenting architectural layers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.027649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:8e5d86e5cd36e6240eddd392366097e1212a899b8ab08c25c251539a083d3c83

Observation a1eb5317-b036-4790-b7bf-b609390b4492 · outbound

This paper cites Assessing the suitability of large language models in generating uml class diagrams as conceptual models.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Assessing the suitability of large language models in generating uml class diagrams as conceptual models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.032834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:7d3bdfb085bc94f17eea49fc05fe29a80ccd8d1a37bab033a5a0604ec4c38b96

Observation 08a4946e-01ff-4797-870a-86ea7f3bfbf0 · outbound

This paper cites On the assessment of generative ai in modeling tasks: an experience report with chatgpt and uml.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation On the assessment of generative ai in modeling tasks: an experience report with chatgpt and uml

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.019308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:c0a553aed4c358f21b1ee0af406b9c7d026b8f0123b38eccdc39ee06a472660a

Observation 14a9cace-8f07-4bc0-9da5-692b80e7f91a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Evaluating Large Language Models Trained on Code

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:30:53.141148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:05f6773c489946a052ab1c32f5af382ab7873018260b2bbc0ab0f1f3bf35841c

Observation de570f53-e4db-454a-ad3c-17013dcae278 · outbound

This paper cites Can llms generate architectural design decisions?-an exploratory empirical study.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Can llms generate architectural design decisions?-an exploratory empirical study

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.030400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:181cd24bd7f26ceab4c843fe4dc13689416f1a0b6a6a4299eea9409391e39603

Observation c5c24daa-1288-4ca6-ba1d-6d139e1ca80c · outbound

This paper cites CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation CrossCodeEval: A Diverse and Multilingual Benchmark for Cross-File Code Completion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.153123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:d1a92ca2eaa68fb05f0542f0df113dda38803f9a3fd929ed15d7172c02ede955

Observation c6cdb260-851c-441f-af74-1efa4759c57f · outbound

This paper cites ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:53.119918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:dce960abdd40ddfed70646fdb39575afe916f7dbfb05db04bfe1b5f168e83889

Observation 2093ceb7-6e83-4e33-8950-c90976c7c51d · outbound

This paper cites Erni and C.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Erni and C

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.024759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:1c4e704091fcce7a8cd3ce18a1df395b6860c74915df172ddb01405501995906

Observation 67dc89eb-c377-4cbf-9f32-f949bd3345fa · outbound

This paper cites an unresolved cited work.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Unresolved cited work

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:30:42.834015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:4739ad0d4873aa1cfba3d856f857cc308adcf49f4d3aac08dde04d49c39dda66

Observation 19e60196-bab4-4112-9814-f85aa389a6d6 · outbound

This paper cites A survey on llm-as-a- judge.The Innovation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation A survey on llm-as-a- judge.The Innovation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.035335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:7c7c278e9b8b4a3fe13ccf5aa98fc306d874a8ad85bea1bb41131d072b785809

Observation 864208b2-b7c8-484a-90e4-0355648f2f74 · outbound

This paper cites MetaGPT: Meta programming for a multi-agent collaborative framework.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation MetaGPT: Meta programming for a multi-agent collaborative framework

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.037781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:979823ecfca233bc4550da5591de5a53fafb225b759f4435263cc40292620445

Observation 04b7dddc-bda8-47ab-a31d-86aa83496536 · outbound

This paper cites Large language models for software engineering: A systematic literature review.ACM Transactions on Software Engineering and Methodology, 33(8):1–79.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Large language models for software engineering: A systematic literature review.ACM Transactions on Software Engineering and Methodology, 33(8):1–79

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.052729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:7e60b3a9c2a04689021bfae1f65f88e250eb4b7f4db0b487d36b985c13f18d11

Observation 0319fbe5-e359-4dbd-a0bb-767511f6469a · outbound

This paper cites Role of ai in requirements engineering.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Role of ai in requirements engineering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.047736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:a6cf452bd553f12b07447035aee979aaaffc0f154fb43d391f69bfa1f0082cbc

Observation e6e759d7-9800-43e2-a6be-6d2279fe4791 · outbound

This paper cites The unified modeling language reference manual.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation The unified modeling language reference manual

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.049946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:31f7f70af57340e00725475ecbe72ca91de0288708a0513dd8008ec520cf98e6

Observation f561b1c9-a849-4f95-a97b-443185be38ed · outbound

This paper cites Testgeneval: A real world unit test generation and test completion benchmark.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Testgeneval: A real world unit test generation and test completion benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.045674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:17838fce8acd9cde31b38d7c2066fd0cb684f4b0310d670864c8d2d5eaef49e4

Observation f9463452-975f-4659-905c-a9c440bb1d46 · outbound

This paper cites Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Survey of hallucination in natural language generation.ACM computing surveys, 55(12):1–38

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.057860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:4fb19d0b8d427fb6e0808450464f326ce59487ec8f311f028bbe72bbca785bdf

Observation d5b74280-14a8-4aa0-ad23-a283df4b56fa · outbound

This paper cites Prompting large language models to tackle the full software development lifecycle: A case study (devbench).

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Prompting large language models to tackle the full software development lifecycle: A case study (devbench)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.055186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:73fbed0df045929871dee5c7acd4099f99b63c8ac92150b71d52a573d97897da

Observation 71a01ba6-55e0-4628-8dbf-1bdcb945c2ee · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Rouge: A package for automatic evaluation of summaries

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.078961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:bad047dc950bcb1118f511d0fb6d00f7c774cbe6fe3a3cfa110c1fd5e64d9cb0

Observation 6da977d5-731e-4e0b-ab3f-4a838a1c59cf · outbound

This paper cites C4 model: a research guide for designing software architectures.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation C4 model: a research guide for designing software architectures

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.076599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:12fd3488fced32df366f8c77e942359db78ecdf0c096e65cc8b78c72d6e9aff1

Observation f6c957cb-3acb-45ff-b1c0-bb24c0c36003 · outbound

This paper cites Rec- ommended practice for architectural description of software intensive systems.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Rec- ommended practice for architectural description of software intensive systems

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.040254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:649ff44a19f872e2da3fbab2c3bf0e57cdb5ebd2fe62f0c9e2a86545c52046fd

Observation bb23b425-4288-411a-b143-b7b846754eca · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Bleu: a method for automatic evaluation of machine translation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.063206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:1c4409182a62be200ac7d1db424ab3f4db7ca1f782d7028f28b40935ef90d017

Observation 9162af84-f866-4733-9947-c930167df4d0 · outbound

This paper cites Sys- tematic literature reviews in software engineering—enhancement of the study selection process using cohen’s kappa statistic.Journal of Systems and Software, 168:110657.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Sys- tematic literature reviews in software engineering—enhancement of the study selection process using cohen’s kappa statistic.Journal of Systems and Software, 168:110657

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.067723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:acb19d1036604b2588aa29ba677c258276839663a1a3d03c9a3a1acc29b41583

Observation 06a649f3-b440-4879-8271-36d2b5dd7774 · outbound

This paper cites Software Architecture Meets LLMs: A Systematic Literature Review.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Software Architecture Meets LLMs: A Systematic Literature Review

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.193676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:4145fb0d127c656bd0207d405d0bc31c16c5981bf6159ab4f9b2b6166b442e6c

Observation 4af3e750-4c75-4789-84a8-2aa44fd6d5a9 · outbound

This paper cites MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:30:53.200344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:ec65c8b4fc686139b8bfb676548b7939ade1757b9b9b911d9ed33faea1296fc9

Observation 0cacff00-a0b5-4b3c-86be-7e00ffac72d5 · outbound

This paper cites Application of the tree-of-thoughts framework to llm-enabled domain modeling.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Application of the tree-of-thoughts framework to llm-enabled domain modeling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.074443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:38f61bd612666ae53868dcec3016508bbcb4f9ad9369b6eb3284e7df398ead0c

Observation d8a70f91-457a-4af0-99d7-cb4000cfe1e7 · outbound

This paper cites Collaborative llm agents for c4 software architecture design automation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Collaborative llm agents for c4 software architecture design automation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.181552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:78232d22bbaf0c2c7afb0e9fbedcb64c3ac44b0f49ba2ec38b1d7631e5e894ff

Observation 46c1c2dc-a90d-43a1-b69c-dd8133990c65 · outbound

This paper cites Contest: A unit test comple- tion benchmark featuring context.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Contest: A unit test comple- tion benchmark featuring context

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.072223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:a4548b98c46cf39cfcdef49a5bdc35fb415e9d87006f1df77b7d97009978b868

Observation 636e62ab-10e8-4820-8303-bbc44aa03ccd · outbound

This paper cites Testeval: Benchmarking large language models for test case generation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Testeval: Benchmarking large language models for test case generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.069913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:cbb54a03a4a977b4b4ffd546ae114a307a8ae956e05e942edc1cc0574edc5143

Observation c48d842a-8a82-40a1-bbb6-c1333ce61a16 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:08:06.116354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:f0727e354c6f6bbbf4b6658855780178f0c5c870b754de5336fb46b0446b360c

Observation 74898c61-df5d-4224-8563-710129465bea · outbound

This paper cites Repocoder: Repository-level code completion through iterative retrieval and generation.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Repocoder: Repository-level code completion through iterative retrieval and generation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.065435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:322a97850cd49a2ce4919a63f5c514c9575b7338ffa0acda9716ab104ab51c12

Observation 9f6baf82-1bc8-46de-baa5-fb76f067c1ed · outbound

This paper cites Towards realistic project-level code generation via multi-agent collaboration and semantic architecture modeling.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Towards realistic project-level code generation via multi-agent collaboration and semantic architecture modeling

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:30:53.167223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:5c9a9025221b17ba5450b92aa3ea1db0849e827075e90aba6670a86c20bcff13

Observation 85d692d5-015c-446c-a72e-0c1982abf198 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.060720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:d9c6559bcf2f049c7c6bfc96f4ee586a1cadf7d2f0d7b2d5013dc8712add962f

Observation 70757e43-0783-4fb0-ba10-7d1c729123f9 · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:33:53.042595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:29:09.202017Z digest=sha256:e706151704270ff755b062952970ac50661f7505ad33ae4780afb49198c04137

Pith citing papers

No inbound Pith citation observations are available.