Pith. sign in

Paper Citation Record · LEDGER

AlignBench: Benchmarking Chinese Alignment of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2311.18743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18743 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:26:06.931661Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca841481-233e-462b-bdbd-1c2da779d8ed · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:08:05.831970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:4f640aeb23156619dab15f132c64513c1fae1257b2e6bbb7984b7bf6b413eb4e

Observation 512de56b-abb6-4b6d-bed1-c997bf629b57 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 175

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:44:38.323496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:d800021bde4b394b116b376ebac256d72ee539cd99c454964ab493ca1167bdc5

Observation b7a622bc-38a5-4a9e-81e6-929238bb9bcd · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.438167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:540e554981e63e7f4a506710a3abdf0f5bf88da2e7950234568236a10527ee24

Observation 0c5619da-cf35-4326-9b10-1b2861cc74ba · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.299351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:c518332e1fbe9d0dd5178591e0cda7a929adc085f136da136217886310c3e8ee

Observation 2effc0c7-cc26-48f1-b2bd-3739b840c7f0 · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:06.931661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:06.931661Z digest=sha256:3fd6a48cdcadc8a0f22cc1dc9ec20de0b5535510db5613c995bd010dada160c7

Observation c96a5d48-19fd-4db4-8fa2-81e2d1274caf · inbound

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese cites this paper.

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:35.123660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:35.123660Z digest=sha256:a99b410e5db57ccdbcccce6995f7c695fd031c98076bdc74827f00f718e48271

Observation df62893e-d5f8-4916-a8f8-3c71b13c46a2 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.837491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.837491Z digest=sha256:cc5f28c7e8f50cfb144aba52e953eab3a90c6ec6348c950c3a0d676e5ad531d9

Observation 5982dec6-ecb1-4708-abd5-cc5daca3049b · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.030027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.030027Z digest=sha256:7616cf22017a5c7d3ae34aaa1eb5a6258a23a98701c8f404c1084fc42c2df7c0

Observation 71fbc989-12a5-4017-a47c-e9e65b85ced1 · inbound

Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use cites this paper.

Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.485869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:27:19.485869Z digest=sha256:5364ab32928434b360760c7978db428c26ebdb1bb96512d52136e0f928a7e7e0

Observation e1774273-c23d-4431-9cdd-c9135588f87a · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:01:09.994861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:3c10b321829cca19988ee99ffa5b578e9e6ac75c678304fb3fdc570b6fb2181f

Observation 2a47dbe2-5407-4ed4-b574-e7ec36c316b5 · inbound

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework cites this paper.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.510620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.510620Z digest=sha256:e97c8d638bb7bea4f66e3e5f90d2ec43a838d65508d3760afb2fbb03cdba7a60

Observation 2285295d-a65f-46bd-b053-39e90300167b · inbound

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards cites this paper.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.444272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.444272Z digest=sha256:757adec31b71924eebe089630e987ee79501517ac3a840e6440384db99962e1c

Observation 6463cf01-e5ad-47ae-b3f7-88993610ba89 · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:22.372244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:22.372244Z digest=sha256:a4123c99715cbfcfd719089eea94cc32e545dd9b15105d4692b1b0e1a5623b7d

Observation 09e90af4-6a09-44df-80c5-dfea3a39aa04 · inbound

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications cites this paper.

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:28.871708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:26:28.871708Z digest=sha256:5f22532a5c8990a1e73cd5e1ae264f7002144d9ec1d3225ca6dda0756691a4a4

Observation c534328b-7f8f-4641-9ce8-0ae001490020 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:29.295048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:29.295048Z digest=sha256:1a0763cb4e78766f9f89f60ecd0d9afb2bd1cabde4845f81f03711eeb5e04a82