Pith. sign in

Paper Citation Record · LEDGER

SedarEval: Automated Evaluation using Self-Adaptive Rubrics

As of 11 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2501.15595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15595 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:13:22.017520Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T22:40:57.908347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:47:25.418736Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afacc28a-b5ad-479f-888b-b869ca940816 · outbound

This paper cites online" 'onlinestring :=.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.770257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.770257Z digest=sha256:c142355985307937f4d880dfae7c6c6349a1c09dc734bc70249555667b0845ea

Observation 0be2adbb-d7fe-413e-b04c-44e007e419dd · outbound

This paper cites write newline.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.776366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.776366Z digest=sha256:1589705dd2a14dcfed2bf7baad72a612ac3128350f711f5df5db6372cb5df208

Observation d2183808-ddcf-44cf-a0fc-aed773d18927 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.782510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.782510Z digest=sha256:55aefd82d0c90423df201fab17f7685d2614e5c6942d149726d711583ab07230

Observation 21330030-275b-4061-8563-a0ecf2d42942 · outbound

This paper cites Qwen Technical Report.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.788524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.788524Z digest=sha256:3f5818de4a1b777589009b40231bde2821c5dd83495b87e6839ae11f28e7acd2

Observation 15844e17-d0fb-4b6b-a82a-46ca5ebf2d15 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.794354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.794354Z digest=sha256:e210ac6680615214f33294678d0c25a190feb6a307571f4c4da16374ab651129

Observation b36ea18e-7d22-4abc-94ec-7fe1e28565a1 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.799924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.799924Z digest=sha256:2268f5471001aafc921415a2efb766bf26df90a482959cbceef35be65b614fb1

Observation 1dc2d5ae-cc34-429b-a023-758493073443 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.805266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.805266Z digest=sha256:28f14552d302dc4c8c57390e832045c3d4811dd1896c99ac2fa0e636f4181fa1

Observation f9fdd0ba-b398-47a2-937c-e6965792c14c · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.822626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.810761Z digest=sha256:d3eefb98f9237902eebf13eb44b52dd974ab54b3e59d5c0e548f59e49b12bc96

Observation 8a9e6d06-1c5d-4b09-9542-eb6cf5d4d1bc · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.816627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.816627Z digest=sha256:14113aef6eff524761266e23199f70621be4fec62f9e90919b17768ea0ad82a0

Observation c6ddccf0-affc-43a9-8e71-b141380a2a31 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.821755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.821755Z digest=sha256:1689a20bed849bd920d507e448a7741a9f0ec349ce8786e0f434e4cb3264e29a

Observation 0506ffbb-76e7-4577-968c-763e8d59b122 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.827306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.827306Z digest=sha256:56892d0be363227e1654bb4c0803bea3ed710cd2ca30f1c5798c87f925d84d60

Observation 5ef75484-31b6-41c9-8b23-b1141bc4eb13 · outbound

This paper cites Prometheus: Inducing Fine-grained Evaluation Capability in Language Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Prometheus: Inducing Fine-grained Evaluation Capability in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.832007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.832007Z digest=sha256:b7a704793c68f733d3d15f1e7c692e46a8fb0fc151517e7293cc6678b38ea87a

Observation dc6fb207-9096-4997-a5f1-528bfcb8ecfc · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.837440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.837440Z digest=sha256:35b8d65dd3a27224d5bb440fe859a7da14c49e61053ee8b56fe540f6598d91a1

Observation d110443c-3249-484f-914d-3430276b69e2 · outbound

This paper cites To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.842619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.842619Z digest=sha256:3f4b03ce9e35e7a5062401f7a92c57fa78babc94574e995d4945c745fd150c8c

Observation 20097df1-f407-460b-9b57-54aedf407901 · outbound

This paper cites Hurdles to Progress in Long-form Question Answering.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Hurdles to Progress in Long-form Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.847773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.847773Z digest=sha256:1ac6a8b07517ab4980fd7707f8ed93e3c72eb05d86bf55d093a3d38825403bc9

Observation ab29fc4b-b4cd-4107-bd80-5174d73e0d14 · outbound

This paper cites Generative Judge for Evaluating Alignment.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Generative Judge for Evaluating Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.853247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.853247Z digest=sha256:d3dc822ad3c662688f8ca54a3aa231d5d32d39bab93b043b7a85d8e20d7a3144

Observation 9edd6644-54cc-4305-b6bb-1c0edf15904f · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 17

Resolution
verified exact
doi, observed 2026-08-10T14:13:22.111833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.858961Z digest=sha256:5183cce91e1e1695fffce119683a3db330e94f6471988698fcccb4b68d90828a

Observation 4e6bbf66-8e4e-400a-bab1-de64c53ce0f6 · outbound

This paper cites Hashimoto.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.864027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.864027Z digest=sha256:8cf0e0363273f6560d0060f07ad72574ffb3ffd22948a573829501b3b2dee1f6

Observation b149cb7b-5fe7-4f7c-8f0c-0d9c36a6e0d7 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.869578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.869578Z digest=sha256:71b0c0ad2b8233fccb325ac7aa4919e34447a4869d29d1c565660e6f7c0f7aa0

Observation caada853-b696-40cf-8891-8ca26600b5bc · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.874624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.874624Z digest=sha256:e12b17d5a6024ef39649229f563b151ab58771cb3b86d39ee3a06c759e6fb83c

Observation ec07fe9e-b3ad-4558-92f6-0412a8666a4d · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.879936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.879936Z digest=sha256:0f1c155320f35c8094837eab1fe092a1405105d88578ff795cc15c5b8a6f2a4d

Observation 0555734d-5b41-49ac-98e6-2f4c4c6260f3 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.885097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.885097Z digest=sha256:304272c9c23bf5308179f7ce46d4d079f29ce3ee3277aee670821afc162d376f

Observation 3eb7ea0c-d841-40e9-bdf6-21452aa0302b · outbound

This paper cites Jointly Measuring Diversity and Quality in Text Generation Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Jointly Measuring Diversity and Quality in Text Generation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.890184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.890184Z digest=sha256:9eca979b0d421554661fc22890ff486cfab77ea91a28f5198f5eff4dc4dc1ea6

Observation 1e713b82-7510-478c-b2c5-f11e011aa874 · outbound

This paper cites GPT-4 Technical Report.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.895039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.895039Z digest=sha256:ef17b5f3f535716fa70ebffcc0915ca35f3db151e6a4fc0e851ce2b64112e3b2

Observation 2597b9e1-2752-4656-b6ed-42128c90bb78 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.900281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.900281Z digest=sha256:606d0d006dccd1390ed76036eec8c38573e27fd80ecb0bd3ae9ef59b164b87d9

Observation a9fe2212-3cf2-4cd3-a828-02ddc9eb5e87 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.905347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.905347Z digest=sha256:1e607d076e6cb28aaa8548a680abf5090a8e795d27af7dc7249dcefda648b009

Observation 0083f9ba-b59e-4cd8-87ea-443ae994392e · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.910721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.910721Z digest=sha256:38450f40bd3e16bad411419f06ddd1dbb28281b4d016f6c89e84c5008b60969e

Observation d62d85ed-850e-4959-9650-24464ccdd08a · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.915799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.915799Z digest=sha256:bce331b72c39c00f2178ffd9481b0066def13d8042481fb5f9493e3cf454ddbb

Observation 61c691ac-3f8d-4677-bafe-3107bd59c2a6 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 29

Resolution
verified exact
doi, observed 2026-08-10T14:13:22.059851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.921031Z digest=sha256:2a66219ab3e2b53469d4ae7452e6bf54fa821b3034e7d1791e7e7894f0ed7ed7

Observation 9b3d69d7-076b-4668-9be5-ae4d369b8ef0 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.762732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.926591Z digest=sha256:fac1dede7e6a6864032c9a05a8d507a8d99db437f64b3c733ba69117366e4b0e

Observation 8c73b962-8878-457d-a139-1d8c3e93c7d0 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.746485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.931499Z digest=sha256:63caaab844f1d4dbc9cfde6d2f03b394768fb2dd3d8875749577114806310120

Observation 41d5405a-e4c3-40c4-af34-c80bde0d1198 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.936511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.936511Z digest=sha256:670a69d56aebe30fc1f7e694bd231558dfb1cc691d66a8eae32c8cb1e2ff1dbb

Observation 2b2633e7-f30a-4b16-929b-9fd1dff4b44a · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.942192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.942192Z digest=sha256:6b3fbf05e7642bd52c735ada9c02db533eadadc774a645a4127168c05e5a5b40

Observation ae11bcc1-d549-4849-b8b8-9977439b5004 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.947802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.947802Z digest=sha256:1bc03de8e883265083ae4ba9d4f95f89a17ee4cdcf470421fd14f09d21f99cab

Observation bcec49dc-f423-497b-b297-6a16880bf4c2 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.718344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.954187Z digest=sha256:8ebc922eb502bc42127bba97f07e65299fa43f77ce6a9a4a462cbecb13c2dfbe

Observation 46964b07-857c-4b76-a9cd-77160b7ab21f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.959034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.959034Z digest=sha256:99f0dd46ae46c46b8aaf1526c881f7be59029df413ba5efec8595db7403e9e54

Observation 9742c1b6-7a03-4915-aae1-11c20db53c59 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.964177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.964177Z digest=sha256:f31cf1922a9d688a029e44e12242dfbd739be662ce6b936ff80bb6b6a65c5768

Observation 4774d543-5df4-4e24-886f-d26bcb351cba · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.969582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.969582Z digest=sha256:ad9d175b6b6cf814ed6313d7b97595f2f3ad4034589dec1f1e593d6eb31a3e6f

Observation c26c8cb1-9deb-44ad-aa7a-70f4d0e034b4 · outbound

This paper cites FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.975242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.975242Z digest=sha256:b0d050cdef969f412823459fbb96eca3454660c0b495acf51939aa5c682f83dd

Observation a197efe6-5edc-43da-a24f-1a4356ac5eb7 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics GLM-130B: An Open Bilingual Pre-trained Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.981049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.981049Z digest=sha256:b0a3d812233f7bd67ca5b5ce1b9744690af90f699fbb8ba4e79d41c047ee9c4c

Observation 1d8fc984-641e-40de-966f-0d5f0b33c132 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics BERTScore: Evaluating Text Generation with BERT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.987198Z digest=sha256:77348d2ef914845ee68a384ec0c9c42e1ef698b7118ab38c196f7771a09a9cdb

Observation 958b4490-cddb-4dc2-a802-66fe1bcae7a5 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.674409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.993292Z digest=sha256:af60873b53b4ffc86fcac87a022c7838f7ad2e02d7e11ed77ff110263b17c710

Observation 9a19d39d-aaeb-4962-9585-617053fe9728 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.000656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.000656Z digest=sha256:13718d183814b657e7a1a76ef8942296a414b6ad74ed0de3e48be7f66550e344

Observation 6033d781-4e14-45f8-a673-6ad63a147e42 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.007889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.007889Z digest=sha256:669d6a37abb47b896c06f81db1147e1169a84fc7dd90b486aa1e27eada8b2dc7

Observation 5f9bd9c4-c29c-48c5-8b43-79ba6ed27c83 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.012661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.012661Z digest=sha256:08ef81557af32563ec167bfa6263c7146fd08122c2d2cfbf4e49dde9b295e0b6

Observation 3355d1db-e358-4897-b05c-d2c1d42b7d02 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.017520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.017520Z digest=sha256:8770da6755bd8c8fdbd36ef77d08dde27911d025b938931a438afacc42caba0d

Pith citing papers

Observation 730f1a8c-26df-4854-9338-52e80b4c4f53 · inbound

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape cites this paper.

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape SedarEval: Automated Evaluation using Self-Adaptive Rubrics

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:27:25.939000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T18:54:30.970241Z digest=sha256:7a25e3e092d3c1fb1c2410ce479f4696fc7f0318eaeda87299d4dcbe6db872ea

Observation 39177722-d2c1-4721-839c-43a2c8b17ccb · inbound

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape cites this paper.

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape SedarEval: Automated Evaluation using Self-Adaptive Rubrics

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.420086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T22:40:57.908347Z digest=sha256:6fa65f195a2b60152ef8510bfec104c8dd45864e45a3769a850748bae9231ea3