Pith. sign in

Paper Citation Record · LEDGER

SedarEval: Automated Evaluation using Self-Adaptive Rubrics

As of 19 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2501.15595.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15595 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:13:22.017520Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T22:40:57.908347Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:47:25.418736Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afacc28a-b5ad-479f-888b-b869ca940816 · outbound

This paper cites online" 'onlinestring :=.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.770257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.770257Z digest=sha256:448e376365cc39cc405431a6e56e546f027226748f87ff6be97792799ea3393e

Observation 0be2adbb-d7fe-413e-b04c-44e007e419dd · outbound

This paper cites write newline.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.776366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.776366Z digest=sha256:fc386aedb88c406992427ecebec47dfffc0ce579ab23b81c3198f3a64096d449

Observation d2183808-ddcf-44cf-a0fc-aed773d18927 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.782510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.782510Z digest=sha256:a2623382884b2833f745d24057e97667a7dee55cb463b52136ebf811e6bcba2d

Observation 21330030-275b-4061-8563-a0ecf2d42942 · outbound

This paper cites Qwen Technical Report.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.788524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.788524Z digest=sha256:7c70b0bb6a244c2ffea114f7f081a5a0bbf870d1ca5498baeec32a047a8e7d92

Observation 15844e17-d0fb-4b6b-a82a-46ca5ebf2d15 · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.794354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.794354Z digest=sha256:b6d1c46b66e0079be5b27cbcc69a62313f407fa000a892ae217cf72f2c234958

Observation b36ea18e-7d22-4abc-94ec-7fe1e28565a1 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.799924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.799924Z digest=sha256:e38b0c0a12c2bc96178869a4350495f3b703d02b1006557be3191045ec189c18

Observation 1dc2d5ae-cc34-429b-a023-758493073443 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.805266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.805266Z digest=sha256:dce76f77b29e52849f468c53d7d5de8a2d51e3f7fd2c08e66afdd79472c9d113

Observation f9fdd0ba-b398-47a2-937c-e6965792c14c · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.822626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.810761Z digest=sha256:3cb3791149b748936ea7b2c6879614736d27467af5cf3c05d4bace7edc441c24

Observation 8a9e6d06-1c5d-4b09-9542-eb6cf5d4d1bc · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.816627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.816627Z digest=sha256:5e0c81f0c74acc4f40695b383579a720748fd414301789be8439f1414b1e9723

Observation c6ddccf0-affc-43a9-8e71-b141380a2a31 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.821755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.821755Z digest=sha256:102942cdf005f309bbdb8462825067e50dc00e4cceee0d7cd50717d56604d506

Observation 0506ffbb-76e7-4577-968c-763e8d59b122 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.827306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.827306Z digest=sha256:2746271ad5535e36db87c0fddf6a2c03c9272e6fbdb8880c4557d6735904800f

Observation 5ef75484-31b6-41c9-8b23-b1141bc4eb13 · outbound

This paper cites Prometheus: Inducing Fine-grained Evaluation Capability in Language Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Prometheus: Inducing Fine-grained Evaluation Capability in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.832007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.832007Z digest=sha256:2c8452d6eada97082092eb8c7d1a2cf11cceca2da099a2cc47b283fe9f9c17e8

Observation dc6fb207-9096-4997-a5f1-528bfcb8ecfc · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.837440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.837440Z digest=sha256:ab644d63fbc21ff27806aed28b8db91e54b75c3156bd029c03be623ae4ad9574

Observation d110443c-3249-484f-914d-3430276b69e2 · outbound

This paper cites To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.842619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.842619Z digest=sha256:e9902683af2912330925291e7e3731ddbfb5755e9b34088bddfddf9bfbc8a969

Observation 20097df1-f407-460b-9b57-54aedf407901 · outbound

This paper cites Hurdles to Progress in Long-form Question Answering.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Hurdles to Progress in Long-form Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.847773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.847773Z digest=sha256:1ea190d743130a5171fde0c019edb5351959b8fc7f95778bdd966f02760defd8

Observation ab29fc4b-b4cd-4107-bd80-5174d73e0d14 · outbound

This paper cites Generative Judge for Evaluating Alignment.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Generative Judge for Evaluating Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.853247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.853247Z digest=sha256:6bec77844d78195adc263d476d9dfc60bc14ee227433a99a52aa7d42e673b468

Observation 9edd6644-54cc-4305-b6bb-1c0edf15904f · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 17

Resolution
verified exact
doi, observed 2026-08-10T14:13:22.111833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.858961Z digest=sha256:6c83eb68ce297abc58981b496a8e2c2600fa260eae264abb32e18335e56ec1c8

Observation 4e6bbf66-8e4e-400a-bab1-de64c53ce0f6 · outbound

This paper cites Hashimoto.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Hashimoto

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.864027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.864027Z digest=sha256:a36280ef412bc7be1a329a73d96a4a8cb42df53ed24a615651173cfda33ff99b

Observation b149cb7b-5fe7-4f7c-8f0c-0d9c36a6e0d7 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.869578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.869578Z digest=sha256:fbe179df11d3aaae0b9ff5771941731ade47310f961b059bda319f13bc07836d

Observation caada853-b696-40cf-8891-8ca26600b5bc · outbound

This paper cites G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.874624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.874624Z digest=sha256:311cbed64d1705288611187157719ef345a4c3790ac33db79d2348d06dd4e1b6

Observation ec07fe9e-b3ad-4558-92f6-0412a8666a4d · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.879936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.879936Z digest=sha256:a8242bdc5c15333a42c027b8acf8576873511dbcbca49aadaf0e1ba4a7651ec5

Observation 0555734d-5b41-49ac-98e6-2f4c4c6260f3 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.885097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.885097Z digest=sha256:82cdd4d0ac3678ebae78e39f9070774f77618d9303d5b45b7072b0b64bb6195f

Observation 3eb7ea0c-d841-40e9-bdf6-21452aa0302b · outbound

This paper cites Jointly Measuring Diversity and Quality in Text Generation Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Jointly Measuring Diversity and Quality in Text Generation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.890184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.890184Z digest=sha256:3fe7ed9b8a02b161e815b4dd55bd2c5bd21e420fe85044f8ddfa11b21cdd7dce

Observation 1e713b82-7510-478c-b2c5-f11e011aa874 · outbound

This paper cites GPT-4 Technical Report.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.895039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.895039Z digest=sha256:706932b8ba97ea9a5da6b15028db7b64ec3172fa3eb65320d3d891d4e4de71fa

Observation 2597b9e1-2752-4656-b6ed-42128c90bb78 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.900281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.900281Z digest=sha256:f090ccb229f2f74fd2d72cf4f45806240d222d36a6d39390d2c751e9616dab36

Observation a9fe2212-3cf2-4cd3-a828-02ddc9eb5e87 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.905347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.905347Z digest=sha256:e9c54fcfe9f9b3694575b3615277b2b6617971ae142e911d29d74b19afaed523

Observation 0083f9ba-b59e-4cd8-87ea-443ae994392e · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.910721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.910721Z digest=sha256:33eae81bc222659e98543ca6dc296353fed9ef415858c4ea3214930c7405b38d

Observation d62d85ed-850e-4959-9650-24464ccdd08a · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.915799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.915799Z digest=sha256:d7562338643649f51bbdfc2782c4cb011e34740b9c78a0aa09fd2aa1bc841cc7

Observation 61c691ac-3f8d-4677-bafe-3107bd59c2a6 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 29

Resolution
verified exact
doi, observed 2026-08-10T14:13:22.059851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.921031Z digest=sha256:97f97685d4c63de80cfb5c4389c9c06bca034dc95cc67fc340b0c035d637193a

Observation 9b3d69d7-076b-4668-9be5-ae4d369b8ef0 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.762732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.926591Z digest=sha256:2d59e991b8d8f788d3981ce630f9e8d4f2238d8a12f8a58354741e18801c73a9

Observation 8c73b962-8878-457d-a139-1d8c3e93c7d0 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.746485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.931499Z digest=sha256:20cf0f1bf52444baa9d14e8fe2e12e27cd3083cceb0d77495b220e9eebe81385

Observation 41d5405a-e4c3-40c4-af34-c80bde0d1198 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.936511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.936511Z digest=sha256:b7af06edd041390bc6f143c10220bfd0f8c970347803cab8b595baaa81a5c72c

Observation 2b2633e7-f30a-4b16-929b-9fd1dff4b44a · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.942192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.942192Z digest=sha256:f185a804b80c71a92a6c0f40a5e33fa05fef36e7cf657e926222b3725ec10c8c

Observation ae11bcc1-d549-4849-b8b8-9977439b5004 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.947802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.947802Z digest=sha256:a8caa972e1003428d1846b19548514de4ae957a27fb55f91fbd4e1631903acc8

Observation bcec49dc-f423-497b-b297-6a16880bf4c2 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.718344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.954187Z digest=sha256:08bbf096b62cea0aa708a083bc8b6cdada7e889d0782c381cd9164d5b5993791

Observation 46964b07-857c-4b76-a9cd-77160b7ab21f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.959034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.959034Z digest=sha256:8ecdceb23669b41dcb3b103edcf2a6d518f7dea04bc193d28d238a53be25a2a3

Observation 9742c1b6-7a03-4915-aae1-11c20db53c59 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.964177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.964177Z digest=sha256:490e6bb4eae1adfdf15952a370f126231f07f9aa40ed617f31e9d1a5dd12ac02

Observation 4774d543-5df4-4e24-886f-d26bcb351cba · outbound

This paper cites PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics PandaLM: An Automatic Evaluation Benchmark for LLM Instruction Tuning Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.969582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.969582Z digest=sha256:371cf7928eef0475b98d29a56411924a7026712968793715535e1d44815a163a

Observation c26c8cb1-9deb-44ad-aa7a-70f4d0e034b4 · outbound

This paper cites FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics FLASK: Fine-grained Language Model Evaluation based on Alignment Skill Sets

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.975242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.975242Z digest=sha256:c885b120d6a553db5abbb58828b3db87c821cfdd77d327201ad01ea3961cd8f3

Observation a197efe6-5edc-43da-a24f-1a4356ac5eb7 · outbound

This paper cites GLM-130B: An Open Bilingual Pre-trained Model.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics GLM-130B: An Open Bilingual Pre-trained Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.981049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.981049Z digest=sha256:a8979ec1032a439d10a0b89a5b9ecf59f6a1fbd754749d3e1800c99b277b56fb

Observation 1d8fc984-641e-40de-966f-0d5f0b33c132 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics BERTScore: Evaluating Text Generation with BERT

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:21.987198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:21.987198Z digest=sha256:0fbc804b77deff5de2398cd3416d4bdcf32b6e640c69565612545d90ee4ddf90

Observation 958b4490-cddb-4dc2-a802-66fe1bcae7a5 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:13:22.674409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-10T14:13:21.993292Z digest=sha256:3ec15abdfbf365fbefa79c79621ddeb64ca58d7667cc11a9ccb90cbc151bb5a8

Observation 9a19d39d-aaeb-4962-9585-617053fe9728 · outbound

This paper cites an unresolved cited work.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.000656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.000656Z digest=sha256:91623a70b0f5af8cf419c45d467c1ccb416cf790907884b52eb3304027c33e20

Observation 6033d781-4e14-45f8-a673-6ad63a147e42 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.007889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.007889Z digest=sha256:1ebce2fec67d85ec61957522dbe806e1496b3754674039cdbb45dfd9363369fc

Observation 5f9bd9c4-c29c-48c5-8b43-79ba6ed27c83 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.012661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.012661Z digest=sha256:43379aaabf57dc5f9287f3073c37186eb8f19bae83e387004eccb80a39fd978e

Observation 3355d1db-e358-4897-b05c-d2c1d42b7d02 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

SedarEval: Automated Evaluation using Self-Adaptive Rubrics JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T14:13:22.017520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:13:22.017520Z digest=sha256:d5da9bf1efe72e7441e23433db341699a93424a7bd4be09077071765b605baa4

Pith citing papers

Observation 730f1a8c-26df-4854-9338-52e80b4c4f53 · inbound

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape cites this paper.

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape SedarEval: Automated Evaluation using Self-Adaptive Rubrics

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:27:25.939000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T18:54:30.970241Z digest=sha256:e7e8926e3118d64853d449f9a95747da589629f3d6d3414c48c1a11f5df68ab7

Observation 39177722-d2c1-4721-839c-43a2c8b17ccb · inbound

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape cites this paper.

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape SedarEval: Automated Evaluation using Self-Adaptive Rubrics

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:47:25.420086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-02T22:40:57.908347Z digest=sha256:e9f8f1b478ac9ef69da84ade0c7628d1f76ccf513ee95b6f40241c70ff1109b8