Pith. sign in

Paper Citation Record · LEDGER

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams

As of 9 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2606.22723.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.22723 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T07:04:21.555971Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6f6dc6d-7fe9-4f01-87ba-2a4a0304bae2 · outbound

This paper cites In: Naldi, M.C., Bianchi, R.A.C.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: Naldi, M.C., Bianchi, R.A.C

Reference 1

Resolution
metadata mismatch
doi, observed 2026-07-01T07:05:27.451600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:1bc0c364cbf135c684638a319726ecd24111ccb099e6f81f7855f4319ebfdddf

Observation 8aaa95cc-5364-4878-9a16-49ab4f223af2 · outbound

This paper cites Educational and Psy- chological Measurement20, 37 – 46 (1960).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Educational and Psy- chological Measurement20, 37 – 46 (1960)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.189627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:621b4a76451aec82c495181638e79d6fbe3db22e81d20a0904d1e021ad388b8e

Observation 6768d605-eaa4-4f0a-bd77-68dfca62459c · outbound

This paper cites an unresolved cited work.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:32:39.184125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:e576f2f3caf0b87964fb95d03b6fef3f00e680b8bf3890265c89681fabec2f65

Observation cf096ed5-4d03-4dff-84e5-49364cdfeaab · outbound

This paper cites In: International Con- ference on Learning Representations (ICLR) (2021).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: International Con- ference on Learning Representations (ICLR) (2021)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.180142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:f647f48dbb9c83f0b3b7b4881389e3e9222cc0ff93ac4036678d0a5525afd430

Observation c50c3c1a-dc09-4438-b367-781a115730b7 · outbound

This paper cites CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams CritiqueLLM: Towards an Informative Critique Generation Model for Evaluation of Large Language Model Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.220496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:f7566936d6df5f8a2e8f904f1b94b8471c1c2bf039729197eb1a0be4496c3e74

Observation ab4cfba3-2939-4087-b9f4-adc36828036e · outbound

This paper cites Sabi\’a-4 technical report.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Sabi\’a-4 technical report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:28.233471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:f2784fcc93271fbfe8496b2a28a3e454aca9259ff41033f7ff36c929494bc213

Observation 4633b72e-c1f7-4624-a852-2eaee7e97732 · outbound

This paper cites Biometrics pp.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Biometrics pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.182134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:2a8b9eb9e5c207a6487da9162703f41848351110d8bc2f39c07927c9d8e232f3

Observation 903c2315-5fbb-4c01-ac25-a02cb6a5e750 · outbound

This paper cites an unresolved cited work.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-07-06T16:32:39.187877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:ec465b945cb121064d9074f15ed37e6ca959ac558157cc8662e083352c080d48

Observation 8cc3dfe2-3c08-4624-9138-c49696b8fe6d · outbound

This paper cites Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.224466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:85306de956f1bc6a6f610928141e80bd4ca18d11c3e010d586b3661914dbf007

Observation 86f13892-1268-4f79-8134-52414b1ea590 · outbound

This paper cites Automatic Legal Writing Evaluation of LLMs.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams Automatic Legal Writing Evaluation of LLMs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T07:05:28.238188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:4025b9e1376c70f6c9a83b6c0b3d8bce777306591c8d2d344327a1e3b8acb8ce

Observation ae62caed-c618-49ef-8860-2ab6a44fa524 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:05:28.228708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:4aa05f538b0ba1ea56f26b8f1187e5e7f257f4f4fff42d0de0ca172f60a0d9de

Observation 61be3dab-fabc-48f8-8dba-bd09c9f92fbd · outbound

This paper cites In: Proceedings of ENIAC (2025).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: Proceedings of ENIAC (2025)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.186002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:6f36cbe69ed8527a1aa0540f7e6baf0bfdfbfe8dc2b904ada9a52f674d1004d9

Observation 12eb00bc-923a-46cb-9586-1717bdcca0e4 · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS) (2023).

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams In: Advances in Neural Information Processing Systems (NeurIPS) (2023)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T16:32:39.178157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:2103994e7d32c80a5ecc668933c66ccfc658b7318539b1bfbe409bd3160f5c0e

Observation bfa5fb53-4e67-4b36-8283-784db1cf8bab · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

BLUEX v2: Benchmarking LLMs on Open-Ended Questions from Brazilian University Entrance Exams AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T07:05:28.242446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:04:21.555971Z digest=sha256:65cd897969999cacf88d15ae228030b6efb96e69b0f9d78bf4dd7571d4ff0c34

Pith citing papers

No inbound Pith citation observations are available.