Pith. sign in

Paper Citation Record · LEDGER

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation

As of 17 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.14256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14256 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T02:44:17.432243Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52876c55-4ea9-4615-bc02-c38753137410 · outbound

This paper cites Security in LLM-as-a-Judge: A Comprehensive SoK.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Security in LLM-as-a-Judge: A Comprehensive SoK

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.210092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.210092Z digest=sha256:7b7768be4bdbc7ca30b1eb9ff1c3dc8587b415ff580a5b299a1c02c2774d61df

Observation 58511f56-6fb5-48e6-81c5-87a22d890e21 · outbound

This paper cites Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form qa.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Reference-guided verdict: Llms-as-judges in automatic evaluation of free-form qa

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.302235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.302235Z digest=sha256:b90301dcf3adabf0f5452b6b73f435d63205ea205d0f374959385afab6bc4333

Observation 25ec6ad5-7450-430d-8bc1-36f5a9a20c14 · outbound

This paper cites Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Reading Between the Pixels: Linking Text-Image Embedding Alignment to Typographic Attack Success on Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.401198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.401198Z digest=sha256:5ce812f76144fbd4e9340447cccecaf29346654fd14cc872d9f144a3747772fd

Observation 7c0bc200-52d7-411b-b170-f737ba754229 · outbound

This paper cites Rethinking fine- tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Rethinking fine- tuning when scaling test-time compute: Limiting confidence improves mathematical reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.549266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.549266Z digest=sha256:3de760e45fb3476c29fa8b66280a3a797aa739fad4f9b0d8ef82d18038b2344b

Observation d430814e-b2b2-4e82-9c3d-bf34a6ec2e5f · outbound

This paper cites Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Every Picture Tells a Dangerous Story: Memory-Augmented Multi-Agent Jailbreak Attacks on VLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.644542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.644542Z digest=sha256:f1de53bfa45792481627e9e567fb56e9e11e5d768598c966f8ce20a46b74a890

Observation 4697ea5a-7ce8-48f8-a8cb-ac68ac45c49e · outbound

This paper cites The side effects of being smart: Safety risks in mllms’ multi-image reasoning.arXiv preprint arXiv:2601.14127, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation The side effects of being smart: Safety risks in mllms’ multi-image reasoning.arXiv preprint arXiv:2601.14127, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.742784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.742784Z digest=sha256:09e75f96efe2f76f3970ca719a95dff5f1b71adb46d0558629c8cc79fec1105c

Observation 9550f4db-fcb5-44ca-aa70-6971f60b464a · outbound

This paper cites Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.880726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.880726Z digest=sha256:4315eb880a6038458115f45e280b0d0f75cdf451e23f301033745085ad8166e1

Observation 0e4b58c6-1b82-4790-b85a-10630cd29029 · outbound

This paper cites Jailbreaking llms & vlms: Mechanisms, evaluation, and unified defense.arXiv preprint arXiv:2601.03594, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Jailbreaking llms & vlms: Mechanisms, evaluation, and unified defense.arXiv preprint arXiv:2601.03594, 2026

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:13.955654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:13.955654Z digest=sha256:4438b5e3ff42b2af6bb3b7dd790ea2ed6eac13c32e16491c938db4f84fe5ed0e

Observation 756d73ff-2ec1-4636-81fd-ecb371ac8b24 · outbound

This paper cites Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Courtroom-Style Multi-Agent Debate with Progressive RAG and Role-Switching for Controversial Claim Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.030820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.030820Z digest=sha256:0a5abab10aea439227e2104606f610a6e133b51ceee59213f62db23edc05693f

Observation 482ba05c-36e7-4bbd-96b1-5fffbe61dc7b · outbound

This paper cites Improv- ing factuality and reasoning in language models through multiagent debate.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Improv- ing factuality and reasoning in language models through multiagent debate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.131820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.131820Z digest=sha256:77ebdb93ff0e7b707fcda96d30458259087ff3a11dc41c99ebb41ae1be2c57f0

Observation 2d181dc7-b008-4baf-9ebb-65978dae824b · outbound

This paper cites Bad students make great teachers: Active learning accelerates large-scale visual understanding.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Bad students make great teachers: Active learning accelerates large-scale visual understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.201724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.201724Z digest=sha256:236c048fccf3f2ed35b7ade255160d6d995f55c1cd2eeec817d004b2779596cc

Observation 15ff5940-23fc-4a59-830b-8f83a5e745a9 · outbound

This paper cites Contextnav: Towards agentic multimodal in-context learning.arXiv preprint arXiv:2510.04560, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Contextnav: Towards agentic multimodal in-context learning.arXiv preprint arXiv:2510.04560, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.292307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.292307Z digest=sha256:2523c7db5bcad5691de49a1e0ee14d08bd6ace08170cedc3af0cfcec60cfee5a

Observation 41636004-8125-44d0-958a-94928a86f81f · outbound

This paper cites Adversarial defense in vision-language models: An overview.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Adversarial defense in vision-language models: An overview

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.406045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.406045Z digest=sha256:86cdb1e1dbdbc877ce300c14fdc2b63e8319ab123931f1d6a175453f59849d5a

Observation 64785b11-9b35-4244-b252-0f932846d848 · outbound

This paper cites DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation DetPO: In-Context Learning with Multi-Modal LLMs for Few-Shot Object Detection

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.501395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.501395Z digest=sha256:9223a56202124f524597c3eb6aa2007247f228d8cca2429ad1e735fb4fd1dc75

Observation dd6e3a12-bcfc-4e2a-bc28-fdd9f16717ec · outbound

This paper cites Debate, deliberate, decide (d3): A cost-aware adversarial framework for reliable and interpretable llm evaluation.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Debate, deliberate, decide (d3): A cost-aware adversarial framework for reliable and interpretable llm evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.602511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.602511Z digest=sha256:314b2d267892080a83399099acdbb9a8edad30de08612873cf96bfaca2f66122

Observation ca22c62f-9c53-4427-ae90-2a3130017d0f · outbound

This paper cites Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration.arXiv preprint arXiv:2512.02530, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Aetheria: A multimodal interpretable content safety framework based on multi-agent debate and collaboration.arXiv preprint arXiv:2512.02530, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.693049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.693049Z digest=sha256:e4c1468d22e24b752793e398356f1d93a1b92abb52c518a68cd433a218481f23

Observation 13e9c64d-93e5-40cb-8bae-88ef7353e82a · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.769734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.769734Z digest=sha256:7f11afd3e7d7b6150c3c870286956da7c017a60f98fc8771760145dfddc2c0b1

Observation d069afdb-8d62-4b5b-bf68-da6366714a31 · outbound

This paper cites A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation A Judge-free LLM Open-ended Generation Benchmark Based on the Distributional Hypothesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.843081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.843081Z digest=sha256:f8ab4f33b0ce5b2ceeee76ec8b31dc971892d1d373c8cf75831d5167587bdc33

Observation a35ee8e3-d5dd-4d57-99ce-0c7df1dfd0e4 · outbound

This paper cites AI safety via debate.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation AI safety via debate

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.904515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.904515Z digest=sha256:78270e13fa6514a847c6252cffbd3ee4bfe3a8dd0312401c26bb3b0599056e92

Observation 91538c90-cdc5-4d1d-be9d-95236c6607c7 · outbound

This paper cites Adversarial attacks on multimodal large language models: A comprehensive survey.arXiv preprint arXiv:2603.27918, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Adversarial attacks on multimodal large language models: A comprehensive survey.arXiv preprint arXiv:2603.27918, 2026

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:14.990323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:14.990323Z digest=sha256:6620a9c5bb9c2927b758b0b54f97cba831d0280cc43c6ad10286785637f1307f

Observation 83e048fd-179d-4526-9624-52b5e4a6387b · outbound

This paper cites Curriculum guided massive multi agent system solving for robust long horizon tasks.arXiv preprint arXiv:2512.08545, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Curriculum guided massive multi agent system solving for robust long horizon tasks.arXiv preprint arXiv:2512.08545, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.074317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.074317Z digest=sha256:cbbd5fe9a6b9fdc71d79c08996be6affa73cdaf99f1fe3e5f0068c3d42af24dd

Observation ddf61226-84f2-4a8b-98c8-c4d3febcff2b · outbound

This paper cites Evaluating nova 2.0 lite model under amazon’s frontier model safety framework.arXiv preprint arXiv:2601.19134, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evaluating nova 2.0 lite model under amazon’s frontier model safety framework.arXiv preprint arXiv:2601.19134, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.141754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.141754Z digest=sha256:2f4e83662a827c85a37b5f94a47f702e030524eef3941c89de2a4a3d1d5aa10d

Observation a5eae34c-4ea6-4f95-971b-48f20d8970f9 · outbound

This paper cites Biasscope: Towards automated detection of bias in llm-as-a-judge evaluation.arXiv preprint arXiv:2602.09383, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Biasscope: Towards automated detection of bias in llm-as-a-judge evaluation.arXiv preprint arXiv:2602.09383, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.193060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.193060Z digest=sha256:7fa48a0e4484dcfc6248b88a7dfce8ed1da0230b4b38e6a1e47a5f48d5fe48e0

Observation 30ae7cc6-cf76-4203-828a-0d13922aa6b6 · outbound

This paper cites T-map: Red-teaming llm agents with trajectory-aware evolutionary search.arXiv preprint arXiv:2603.22341, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation T-map: Red-teaming llm agents with trajectory-aware evolutionary search.arXiv preprint arXiv:2603.22341, 2026

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.275900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.275900Z digest=sha256:e4928e8b70d28db807fe9550717ea2848ecf299536b1b81eea97bbc1a60afecb

Observation 9137ca88-9e26-4c79-86cf-143d26060fda · outbound

This paper cites THINKSAFE: Self-Generated Safety Alignment for Reasoning Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation THINKSAFE: Self-Generated Safety Alignment for Reasoning Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.360226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.360226Z digest=sha256:2e3870f705c6e9104e20a49277a302a5972be393d615199e5d97c2d6b37a097b

Observation 923205e8-ed21-4f3d-8ddd-1697e24ec7a1 · outbound

This paper cites Holisafe: Holistic safety benchmarking and modeling for vision- language model.arXiv preprint arXiv:2506.04704, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Holisafe: Holistic safety benchmarking and modeling for vision- language model.arXiv preprint arXiv:2506.04704, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.415061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.415061Z digest=sha256:a8cc1713467cb61da8097d881172d689a3348823cabee904493f56ee5cf6a974

Observation 07572d23-b8af-4145-a2de-79170a537b03 · outbound

This paper cites an unresolved cited work.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.502229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.502229Z digest=sha256:596d39d33b9d7c18ca4198358a5b26ba7b96abdb10f57ef4212859ad48958cc1

Observation 9594631e-637b-42b6-9e7e-8911513f7531 · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation From generation to judg- ment: Opportunities and challenges of llm-as-a-judge

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.559113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.559113Z digest=sha256:1f01fc7314ffe2d60af1f274cc1c458e597c1dcbf018af2a9cf3824427e3cd07

Observation f808fac0-18f9-4a75-be1a-ba4060ee7ead · outbound

This paper cites Evaluating Scoring Bias in LLM-as-a-Judge.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evaluating Scoring Bias in LLM-as-a-Judge

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.637387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.637387Z digest=sha256:0c094674dcd7badb373638f41ec943448b2b18d969b14527793c0f51d9b84b46

Observation 9d216852-3fe2-4f35-8201-37ac60540127 · outbound

This paper cites Who judges the judge? llm jury-on-demand: Building trustworthy llm evaluation systems.arXiv preprint arXiv:2512.01786, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Who judges the judge? llm jury-on-demand: Building trustworthy llm evaluation systems.arXiv preprint arXiv:2512.01786, 2025

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.711786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.711786Z digest=sha256:d01f14520b6e48df40a862c900766e6dd8eede15e0903028d3ccf7ed2031b02c

Observation 67fad8d2-9f74-4a5c-82a7-14a406e4c347 · outbound

This paper cites Benchmark test-time scaling of general llm agents.arXiv preprint arXiv:2602.18998, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Benchmark test-time scaling of general llm agents.arXiv preprint arXiv:2602.18998, 2026

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.789634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.789634Z digest=sha256:9fadc7ca917c75c21527ae9a611fa7f3a013e593507c481bedc8dff099a46c9b

Observation 1df74600-1e78-46b9-8622-439d85ef6787 · outbound

This paper cites Elhplan: Efficient long-horizon task planning for multi-agent collaboration.arXiv preprint arXiv:2509.24230, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Elhplan: Efficient long-horizon task planning for multi-agent collaboration.arXiv preprint arXiv:2509.24230, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.871821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.871821Z digest=sha256:98375da1c1257aa97634ffe7046a1af9e2d6071133595e99fb595a39b6b9b7bf

Observation 8ab350e0-6b97-41c7-ab24-217c230c7269 · outbound

This paper cites Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Examining LLMs' Uncertainty Expression Towards Questions Outside Parametric Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:15.960327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:15.960327Z digest=sha256:668ababd78e981d34977b0d1ee3bf732a1aed4dda19450dcf4b9fde91fa53197

Observation dd68d04f-a297-440d-af26-260c3a2cfa9d · outbound

This paper cites WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation WebCoach: Self-Evolving Web Agents with Cross-Session Memory Guidance

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.036683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.036683Z digest=sha256:79314472055edcecb2c6fccc5d698f180e774b2b24f76df79355a1d3dfa0d6c3

Observation 7ce210d1-3829-4b1c-812a-83df63e4f122 · outbound

This paper cites Mosaic: Modeling social ai for content dissemination and regulation in multi-agent simulations.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Mosaic: Modeling social ai for content dissemination and regulation in multi-agent simulations

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.085510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.085510Z digest=sha256:6dc467db08089a5f23e39b07578e17a6ba9bf86ed6e462757b9d82064aefb6e8

Observation 5ab79f7c-a3bc-497f-8223-45b369e101bb · outbound

This paper cites Mtmcs-bench: Evaluating contextual safety of multimodal large language models in multi-turn dialogues.arXiv preprint arXiv:2601.06757, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Mtmcs-bench: Evaluating contextual safety of multimodal large language models in multi-turn dialogues.arXiv preprint arXiv:2601.06757, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.135485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.135485Z digest=sha256:e5564ad26834d45d1f9bcac676fec325c5a5efe5196f2495f7c4f6013b94d591

Observation 531fa067-92e2-407c-a07f-12b8772125dc · outbound

This paper cites Ai debate aids assessment of controversial claims.arXiv preprint arXiv:2506.02175, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Ai debate aids assessment of controversial claims.arXiv preprint arXiv:2506.02175, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.185073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.185073Z digest=sha256:4025197cae692c300f4d8ffd2017d22550157027e11776387b6e38b8b3f217bb

Observation 9236d518-02f5-4c78-86c7-a3e914e97b47 · outbound

This paper cites X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.255412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.255412Z digest=sha256:4004e9ca9462ee6cac3eba4b057f0630a06fda01062089feeab99f4d226fffa6

Observation ffb7c359-7219-48e5-bcf7-2e1325c7a3cd · outbound

This paper cites Disc-amc: Token-and parameter-efficient discretized statistics in-context automatic modulation classification.arXiv preprint arXiv:2510.00316, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Disc-amc: Token-and parameter-efficient discretized statistics in-context automatic modulation classification.arXiv preprint arXiv:2510.00316, 2025

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.307323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.307323Z digest=sha256:b32ece57e1e7e7f0d354e07985caebacecd1f81b2c83ae686fd3cb2a13cd87cf

Observation 7058828b-0ee1-4b46-a2c9-16c53fc97cb7 · outbound

This paper cites Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Interaction Protocol Shapes Moral Judgment in Multi-Agent Debate

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.371761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.371761Z digest=sha256:9a5981e9c2a808f7617775b11704ec2679d876f5863b9b32a569b56920eaee38

Observation 95cefedf-3041-4473-b764-6c41c76f0066 · outbound

This paper cites Assessment of Multimodal Large Language Models in Alignment with Human Values.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Assessment of Multimodal Large Language Models in Alignment with Human Values

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.420847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.420847Z digest=sha256:3e29f6ee5a3762dc68271f2817d9fdae12494d3641dab7a7157870ecbe718b1f

Observation b3ab628d-ad3d-427c-a6a4-17e7ae4e68e0 · outbound

This paper cites Llm-as-a-judge for time series explanations.arXiv preprint arXiv:2604.02118, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Llm-as-a-judge for time series explanations.arXiv preprint arXiv:2604.02118, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.446615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.446615Z digest=sha256:808810b9cd388f415a26859bf0eedbfe3680de9c13234c804c45d056f204f287

Observation 120bdebe-f285-4bb1-bb70-5a21c75d58a0 · outbound

This paper cites Unigame: Turning a unified multimodal model into its own adversary.arXiv preprint arXiv:2511.19413, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Unigame: Turning a unified multimodal model into its own adversary.arXiv preprint arXiv:2511.19413, 2025

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.513676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.513676Z digest=sha256:44706f0b578d8f1e965aef00ae4515b4c7b2b8e9fb391f3b0114017b34b16377

Observation b5435b84-4cfe-4289-a100-9e21c2b2aba3 · outbound

This paper cites Supporting human raters with the detection of harmful content using large language models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Supporting human raters with the detection of harmful content using large language models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.574250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.574250Z digest=sha256:d4c8198b720e98db901955e2603217e82b5a82d4c914894731f6e487dea7cf98

Observation 26823bb7-34c3-4c03-b006-82a95b9ab134 · outbound

This paper cites Automated concept discovery for llm-as-a-judge preference analysis.arXiv preprint arXiv:2603.03319, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Automated concept discovery for llm-as-a-judge preference analysis.arXiv preprint arXiv:2603.03319, 2026

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.631763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.631763Z digest=sha256:ed14d0b19d1bae2dd8667e506c451788b7dced4ec7c125cb8501c3f76f90c71a

Observation ae09ccf1-033f-4661-b6d9-4cdeb3e18042 · outbound

This paper cites Can llm agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Can llm agents really debate? a controlled study of multi-agent debate in logical reasoning.arXiv preprint arXiv:2511.07784, 2025

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.691039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.691039Z digest=sha256:e63130f886c10b3870172b44d4ea5cd015a942ceeb8096315baa5864cfcc815f

Observation 9921a8aa-53d0-4ed6-8c6d-f66a804b1077 · outbound

This paper cites OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation OutSafe-Bench: A Benchmark for Multimodal Offensive Content Detection in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.755987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.755987Z digest=sha256:2ec59d96c18d1db38847f8eefaf3b992dc4c63fa1636164cbd661256016b6928

Observation 808993bb-8fe5-42ae-b424-758395746a7d · outbound

This paper cites Llama-3.1-foundationai- securityllm-reasoning-8b technical report.arXiv preprint arXiv:2601.21051, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Llama-3.1-foundationai- securityllm-reasoning-8b technical report.arXiv preprint arXiv:2601.21051, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.818828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.818828Z digest=sha256:308750a9fd265c2b2132a842061ff9abbc1862bb16c1f68a270acf7b2c4e1d16

Observation 243d2632-a2f4-454d-a540-90894987e1cb · outbound

This paper cites When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.888129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.888129Z digest=sha256:cb6b51d9af6b1cf701d7baa9f36c8b3701183778f300cd8ab5977b0f09aaf786

Observation d7e7d1ef-c352-43da-8b8c-b317f6c69775 · outbound

This paper cites Evolving contextual safety in multi-modal large language models via inference-time self-reflective memory.arXiv preprint arXiv:2603.15800, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Evolving contextual safety in multi-modal large language models via inference-time self-reflective memory.arXiv preprint arXiv:2603.15800, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:16.956146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:16.956146Z digest=sha256:548a41ace788700336cb9fadddc70d44921c55208cab7ac13b20bebdb8414567

Observation 91a72551-1176-4d1c-9b5d-d29132eb03b7 · outbound

This paper cites Visual exclusivity attacks: Automatic multimodal red teaming via agentic planning.arXiv preprint arXiv:2603.20198, 2026.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Visual exclusivity attacks: Automatic multimodal red teaming via agentic planning.arXiv preprint arXiv:2603.20198, 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.016819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.016819Z digest=sha256:98ad0399a2ca0218d8c8dd25072178e58d3f58349fa674c7b1f927e0c9be9a69

Observation b2f95ed4-09d2-4d0e-a9f8-f48b33a53814 · outbound

This paper cites Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation Amulet: ReAlignment During Test Time for Personalized Preference Adaptation of LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.054063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.054063Z digest=sha256:99c13fdc3435e593653c9d03c78b67705f9a6b76529389204e629c105c4a4d72

Observation bc460a0f-f52e-4af6-8b41-e9d10122f13f · outbound

This paper cites inhumane conditions.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation inhumane conditions

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.106150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.106150Z digest=sha256:10c8478fc0d8c361438e39078e62913175d0606ecf162a299409f08d502d5e51

Observation 24c00459-75db-4a1e-8d9b-515db75a8ca2 · outbound

This paper cites {policy_text}.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation {policy_text}

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.183685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.183685Z digest=sha256:7d392873d39b640060124a53f9ae5912f4c2988ce9e1b825aef44fda6a9877a1

Observation 4aaef1b5-e16b-45b9-b010-037a7d8986f6 · outbound

This paper cites This synthesizes a complete context block explicitly documenting the distinct arguments for and against each candidate classification label.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation This synthesizes a complete context block explicitly documenting the distinct arguments for and against each candidate classification label

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.271516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.271516Z digest=sha256:a5f37b7a3117c1eab9ecb6d72a0fe8d88bded8f9f31853c95de2979c53610555

Observation 0af11b24-c32d-4db8-bc87-8fd0451384ff · outbound

This paper cites {policy_text}.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation {policy_text}

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.347594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.347594Z digest=sha256:5bcf5d334b83e60ce464113f8d29204cb948c90971dc93d0e996c65d5dda6156

Observation a4c2016a-a3d2-46bc-b8c6-7d70ce6315a9 · outbound

This paper cites If any dissent continues even after the debate round, the system automatically categorizes the instance as a deadlock and escalates it to the Level II Jury Committee.

Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation If any dissent continues even after the debate round, the system automatically categorizes the instance as a deadlock and escalates it to the Level II Jury Committee

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:17.432243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T02:44:17.432243Z digest=sha256:c97aba03235fa4551e4cf53998a336c23769f5b29a54e17edbdc6da2244fefff

Pith citing papers

No inbound Pith citation observations are available.