Pith. sign in

Paper Citation Record · LEDGER

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

As of 22 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 4 inbound Pith citation observations for arXiv:2501.00418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00418 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:17.294000Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:23:15.199636Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5c8e58f1-382a-48e8-a37b-39e514e23e41 · outbound

This paper cites Deep learning with differential privacy.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Deep learning with differential privacy

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.095613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.095613Z digest=sha256:89c4def2501fcd7e2e891ff1ced17a48dea4c5f0edcdb94c2f62c49dbab97007

Observation 4cf6d8b4-e7da-4200-af45-8fc4392de94a · outbound

This paper cites Types of Out-of-Distribution Texts and How to Detect Them.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Types of Out-of-Distribution Texts and How to Detect Them

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.100416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.100416Z digest=sha256:4ed1fafe5b64d4d179332148571f186cfe592e6d1bb4ef9243168f02a9251f52

Observation ba0e8d9c-0793-4571-84e6-d83b7130f7fd · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.829745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.105045Z digest=sha256:4d829dfacc92e7664f9b3f10e650d1da94058546bf3017ea4c29aa2ec1043795

Observation 5db256fe-b703-400a-97f5-c90c9f1b2276 · outbound

This paper cites Man is to computer programmer as woman is to homemaker? debiasing word embeddings.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Man is to computer programmer as woman is to homemaker? debiasing word embeddings

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.110022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.110022Z digest=sha256:10f250ef57df5c8134a6c14dc4ac34fea1a11e7fd8f2981939e6738f17908238

Observation c0e61272-cdf4-4e70-823d-b3b53a899bac · outbound

This paper cites Language models are realistic tabular data generators.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Language models are realistic tabular data generators

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.806767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.114332Z digest=sha256:9d371d6e92f28596e2552a3365e1b83bf9595e868edfd2384d250adc4d39a225

Observation 598298da-00ab-43c7-89f5-e8b19dc89534 · outbound

This paper cites Generating Sentences from a Continuous Space.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Generating Sentences from a Continuous Space

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.119094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.119094Z digest=sha256:d8815bd67628b17bda766c3aac7aefdd60068cf9adb6b413b62076e51e385e2c

Observation 1bd0c2b0-c174-433c-9f08-5cd6a0801e6e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.123720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.123720Z digest=sha256:93d1a0896366a7d28a34b350ccf5993ae89789b4fe5fc7fcade46a9b252b2117

Observation 5c029ba4-6d01-4cfe-abc5-703bb9c29312 · outbound

This paper cites Weak-to-strong generalization: Eliciting strong capabilities with weak supervision.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Weak-to-strong generalization: Eliciting strong capabilities with weak supervision

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.792252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.127839Z digest=sha256:818b4b5e150070ace519400e721309e932ae573f5c2ac79ccbd2bd88f28441ad

Observation 1307bddd-955d-4ba4-8c00-544026da2576 · outbound

This paper cites Extracting training data from large language models.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Extracting training data from large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.132106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.132106Z digest=sha256:bf6ba8664c538ca7a4c1e74cde2e7c60317a1323a3fa576948455c8decf8ac02

Observation 8e9508fb-2ae6-4a58-b8c0-83135f298c20 · outbound

This paper cites Retiring adult: New datasets for fair machine learning.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Retiring adult: New datasets for fair machine learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.764110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.136975Z digest=sha256:aa1c7dee12cf3460daf3637f5760b9ad054d9cd27bd5155a22eabbec15f0288c

Observation b9f4d537-f9de-4555-95f8-77fe4e3e8bb7 · outbound

This paper cites Calibrating noise to sensitivity in private data analysis.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Calibrating noise to sensitivity in private data analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.141214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.141214Z digest=sha256:b0252041833ffb49b2701eb99ac3fd3f64e018db2f432745bcf5f62e6c6c4f0b

Observation 654ac7ef-24dd-4ab6-924a-297dc1ea55fa · outbound

This paper cites BAE: BERT-based adversarial examples for text classification.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models BAE: BERT-based adversarial examples for text classification

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.147521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.147521Z digest=sha256:e34e3d38db807b2fe0f2346012b3936503fd31ba052c50ee477652c6be85e2dc

Observation d1b1219e-6e47-4f6c-a5af-5027b6fe9d9b · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Explaining and Harnessing Adversarial Examples

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.152544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.152544Z digest=sha256:67780d54f5221f3c3ef77344246ca2223887c18b6e0d078c2e46b4ae065ae317

Observation dbf4754a-ae37-4c36-be28-9774298c1750 · outbound

This paper cites Reducing sentiment bias in language models via counterfactual evaluation.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Reducing sentiment bias in language models via counterfactual evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.157126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.157126Z digest=sha256:53c4bea2adeddbcc79f88658e05bf7fca40f70a0ebc5b968013467e000f580e1

Observation 944c1e52-49d2-41d5-9656-02dc8312debf · outbound

This paper cites Students parrot their teachers: Membership inference on model distillation.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Students parrot their teachers: Membership inference on model distillation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.740501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.161816Z digest=sha256:34d509dc23b34af9373e9fd8d000f98754e0f13a797024494e80454e61812c37

Observation 2e93573f-77c8-4ee2-9961-923a8aea1940 · outbound

This paper cites Is BERT really robust? A strong baseline for natural language attack on text classification and entailment.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Is BERT really robust? A strong baseline for natural language attack on text classification and entailment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.166633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.166633Z digest=sha256:4289c56970000155a17cd677d145df77b3548342f5faf3c486e91f38ed012d8c

Observation 39542e87-0d8b-4bb7-a84f-082094c8201c · outbound

This paper cites The enron corpus: A new dataset for email classification research.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models The enron corpus: A new dataset for email classification research

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.720294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.172051Z digest=sha256:5b624d215c5149c8c85eed002e75dbecf776e1a485c1a664aaa9cd2896656ece

Observation 9f46bc3b-ce1d-4d47-b89d-14a4c7096c61 · outbound

This paper cites Certified robustness to adversarial examples with differential privacy.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Certified robustness to adversarial examples with differential privacy

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.706453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.177831Z digest=sha256:193aae44c03cd18b32f0813a31b109a5e009c055d48e69f4f04b74b0c492835b

Observation 6a57a0d9-d202-4e51-b7fb-0dd768117b05 · outbound

This paper cites Gaussian membership inference privacy.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Gaussian membership inference privacy

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.693337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.183166Z digest=sha256:b2a7ee6170ec54767594efa3af3dda6efdcac554e7da00cb151d6124dc558659

Observation cf75afa3-cb55-4c85-9c1b-76e55e8cc6c3 · outbound

This paper cites Certified adversarial robustness with additive noise.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Certified adversarial robustness with additive noise

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.678935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.187765Z digest=sha256:5c6e88509cd33a8dd8414a16992374d330a81bea0349d57f91d5c9d40ddf3f6f

Observation 9296abc1-78f1-40a0-9663-b3b4e0e407e6 · outbound

This paper cites BERT-ATTACK: Adversarial attack against BERT using BERT.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models BERT-ATTACK: Adversarial attack against BERT using BERT

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.192121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.192121Z digest=sha256:cf5ce9197e3df94151ca50853f4457e8e2e8e882ff556ae861b79fe8986450e2

Observation e74d4151-545e-451a-bf51-b6dc2f8347b9 · outbound

This paper cites Focal Loss for Dense Object Detection.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Focal Loss for Dense Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.197196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.197196Z digest=sha256:ef114785aff71ba6e195628920d59fae582ca9f55edbb4efa542531d64400f37

Observation cbadb186-30f5-4f49-8d96-d70ae8a0bd3e · outbound

This paper cites Towards Deep Learning Models Resistant to Adversarial Attacks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Towards Deep Learning Models Resistant to Adversarial Attacks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.201696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.201696Z digest=sha256:142fdc0b6ea0d05c2dfb7cbebdc18f30fbfcba229fa06707205bc7b4dbf0025f

Observation 1d5fbc65-ecf1-47a6-88d4-3ec3c324e019 · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Towards deep learning models resistant to adversarial attacks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.666263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.205854Z digest=sha256:ca54839878060e72a90837cc8ea6490b36c21f7ff7f7ea22686ca8b98951b884

Observation 6cebab5c-3026-47af-936f-899534ed54f0 · outbound

This paper cites Repeated knowledge distillation with confidence masking to mitigate membership inference attacks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Repeated knowledge distillation with confidence masking to mitigate membership inference attacks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.652549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.209968Z digest=sha256:07914c9268601ee21853fbd8f148bf6775ea69020b4a7631eb9b4e7a1d6a055a

Observation cd4adfea-f16a-46a0-a719-993bc5f431f0 · outbound

This paper cites Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Semi-supervised Knowledge Transfer for Deep Learning from Private Training Data

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.214013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.214013Z digest=sha256:de9d80c491e3f3704d4045847ff63f22fb3cc5270cc7135f4a76b8a4c1b10e33

Observation 1ff2eb1f-94d6-4b69-a25e-747a0aca3286 · outbound

This paper cites Language models are unsupervised multitask learners.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Language models are unsupervised multitask learners

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.219018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.219018Z digest=sha256:a9e02a54efe112edaa0906a20970c1933ef797dfd07b62418400f92d1ca8cbd3

Observation 2d320f86-b5d6-474a-abff-b480e52be49d · outbound

This paper cites Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2024.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Are emergent abilities of large language models a mirage? Advances in Neural Information Processing Systems, 36, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.223166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.223166Z digest=sha256:cfcb3ffba3c6caf0a8023722bf6131a153b1b0e6096d1bb3531c1b647dab7e5d

Observation fde160cf-5a22-4a04-b497-a4af7dffca52 · outbound

This paper cites Membership privacy for machine learning models through knowledge transfer.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Membership privacy for machine learning models through knowledge transfer

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.625008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.227063Z digest=sha256:9e088a3b2da027083c92d2a9b5f496fa49d22e7db668ade2293866e18f71f894

Observation 136abd5a-e8d0-448d-9c63-0776758c3d29 · outbound

This paper cites Intriguing properties of neural networks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Intriguing properties of neural networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.231621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.231621Z digest=sha256:8cd023a71af4017d9455154c06d85d52c4ddf590c12f766e945974ed8312211a

Observation 4b50bbef-4638-454a-95e1-9126f9b9cd41 · outbound

This paper cites Rethinking the inception architecture for computer vision.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Rethinking the inception architecture for computer vision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.235789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.235789Z digest=sha256:f9cf7c5b3ce07336033666f49c610bc3f10d0075ea38edad455a8b65b81041d7

Observation 56a17418-ccb2-4e3e-9145-eb28230bb3fa · outbound

This paper cites Mitigating membership inference attacks by {Self-Distillation} through a novel ensemble architecture.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Mitigating membership inference attacks by {Self-Distillation} through a novel ensemble architecture

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.600289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.239929Z digest=sha256:021bba1a1a6ecebeae571a7ac6b2852a9a6bb9a9566108e8044a86a382e2ced2

Observation 79c8cc56-cfa2-4566-87af-1c26d0599b26 · outbound

This paper cites T3: Tree-autoencoder constrained adversarial text generation for targeted attack.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models T3: Tree-autoencoder constrained adversarial text generation for targeted attack

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.244405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.244405Z digest=sha256:c6398adc661eb8a3189069afd74150b7379295658637f68c715e1b53718ce208

Observation 9b50b6b4-6224-4985-8d20-3b03d2c13926 · outbound

This paper cites Adversarial GLUE: A multi-task benchmark for robustness evaluation of language mod- els.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Adversarial GLUE: A multi-task benchmark for robustness evaluation of language mod- els

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.570680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.253890Z digest=sha256:233b674780708f4ab3e8282bf49505108cade83ed93f311849ae35c91d7ed098

Observation 95f7a2a4-031b-42ae-a276-03d4b2bf2b54 · outbound

This paper cites Decodingtrust: A comprehensive assessment of trustworthiness in gpt models.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Decodingtrust: A comprehensive assessment of trustworthiness in gpt models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.556048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.258921Z digest=sha256:b4c5d95dce6cfbf83a2ce66cbdd80efc8085180d84923e6d8efdfdba583ef5c0

Observation b527d9bf-42fb-4362-b7fc-48c93c33d6df · outbound

This paper cites EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.264471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.264471Z digest=sha256:b17471cdbc5a8e12eea7b09c3bd5960f2fe96f9748d0634a40ca8bee5be74edc

Observation 4b9fdc43-8667-428d-b59f-596aafea6ff5 · outbound

This paper cites Emergent Abilities of Large Language Models.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Emergent Abilities of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.268642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.268642Z digest=sha256:39aa9f1b51ab0f9159b53afd96586267a873f8da02cfeb074fb4a2d4d2d2d109

Observation 571b1777-95d2-47a5-81e0-8cf3004c8391 · outbound

This paper cites Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.542888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.272777Z digest=sha256:e55ed3b5ababa8646b659cbfc67b9cc0c7b56395bbcfd0676cd1bf3a49a7e6e4

Observation ca347c03-fe7f-45e7-8344-5967c4a4ee8c · outbound

This paper cites Fairness constraints: Mechanisms for fair classification.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Fairness constraints: Mechanisms for fair classification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.528494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.276628Z digest=sha256:0d68ac8a76326790a95f1efeab132ae0e47b8fa0e1a8769fce7313f79f876b3c

Observation 7d72000f-3ebc-449d-83d3-2bcffab1ab4c · outbound

This paper cites Word-level textual adversarial attacking as combinatorial optimization.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Word-level textual adversarial attacking as combinatorial optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.280424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.280424Z digest=sha256:802f46d6b8c6383434e86b8c6bf5ef9447848f8a293668896956c300f2db1d65

Observation eae24b9f-29e0-4a75-a1bc-93bc9519a2fb · outbound

This paper cites Gender bias in coreference resolution: Evaluation and debiasing methods.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Gender bias in coreference resolution: Evaluation and debiasing methods

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.284964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.284964Z digest=sha256:b4a4ceb5833ae7029be37612ea5438b77d04478a0534265c3262c80cf5b26a4f

Observation c77db318-aa73-46a4-944e-4f3fd6db689e · outbound

This paper cites Resisting membership inference attacks through knowledge distillation.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Resisting membership inference attacks through knowledge distillation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:56:17.514659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.289433Z digest=sha256:79404ef03fb082cc6c0540b50e4050000cd463eebad62e75c452fd1c3b42f8cc

Observation 3bc2b9de-59b9-47d6-b537-9db6bc2b7ee4 · outbound

This paper cites FreeLB: Enhanced Adversarial Training for Natural Language Understanding.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models FreeLB: Enhanced Adversarial Training for Natural Language Understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:17.294000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:17.294000Z digest=sha256:9d36377a4ac171f1675da4de59e9340ccddace5f50f6555f3f5ac59057292f28

Observation b4588d93-7f23-4a3b-8257-e29287254382 · outbound

This paper cites an unresolved cited work.

Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models Unresolved cited work

Reference 495

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:56:17.585753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:56:17.248663Z digest=sha256:abe45834d6ea41f6d856a4e76857c93d45ad403fd0d0ab586ec84e14ee374b22

Pith citing papers

Observation b6afeab1-5c4c-4478-b064-1b062820a5b8 · inbound

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration cites this paper.

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T15:23:15.199636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:23:15.199636Z digest=sha256:e14eecb1e57c526e4786e65e5c3e90d5ef2caa28b06711e760d0a51b09f8da37

Observation 6272a041-0594-41f9-9313-7adfab482516 · inbound

On Weak-to-Strong Generalization and f-Divergence cites this paper.

On Weak-to-Strong Generalization and f-Divergence Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:46.722521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:17:46.722521Z digest=sha256:dc922b1799a585e9936276386d3a9d9e5c861fc50f51f2312385c5417c571dbc

Observation 3ba96de3-5d0a-4103-b537-81e561435348 · inbound

On the Blessing of Pre-training in Weak-to-Strong Generalization cites this paper.

On the Blessing of Pre-training in Weak-to-Strong Generalization Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:08.625742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T14:59:19.883399Z digest=sha256:f46708d9cdf0b4bb5c54185182d4f49a816fa0bad2f7eb3164030a1a7cd59cca

Observation ddc2a02a-6551-4dd4-a954-018a6182b1af · inbound

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher cites this paper.

Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher Generalizing Trust: Weak-to-Strong Trustworthiness in Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:42:24.943936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-28T17:35:24.108984Z digest=sha256:0c3b06bb5107f27c1dc5e38d0c0ddb86bc03e7ea14ddf607ea4d715df1b2ebe7