Pith. sign in

hub Mixed citations

OpenAI blog , volume=

Mixed citation behavior. Most common role is background (57%).

88 Pith papers citing it
Background 57% of classified citations

hub tools

citation-role summary

background 6 method 1

citation-polarity summary

claims ledger

  • background org/10.18653/v1/ 2025.emnlp-main.999 Portelance, E., Duan, Y., Frank, M. C., & Lupyan, G. (2023). Predicting age of acquisition for children's early vocabulary in five languages using language model surprisal.Cognitive Science,47(9), e13334. Portelance,E.,&Jasbi,M.(2024).Therolesofneuralnetworks inlanguageacquisition.LanguageandLinguisticsCompass, 18(6), e70001. https://doi.org/https://doi.org/10.1111/lnc3. 70001 Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Langu
  • background out(x 0, x1)pairs and construct flow states at 20 randomly sampled timesteps and the clean target (t=1.0), producing4 200total evaluations per source distribution. We computeEϕ(xt)and report binned mean energy, monotonicity, rank correlation, and aggregate metrics in tables 9 and 10. Table 9Time-energy monotonicity.Mean energyE ϕ(xt)per time bin; both source distributions. Time bin Src[0,.1) [.1,.2) [.2,.3) [.3,.4) [.4,.5) [.5,.6) [.6,.7) [.7,.8) [.8,.9) [.9,1)t=1 Uni. ¯Eϕ −0.12−0.57−0.75−0.90−1
  • method Let X be a random variable taking values in [−B, B]. Assume thatE  e−ηX ≤1, then E[X2]≤4 1 η +B  E[X]. Lemma 6(Freedman's inequality (Beygelzimer et al., 2011)).Let (Xt)t≤T be a real-valued mar- tingale difference sequence adapted to a filtration(Ft)t≤T . If |Xt| ≤R almost surely, then for any η∈(0,1/R), with probability at least1−δ, TX t=1 Xt ≤η TX t=1 Et−1  X2 t  + log(δ−1) η .(43) The following result is a consequence of Lemma 6. Lemma 7.Let (Xt)t≤T be a sequence of random variables ada
  • background Problem Formulation Let the prompt space be P ⊂ {V,V 2, . . .} where V is model vocabulary, the protected agent prompt be x∈ P and θ be model parameters.Our objectiveis to identify another prompt ˜x∈ P that satisfy the following criteria: (1)Obfuscation (C2): ˜x is distant from x in prompt space; (2)Usability (C3): ˜x is functionally equivalent to x; (3) Non-portability (C4): the utility objective is satisfied only on the target LLM, not other LLMs. 3.2. Theoretical Motivation In this subsection
  • background where ˆ𝜶 ℎ∈R𝑑×𝐶ℎ𝑖𝑑 , ˆ𝜸ℎ , ˆ𝜷 ℎ∈R𝑑𝜙×𝐶ℎ𝑖𝑑 are learnable parameters, and𝐶ℎ𝑖𝑑 is the hidden dimensional- ity of the cross-attention. This allows the model to attend to specific medical history with causality information in the latentˆywhen synthesizing the DT. The objective is to predict the noise𝜖added to the latent representation𝑧𝑡 ℒLDM =E 𝑧∼ℰ(M),𝑡,𝜖,𝑦  ∥𝜖−𝜖𝜃(𝑧𝑡 , 𝑡,ˆy)∥2 (10) where 𝜖𝜃 is the MLP U-Net. During inference, we sample a latent𝑧0 conditioned on the patient's history and decode it us
  • background Ego Network Model (b) Interaction Record from A to B Support Clique … … Like Comment Retweet 40% 40% 10% (a) Tie Strength CoT Training Dataset A CoT Construction User A User B (a) Multi-hop Following Relationship Sampling (b) Graph-to-(sequential) Text Encoding User B User A User C User D User E User F User A User B User B User C … User E User F [Hop 2] …… [Hop 5] ……… CoT Construction Module 1 Module 2 Step 1: Community initialization Step 2: Follow relationship generation Step 3: Interaction Be

co-cited works

representative citing papers

Beyond Heuristic Tuning: Power-Calibrated LLM Watermarking

stat.ML · 2026-07-06 · accept · novelty 7.0

A power-calibrated statistical framework gives closed-form links from KGW watermark parameters (γ, δ) to detection power and KL distortion, enabling principled Pareto-optimal selection.

Probabilistic Attribution For Large Language Models

cs.CL · 2026-05-20 · unverdicted · novelty 7.0

Develops a model-agnostic attribution score as the log-ratio of conditional response probabilities with and without a marginalized prompt token, derived via Bayes inversion of next-token distributions, and relates it to conditional entropies.

When Does Model Collapse Occur in Structured Interactive Learning?

cs.LG · 2026-05-19 · unverdicted · novelty 7.0

Model collapse occurs in structured interactive learning if and only if the directed interaction graph satisfies a specific topological condition, with finite-sample guarantees for linear regression and asymptotic results for M-estimators.

Self-Improvement for Fast, High-Quality Plan Generation

cs.AI · 2026-05-05 · unverdicted · novelty 7.0

Self-improvement of a decoder-only transformer yields plans averaging 30% shorter than a source symbolic planner, over 80% optimal where known, with sub-exponential latency scaling.

Self-Rewarding Language Models

cs.CL · 2024-01-18 · conditional · novelty 7.0

Iterative self-rewarding via LLM-as-Judge in DPO training on Llama 2 70B improves instruction following and self-evaluation, outperforming GPT-4 on AlpacaEval 2.0.

citing papers explorer

Showing 50 of 88 citing papers.