{"work":{"id":"31dc92c3-2fc3-432b-8d63-b7ee13f53a9c","openalex_id":null,"doi":null,"arxiv_id":null,"raw_key":"raw:9ec4527a71eb2b7ff063a889","title":"OpenAI blog , volume=","authors":null,"authors_text":"Language models are unsupervised multitask learners , author=","year":null,"venue":null,"abstract":null,"external_url":null,"cited_by_count":null,"metadata_source":"raw_reference","metadata_fetched_at":"2026-07-11T03:57:49.142954+00:00","pith_arxiv_id":null,"created_at":"2026-05-10T23:33:01.612684+00:00","updated_at":"2026-07-11T03:57:49.142954+00:00","title_quality_ok":false,"display_title":"OpenAI blog , volume=","render_title":"OpenAI blog , volume="},"hub":{"state":{"work_id":"31dc92c3-2fc3-432b-8d63-b7ee13f53a9c","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":88,"external_cited_by_count":null,"distinct_field_count":15,"first_pith_cited_at":"2021-04-18T08:44:56+00:00","last_pith_cited_at":"2026-07-08T21:34:49+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T04:09:28.557684+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"background","n":6},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":4},{"context_polarity":"unclear","n":2},{"context_polarity":"use_method","n":1}],"runs":{"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-22T20:43:52.924177+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":26},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":21},{"title":"Advances in neural information processing systems , volume=","work_id":"12f5a236-ef7a-4d13-b4de-b51465a6f977","shared_citers":20},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":16},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":14},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":14},{"title":"OPT: Open Pre-trained Transformer Language Models","work_id":"d7ff3b21-1fff-4cf4-952a-4714e3ef2307","shared_citers":13},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":13},{"title":"Advances in neural information processing systems , volume=","work_id":"a1fd09f1-b62b-4aca-a5ef-dd2b50ad08b5","shared_citers":12},{"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","shared_citers":12},{"title":"Training Compute-Optimal Large Language Models","work_id":"b2faf28d-86b7-429c-bc42-469458efc246","shared_citers":12},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":11},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":11},{"title":"BLOOM: A 176B-Parameter Open-Access Multilingual Language Model","work_id":"337ba690-f35d-4154-9450-8edf4bc9f488","shared_citers":10},{"title":"On the Opportunities and Risks of Foundation Models","work_id":"a18039e9-928d-47c9-a836-32656a71bf71","shared_citers":10},{"title":"Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback","work_id":"a1f2574b-a899-4713-be60-c87ba332656c","shared_citers":10},{"title":"Advances in Neural Information Processing Systems , volume=","work_id":"bd577a47-49ef-4f8a-81ca-86cd88b71479","shared_citers":9},{"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","shared_citers":9},{"title":"Language models are unsupervised multitask learners","work_id":"175d0d95-85f4-480c-b6d3-53a7e52df1e4","shared_citers":9},{"title":"Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism","work_id":"c888e6d1-0b1d-43d6-9ef5-f0912a0efa1b","shared_citers":9},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":9},{"title":"SC20: International Conference for High Performance Computing, Networking, Storage and Analysis , pages=","work_id":"f62f435f-90be-45ec-bdea-343bdc2734f6","shared_citers":9},{"title":"2016 , publisher=","work_id":"cf0899e0-53ee-4591-aae4-f38fa5ac12ad","shared_citers":8},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":8}],"time_series":[{"n":1,"year":2021},{"n":2,"year":2022},{"n":13,"year":2023},{"n":7,"year":2024},{"n":1,"year":2025},{"n":42,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds","primary_cat":"cs.LG","context_text":"Let X be a random variable taking values in [−B, B]. Assume thatE \u0002 e−ηX\u0003 ≤1, then E[X2]≤4 \u00121 η +B \u0013 E[X]. Lemma 6(Freedman's inequality (Beygelzimer et al., 2011)).Let (Xt)t≤T be a real-valued mar- tingale difference sequence adapted to a filtration(Ft)t≤T . If |Xt| ≤R almost surely, then for any η∈(0,1/R), with probability at least1−δ, TX t=1 Xt ≤η TX t=1 Et−1 \u0002 X2 t \u0003 + log(δ−1) η .(43) The following result is a consequence of Lemma 6. Lemma 7.Let (Xt)t≤T be a sequence of random variables adapted to a filtration (Ft)t≤T . If 0≤X t ≤Ralmost surely, then with probability at least1−δ, TX t=1 Xt ≤(1 +ε) TX t=1 Et−1[Xt] + R ε log(δ−1),(44) and also with probability at least1−δ, TX t=1 Et−1[Xt]≤(1 +ε) TX t=1 Xt + (1 +ε) 2R ε log(δ−1).(45)","citing_arxiv_id":"2605.12316"}]},"error":null,"updated_at":"2026-05-22T20:43:57.898195+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-22T20:43:57.790986+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"OpenAI blog , volume=","claims":[{"claim_text":"org/10.18653/v1/ 2025.emnlp-main.999 Portelance, E., Duan, Y., Frank, M. C., & Lupyan, G. (2023). Predicting age of acquisition for children's early vocabulary in five languages using language model surprisal.Cognitive Science,47(9), e13334. Portelance,E.,&Jasbi,M.(2024).Therolesofneuralnetworks inlanguageacquisition.LanguageandLinguisticsCompass, 18(6), e70001. https://doi.org/https://doi.org/10.1111/lnc3. 70001 Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Langu","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"out(x 0, x1)pairs and construct flow states at 20 randomly sampled timesteps and the clean target (t=1.0), producing4 200total evaluations per source distribution. We computeEϕ(xt)and report binned mean energy, monotonicity, rank correlation, and aggregate metrics in tables 9 and 10. Table 9Time-energy monotonicity.Mean energyE ϕ(xt)per time bin; both source distributions. Time bin Src[0,.1) [.1,.2) [.2,.3) [.3,.4) [.4,.5) [.5,.6) [.6,.7) [.7,.8) [.8,.9) [.9,1)t=1 Uni. ¯Eϕ −0.12−0.57−0.75−0.90−1","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Let X be a random variable taking values in [−B, B]. Assume thatE \u0002 e−ηX\u0003 ≤1, then E[X2]≤4 \u00121 η +B \u0013 E[X]. Lemma 6(Freedman's inequality (Beygelzimer et al., 2011)).Let (Xt)t≤T be a real-valued mar- tingale difference sequence adapted to a filtration(Ft)t≤T . If |Xt| ≤R almost surely, then for any η∈(0,1/R), with probability at least1−δ, TX t=1 Xt ≤η TX t=1 Et−1 \u0002 X2 t \u0003 + log(δ−1) η .(43) The following result is a consequence of Lemma 6. Lemma 7.Let (Xt)t≤T be a sequence of random variables ada","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Problem Formulation Let the prompt space be P ⊂ {V,V 2, . . .} where V is model vocabulary, the protected agent prompt be x∈ P and θ be model parameters.Our objectiveis to identify another prompt ˜x∈ P that satisfy the following criteria: (1)Obfuscation (C2): ˜x is distant from x in prompt space; (2)Usability (C3): ˜x is functionally equivalent to x; (3) Non-portability (C4): the utility objective is satisfied only on the target LLM, not other LLMs. 3.2. Theoretical Motivation In this subsection","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"where ˆ𝜶 ℎ∈R𝑑×𝐶ℎ𝑖𝑑 , ˆ𝜸ℎ , ˆ𝜷 ℎ∈R𝑑𝜙×𝐶ℎ𝑖𝑑 are learnable parameters, and𝐶ℎ𝑖𝑑 is the hidden dimensional- ity of the cross-attention. This allows the model to attend to specific medical history with causality information in the latentˆywhen synthesizing the DT. The objective is to predict the noise𝜖added to the latent representation𝑧𝑡 ℒLDM =E 𝑧∼ℰ(M),𝑡,𝜖,𝑦 \u0002 ∥𝜖−𝜖𝜃(𝑧𝑡 , 𝑡,ˆy)∥2\u0003 (10) where 𝜖𝜃 is the MLP U-Net. During inference, we sample a latent𝑧0 conditioned on the patient's history and decode it us","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"Ego Network Model (b) Interaction Record from A to B Support Clique … … Like Comment Retweet 40% 40% 10% (a) Tie Strength CoT Training Dataset A CoT Construction User A User B (a) Multi-hop Following Relationship Sampling (b) Graph-to-(sequential) Text Encoding User B User A User C User D User E User F User A User B User B User C … User E User F [Hop 2] …… [Hop 5] ……… CoT Construction Module 1 Module 2 Step 1: Community initialization Step 2: Follow relationship generation Step 3: Interaction Be","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks OpenAI blog , volume= because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (6 contexts).","role_counts":[{"n":6,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-05-22T20:43:52.927744+00:00"}},"summary":{"title":"OpenAI blog , volume=","claims":[{"claim_text":"org/10.18653/v1/ 2025.emnlp-main.999 Portelance, E., Duan, Y., Frank, M. C., & Lupyan, G. (2023). Predicting age of acquisition for children's early vocabulary in five languages using language model surprisal.Cognitive Science,47(9), e13334. Portelance,E.,&Jasbi,M.(2024).Therolesofneuralnetworks inlanguageacquisition.LanguageandLinguisticsCompass, 18(6), e70001. https://doi.org/https://doi.org/10.1111/lnc3. 70001 Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Langu","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"out(x 0, x1)pairs and construct flow states at 20 randomly sampled timesteps and the clean target (t=1.0), producing4 200total evaluations per source distribution. We computeEϕ(xt)and report binned mean energy, monotonicity, rank correlation, and aggregate metrics in tables 9 and 10. Table 9Time-energy monotonicity.Mean energyE ϕ(xt)per time bin; both source distributions. Time bin Src[0,.1) [.1,.2) [.2,.3) [.3,.4) [.4,.5) [.5,.6) [.6,.7) [.7,.8) [.8,.9) [.9,1)t=1 Uni. ¯Eϕ −0.12−0.57−0.75−0.90−1","claim_type":"background","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Let X be a random variable taking values in [−B, B]. Assume thatE \u0002 e−ηX\u0003 ≤1, then E[X2]≤4 \u00121 η +B \u0013 E[X]. Lemma 6(Freedman's inequality (Beygelzimer et al., 2011)).Let (Xt)t≤T be a real-valued mar- tingale difference sequence adapted to a filtration(Ft)t≤T . If |Xt| ≤R almost surely, then for any η∈(0,1/R), with probability at least1−δ, TX t=1 Xt ≤η TX t=1 Et−1 \u0002 X2 t \u0003 + log(δ−1) η .(43) The following result is a consequence of Lemma 6. Lemma 7.Let (Xt)t≤T be a sequence of random variables ada","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"},{"claim_text":"Problem Formulation Let the prompt space be P ⊂ {V,V 2, . . .} where V is model vocabulary, the protected agent prompt be x∈ P and θ be model parameters.Our objectiveis to identify another prompt ˜x∈ P that satisfy the following criteria: (1)Obfuscation (C2): ˜x is distant from x in prompt space; (2)Usability (C3): ˜x is functionally equivalent to x; (3) Non-portability (C4): the utility objective is satisfied only on the target LLM, not other LLMs. 3.2. Theoretical Motivation In this subsection","claim_type":"background","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"where ˆ𝜶 ℎ∈R𝑑×𝐶ℎ𝑖𝑑 , ˆ𝜸ℎ , ˆ𝜷 ℎ∈R𝑑𝜙×𝐶ℎ𝑖𝑑 are learnable parameters, and𝐶ℎ𝑖𝑑 is the hidden dimensional- ity of the cross-attention. This allows the model to attend to specific medical history with causality information in the latentˆywhen synthesizing the DT. The objective is to predict the noise𝜖added to the latent representation𝑧𝑡 ℒLDM =E 𝑧∼ℰ(M),𝑡,𝜖,𝑦 \u0002 ∥𝜖−𝜖𝜃(𝑧𝑡 , 𝑡,ˆy)∥2\u0003 (10) where 𝜖𝜃 is the MLP U-Net. During inference, we sample a latent𝑧0 conditioned on the patient's history and decode it us","claim_type":"background","confidence":0.6,"evidence_strength":"citation_context"},{"claim_text":"Ego Network Model (b) Interaction Record from A to B Support Clique … … Like Comment Retweet 40% 40% 10% (a) Tie Strength CoT Training Dataset A CoT Construction User A User B (a) Multi-hop Following Relationship Sampling (b) Graph-to-(sequential) Text Encoding User B User A User C User D User E User F User A User B User B User C … User E User F [Hop 2] …… [Hop 5] ……… CoT Construction Module 1 Module 2 Step 1: Community initialization Step 2: Follow relationship generation Step 3: Interaction Be","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks OpenAI blog , volume= because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (6 contexts).","role_counts":[{"n":6,"context_role":"background"},{"n":1,"context_role":"method"}]},"graph":{"co_cited":[{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":26},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":21},{"title":"Advances in neural information processing systems , volume=","work_id":"12f5a236-ef7a-4d13-b4de-b51465a6f977","shared_citers":20},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":16},{"title":"Decoupled Weight Decay Regularization","work_id":"07ef7360-d385-4033-83f7-8384a6325204","shared_citers":14},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":14},{"title":"OPT: Open Pre-trained Transformer Language Models","work_id":"d7ff3b21-1fff-4cf4-952a-4714e3ef2307","shared_citers":13},{"title":"Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge","work_id":"28ea1282-d657-4c61-a83c-f1249be6d6b1","shared_citers":13},{"title":"Advances in neural information processing systems , volume=","work_id":"a1fd09f1-b62b-4aca-a5ef-dd2b50ad08b5","shared_citers":12},{"title":"Scaling Laws for Neural Language Models","work_id":"b7dd8749-9c45-4977-ab9b-64478dce1ae8","shared_citers":12},{"title":"Training Compute-Optimal Large Language Models","work_id":"b2faf28d-86b7-429c-bc42-469458efc246","shared_citers":12},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":11},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":11},{"title":"BLOOM: A 176B-Parameter Open-Access Multilingual Language Model","work_id":"337ba690-f35d-4154-9450-8edf4bc9f488","shared_citers":10},{"title":"On the Opportunities and Risks of Foundation Models","work_id":"a18039e9-928d-47c9-a836-32656a71bf71","shared_citers":10},{"title":"Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback","work_id":"a1f2574b-a899-4713-be60-c87ba332656c","shared_citers":10},{"title":"Advances in Neural Information Processing Systems , volume=","work_id":"bd577a47-49ef-4f8a-81ca-86cd88b71479","shared_citers":9},{"title":"Language Models are Few-Shot Learners","work_id":"214732c0-2edd-44a0-af9e-28184a2b8279","shared_citers":9},{"title":"Language models are unsupervised multitask learners","work_id":"175d0d95-85f4-480c-b6d3-53a7e52df1e4","shared_citers":9},{"title":"Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism","work_id":"c888e6d1-0b1d-43d6-9ef5-f0912a0efa1b","shared_citers":9},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":9},{"title":"SC20: International Conference for High Performance Computing, Networking, Storage and Analysis , pages=","work_id":"f62f435f-90be-45ec-bdea-343bdc2734f6","shared_citers":9},{"title":"2016 , publisher=","work_id":"cf0899e0-53ee-4591-aae4-f38fa5ac12ad","shared_citers":8},{"title":"Adam: A Method for Stochastic Optimization","work_id":"1910796d-9b52-4683-bf5c-de9632c1028b","shared_citers":8}],"time_series":[{"n":1,"year":2021},{"n":2,"year":2022},{"n":13,"year":2023},{"n":7,"year":2024},{"n":1,"year":2025},{"n":42,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"Autoregressive Learning in Joint KL: Sharp Oracle Bounds and Lower Bounds","primary_cat":"cs.LG","context_text":"Let X be a random variable taking values in [−B, B]. Assume thatE \u0002 e−ηX\u0003 ≤1, then E[X2]≤4 \u00121 η +B \u0013 E[X]. Lemma 6(Freedman's inequality (Beygelzimer et al., 2011)).Let (Xt)t≤T be a real-valued mar- tingale difference sequence adapted to a filtration(Ft)t≤T . If |Xt| ≤R almost surely, then for any η∈(0,1/R), with probability at least1−δ, TX t=1 Xt ≤η TX t=1 Et−1 \u0002 X2 t \u0003 + log(δ−1) η .(43) The following result is a consequence of Lemma 6. Lemma 7.Let (Xt)t≤T be a sequence of random variables adapted to a filtration (Ft)t≤T . If 0≤X t ≤Ralmost surely, then with probability at least1−δ, TX t=1 Xt ≤(1 +ε) TX t=1 Et−1[Xt] + R ε log(δ−1),(44) and also with probability at least1−δ, TX t=1 Et−1[Xt]≤(1 +ε) TX t=1 Xt + (1 +ε) 2R ε log(δ−1).(45)","citing_arxiv_id":"2605.12316"}]},"authors":[]}}