Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T07:15:23.725397Z
Paper Citation Record · LEDGER
As of 24 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 100 inbound Pith citation observations for arXiv:2210.09261.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-11T07:15:23.725397Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T12:06:23.497019Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
42 of 42 outbound references displayed
External citation measurements
43
pith, observed 2026-08-05T02:28:24.338817Z
Observation 139f30fc-d28b-4343-9148-bd909e57aa28 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Program Synthesis with Large Language Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation acddba65-daf0-40f7-9100-e30585c1b417 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Language models are few-shot learners
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6dccc3fc-f32c-40ab-ab2f-3159e18ec2cb · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Evaluating Large Language Models Trained on Code
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 79ccd860-217a-4c3a-8d1b-d8ef05da769c · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Binding Language Models in Symbolic Languages
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0c0b945-7a89-40f8-ba73-fb912374962d · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them PaLM: Scaling Language Modeling with Pathways
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f005f24f-786a-4632-8139-017886228b40 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a760d8ff-71d1-4099-a915-cc4a0e1b9a81 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them BERT: Pre-training of deep bidirectional transformers for language understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9606bdc8-4492-48c0-8535-90d7666127eb · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8c667d39-a06e-458a-9e3e-ff1fa6290db1 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Predictability and surprise in large generative models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ec7c3819-619a-4d5f-b441-51e8cf0f0afd · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them doi: 10.18653/v1/2022.lnls-1.4
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation db24b1b4-0fcd-41f4-a370-2b0ba4560c7e · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Training Compute-Optimal Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 51b424ed-a4fa-4891-997c-a44498ae50b7 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Language Models (Mostly) Know What They Know
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8eb2db11-1a53-478a-8d23-97021e0b0107 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Large Language Models are Zero-Shot Reasoners
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 305a21fa-7344-4800-9742-eb55c66f988b · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Can language models learn from explanations in context?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 93c3de3b-3f6b-41a5-861f-4c7873294fff · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them The power of scale for parameter-efficient prompt tuning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4b1ba2fe-294d-417d-89ea-e4648efad14b · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Making Large Language Models Better Reasoners with Step-Aware Verifier
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6fbdbbb5-6289-4254-b8bc-7fd1f4492dbb · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d5a89317-0e3c-4829-95fd-1395f931bd28 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them AmbiPun: Generating Humorous Puns with Ambiguous Context
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd7ee3ad-eaaa-4649-9e01-5e711d4c4ba8 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Show Your Work: Scratchpads for Intermediate Computation with Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f7d7a848-228f-4eed-9d67-9df47b46083c · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Training language models to follow instructions with human feedback
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f1cacc23-6704-4283-9fb3-2605388a6c9e · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Scaling Language Models: Methods, Analysis & Insights from Training Gopher
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2beb95dc-4a3d-4c72-9b2e-a45488c0319c · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Language Models are Multilingual Chain-of-Thought Reasoners
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation ca686e16-f7fd-4a9b-9aa8-209b6866ec99 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 44f3be94-b7ff-4f14-bad4-572625b01e0e · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Supervising Model Attention with Human Explanations for Robust Natural Language Inference
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0761b1ec-59cb-4d05-a049-06bc526ccb92 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 03afd620-e358-4239-917f-2c503614cdc8 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them On the machine learning of ethical judgments from natural language
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 531a4e0d-1024-4c0e-9a46-b30274c20eab · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b8965c91-0608-4935-b298-9c5ca332b766 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 468806fe-57e1-4417-9b95-ec668a8d4753 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Do Prompt-Based Models Really Understand the Meaning of their Prompts?
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c90b0d1-3d6a-495e-bcc8-4df51c5342b9 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Finetuned language models are zero-shot learners
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 480e90e9-1442-463f-8a75-140e69f51cf3 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Emergent abilities of large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d9c1d26c-5fa7-4781-868c-27109588a307 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them In: Zong, C., Xia, F., Li, W., Navigli, R
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 9c32e6b1-ad75-46d8-8758-c7cbc61597ac · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them An Explanation of In-context Learning as Implicit Bayesian Inference
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8149b2b4-0d23-40a7-a5b1-41952995e9f6 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c31f716e-99c9-4921-9a0c-e2f375813af6 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them The concert was scheduled to be on 06/01/1943, but was delayed by one day to today. What is the date yesterday in MM/DD/YYYY?
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 11a821fa-b545-45ed-b203-483b607a5982 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them If today is Christmas Eve of 1937, then today's date is December 24
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b4c151a8-9de5-422f-8989-f97b03f14ef9 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them So the answer is (D)
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2286532c-d893-4374-bff4-4bdadecce391 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them What is the date tomorrow in MM/DD/YYYY? Options: (A) 01/11/1961 (B) 01/03/1963 (C) 01/18/1961 (D) 10/14/1960 (E) 01/03/1982 (F) 12/03/1960 A: Let's think step by step
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c0534341-9d86-413c-b147-6a4f60eef4a8 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them [ { [". We will need to pop out
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2c00ffa0-59d2-42f5-bc6f-31b70274fa5f · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them So the answer is (C)
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 348249cf-404c-401a-ac50-e7d37b76fc31 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Amongst all the options, the only movie similar to these ones seems to be Forrest Gump (comedy, drama, romance; 1994)
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e5f91c67-0523-489a-9ca1-6afff9eb5583 · outbound
Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them So the answer is (D)
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f0333b25-5c63-40e7-8461-841883c1209d · inbound
Emergent Abilities of Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3b7da897-ea02-44ca-960c-b833576979fc · inbound
Large Language Models Are Human-Level Prompt Engineers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7b0cca8d-9e7c-4ed4-bb13-95e217394241 · inbound
Galactica: A Large Language Model for Science Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bddf11af-ec0a-4edc-b1b5-be632b0b97dc · inbound
Galactica: A Large Language Model for Science Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 239
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b6d7bc77-56a8-441a-870b-61d95b38c45f · inbound
PAL: Program-aided Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3df3f9de-a657-4dc7-bb3f-ff7ba0676cc3 · inbound
Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2676db0a-d9b8-45b7-ba3b-7aca5b8e5b5f · inbound
The Flan Collection: Designing Data and Methods for Effective Instruction Tuning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c73acf03-6702-4408-bbb9-ace0c7aded1d · inbound
ART: Automatic multi-step reasoning and tool-use for large language models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 139
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 68cb8867-7fa5-4d80-b8cf-4384ca190bef · inbound
BloombergGPT: A Large Language Model for Finance Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 111
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c91e47a0-7d56-4fc0-a359-af894c635074 · inbound
Teaching Large Language Models to Self-Debug Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 126
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e1bbe217-6aa7-4138-a51f-bbc01a072d8a · inbound
PaLM 2 Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f6b9d02f-a752-4d8e-907f-78e4dad68e3e · inbound
Orca: Progressive Learning from Complex Explanation Traces of GPT-4 Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6a15a149-943a-447b-8420-a36dedb74cb0 · inbound
Simple synthetic data reduces sycophancy in large language models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation b085600b-ffa1-4f92-9e5a-ddb97706d2fa · inbound
Large Language Models as Optimizers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2a8951ca-0d3b-4519-92c1-079ad2a6ba44 · inbound
MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2452ba05-f82e-449f-849f-ab857b1dfd5b · inbound
EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 174
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 58747a67-a160-4111-b430-6ef360966dbd · inbound
Baichuan 2: Open Large-scale Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6c697736-8f0d-4ef0-ad6e-9d587565a606 · inbound
Mistral 7B Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8c92b47f-a279-4e35-9317-1968abfce620 · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c4ff6870-4368-4756-89e3-0147b47cd79b · inbound
Gemini: A Family of Highly Capable Multimodal Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation faedab8f-a203-430e-adf0-ff7ba29349b1 · inbound
DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation e9d23976-1bd5-4960-8377-340f7f31e2bf · inbound
Mixtral of Experts Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 37cde923-24da-4194-b6d7-50ef693b4a29 · inbound
DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cf57d103-882b-41a8-b1b5-8d95b2eea44f · inbound
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation adba7707-e189-428a-8f4c-33132d408ad4 · inbound
DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a434e9a6-928c-4296-aba4-a533e097827c · inbound
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 8072bf01-20ef-4a54-8a5f-ed70093276c1 · inbound
Yi: Open Foundation Models by 01.AI Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0b330107-2b0b-4a25-9099-5b686843acc6 · inbound
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 63d22337-1b03-41f1-8275-a38837e26b52 · inbound
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 049ef5fc-793f-4d5f-b474-063a1f4ffff8 · inbound
Lessons from the Trenches on Reproducible Evaluation of Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 122
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 825355f2-d67c-41ec-a68a-825ff5a528a9 · inbound
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0d258557-47dd-4de1-85a6-5c5b69fa9df4 · inbound
DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation de55283b-4306-4471-b9bd-87a58a89dd0a · inbound
LiveBench: A Challenging, Contamination-Limited LLM Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 07bff1ea-da89-472d-ad0b-613b07be387b · inbound
Reinforcement Learning for LLM Post-Training: A Survey Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c50530f3-288b-4c2e-a742-ac28c1f143b2 · inbound
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 271
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 60b36141-770e-44b0-8135-0e137ef45a1e · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 42c61861-9c70-4723-982d-d25e9f19a944 · inbound
AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cfb8def-264d-444d-84a0-a691b3a8f32a · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2798ba0c-1687-4752-802c-a3c2dcf5fba7 · inbound
Ultra-Sparse Memory Network Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42aa4d11-1f1f-458a-b147-2bdf34e8ff59 · inbound
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abd24fd7-f66e-41d7-8f77-b9a025491119 · inbound
A Survey on Human-Centric LLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2ee037-24a5-499c-8b7d-106542d005ee · inbound
INTELLECT-1 Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d27fd8c6-5f55-4951-986c-35a6f51d64c6 · inbound
REVOLVE: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a1352d5-c874-4066-9051-63d8d304021d · inbound
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da0a4435-febc-4a1d-9d56-ecac1858699a · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 225
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c18562d5-1b2c-48ca-959c-f8b2d384699c · inbound
Multi-Party Supervised Fine-tuning of Language Models for Multi-Party Dialogue Generation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c89f66-151c-430b-b17b-827b4403e3d3 · inbound
Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8447ce42-b0ea-44f2-adf6-2e3dd18b9bb8 · inbound
From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c037a59-b852-4e65-be2d-3d774f1c1c0f · inbound
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b93d489-5b7d-40f2-9c71-449c20c378d5 · inbound
Codenames as a Benchmark for Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21adfeca-1a8f-4787-92bc-2f1998ab195b · inbound
C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63fe2f30-9075-4174-899b-f79055ea0f04 · inbound
Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb619b33-7315-48c9-baeb-8f9a99d883b3 · inbound
Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d04ef4d7-7c82-4939-a46a-3e5571cc9b40 · inbound
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d0ece0-5b9d-4acd-89ef-12140dc30d40 · inbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99eb43db-ffe6-4276-9c02-268a8ee5af72 · inbound
ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd0be1f3-7eab-4981-b200-e6298c3800e8 · inbound
Language Models as Continuous Self-Evolving Data Engineers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e92359ad-a340-49ab-85d3-99b5a7882fcc · inbound
Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36bc1ab2-a1db-49a9-a90f-bcb34ab95f26 · inbound
Multi-matrix Factorization Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88b4f048-4e97-4b13-b09a-0352066c6f09 · inbound
Aligning Large Language Models for Faithful Integrity Against Opposing Argument Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3590d3-ca72-456d-945f-488df1a5c8f7 · inbound
InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7f41a0a-449a-4f73-bbb6-ff013a3286e5 · inbound
A Survey on Large Language Models with some Insights on their Capabilities and Limitations Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 218
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c19a6f9-f116-45e4-8275-9764fc426a7e · inbound
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df76838-29dd-462a-9686-812e9eef82e6 · inbound
TAPO: Task-Referenced Adaptation for Prompt Optimization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5461f440-0748-4495-9662-88d161f2d8f0 · inbound
ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfcacde6-05fc-4e24-8076-b7417dd6e3bf · inbound
Domain Adaptation of Foundation LLMs for e-Commerce Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208f4aa8-0c4b-45f6-8b11-a4772b66bdd9 · inbound
DNA 1.0 Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbd223b4-5fec-4f2a-a188-f0e9c7955c2d · inbound
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e1237a-0aaf-425e-955c-4d8a3b1f9292 · inbound
Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89aae155-c70d-49c6-a29d-b73f2fcc3f50 · inbound
Spurious Forgetting in Continual Learning of Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5bb4b6-fd4f-4e69-9f89-70d457fa39e0 · inbound
Adaptive Testing for LLM-Based Applications: A Diversity-based Approach Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fb2f21f-e0ef-42ee-8ea8-e223838bbc8b · inbound
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cae2543c-5c9c-464f-9500-08e524451f5c · inbound
On the Reasoning Capacity of AI Models and How to Quantify It Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f12650d-6bb5-41b9-b0a4-f2d72c38eb7e · inbound
StaICC: Standardized Evaluation for Classification Task in In-context Learning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f144c8-4bea-4f08-9802-e3d79dc3f769 · inbound
LLM-AutoDiff: Auto-Differentiate Any LLM Workflow Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45c70adb-9721-44a4-a224-80e0b12dad0b · inbound
BTS: Harmonizing Specialized Experts into a Generalist LLM Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13ff5323-6dbe-4dc2-9d6f-bc7aaffc4c21 · inbound
Process Reinforcement through Implicit Rewards Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1568f826-7b6a-4846-b1cf-8bc21689a904 · inbound
Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2ff57f1f-cd73-4af4-902b-9fd4fde3f128 · inbound
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 5c3ee4f6-2317-4594-92e0-52fca4eafe53 · inbound
Minerva: A Programmable Memory Test Benchmark for Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cecd3773-fd12-44e2-90f6-cbc869d0af81 · inbound
It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9b77044-99b1-45fe-9c5d-727777069825 · inbound
CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f716b3-e7bf-4484-b15f-2057fbb985da · inbound
Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a928af93-6277-4e58-a27c-acd6cafd6f1f · inbound
KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2c5ac7a-a6e2-4931-a46e-a312c619e32d · inbound
CryptoX : Compositional Reasoning Evaluation of Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83636a88-52f1-4891-b66d-98e554e2fb56 · inbound
SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6921b2ba-1521-4e3d-8b4b-623878249861 · inbound
Salamandra Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 186
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45500c75-d4ab-429d-81d9-7d08780ac842 · inbound
RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e11429-2c0e-4206-8140-eadeb4db496b · inbound
Large Language Diffusion Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation f88203ad-f0c2-4061-94a4-2b92f6933b7b · inbound
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 239
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a083efec-094a-4cb1-925f-85ccbfbf1517 · inbound
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation cfedbd3e-8392-472e-8469-63f2cc848781 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 113
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7867c7a8-2bc5-40d9-92e9-293fb68fe583 · inbound
MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0961d143-6f64-4c50-985e-aec82a16c75b · inbound
FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1652c657-4c63-43ea-8e8d-05ef74895200 · inbound
Trillion 7B Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d90395e-ccc4-4a76-8e28-8fe87d62527f · inbound
Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187cc37e-634a-4aa6-92e8-8c403a0c25f0 · inbound
CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b2eb093-5e1e-4e85-aee7-f00ab4adabc8 · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 050e536f-97e7-4336-ac09-e6c678f15308 · inbound
Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6519e7-9c1c-4eed-80db-d4a23fefdf07 · inbound
100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.