Pith. sign in

Paper Citation Record · LEDGER

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

As of 24 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 100 inbound Pith citation observations for arXiv:2210.09261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.09261 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T07:15:23.725397Z

measured 142 of 142 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 100 of 283 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:06:23.497019Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact20
  • verified fuzzy15
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch6

External citation measurements

43
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 139f30fc-d28b-4343-9148-bd909e57aa28 · outbound

This paper cites Program Synthesis with Large Language Models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Program Synthesis with Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:24.027742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:f4473c20f440bed480b9f0d4d1da58646f5789a654ea3db38b5fcdf2a723b41f

Observation acddba65-daf0-40f7-9100-e30585c1b417 · outbound

This paper cites Language models are few-shot learners.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Language models are few-shot learners

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.120171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:fc727551bc14332b9fdd064cda36812a2f258db22e95aae6f9dbf41c37fdf08c

Observation 6dccc3fc-f32c-40ab-ab2f-3159e18ec2cb · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Evaluating Large Language Models Trained on Code

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.837037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:5effb5b48cddd0b1ecd470206a32aa86c64f54610db429765da36520726ae929

Observation 79ccd860-217a-4c3a-8d1b-d8ef05da769c · outbound

This paper cites Binding Language Models in Symbolic Languages.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Binding Language Models in Symbolic Languages

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:23.845653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:f705493896e0c83110656d931fc5655793b1bb6ff8206513e0eafe3d3c2585da

Observation c0c0b945-7a89-40f8-ba73-fb912374962d · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them PaLM: Scaling Language Modeling with Pathways

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.813765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:99ec5d43651c2cc3a9f76b336282b3f97b559291db4e22e29eca91c41c4691bf

Observation f005f24f-786a-4632-8139-017886228b40 · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:23.821018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:79cb0cf2e2c05cbd6049e14d78cc6c6857909efcd0d8345527f44829fc6e72da

Observation a760d8ff-71d1-4099-a915-cc4a0e1b9a81 · outbound

This paper cites BERT: Pre-training of deep bidirectional transformers for language understanding.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them BERT: Pre-training of deep bidirectional transformers for language understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.066731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:3bc14844dd2249496d89aa4649df9aa754a431eb1789484585fc7029e603a9a2

Observation 9606bdc8-4492-48c0-8535-90d7666127eb · outbound

This paper cites BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them BERT : Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
metadata mismatch
doi, observed 2026-05-11T07:15:23.807295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:07461f9d6d6c9eefd3d1743330e18078e20287952350ec1d0043857f2615c786

Observation 8c667d39-a06e-458a-9e3e-ff1fa6290db1 · outbound

This paper cites Predictability and surprise in large generative models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Predictability and surprise in large generative models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.061553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:2f8821faabe3eb967d09aa5f677924571b99e42bb2d8fcc5f91bc61a6f468613

Observation ec7c3819-619a-4d5f-b441-51e8cf0f0afd · outbound

This paper cites doi: 10.18653/v1/2022.lnls-1.4.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them doi: 10.18653/v1/2022.lnls-1.4

Reference 10

Resolution
verified exact
doi, observed 2026-05-11T07:15:23.794252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:300a432353db518077b3a60c228854018b23596856619dd04617b1e63a9c36de

Observation db24b1b4-0fcd-41f4-a370-2b0ba4560c7e · outbound

This paper cites Training Compute-Optimal Large Language Models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Training Compute-Optimal Large Language Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.829606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:580fe87602027dbec0f9eda84c2520bdb8a6b7271c8e501302495d0e9a55ee4a

Observation 51b424ed-a4fa-4891-997c-a44498ae50b7 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Language Models (Mostly) Know What They Know

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:24.038913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:f87789d67edbec8d478763660ededacc02a46a63da40302b018c7f75f8f8b55a

Observation 8eb2db11-1a53-478a-8d23-97021e0b0107 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Large Language Models are Zero-Shot Reasoners

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:04:13.904485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:3bcb8882f6fedefe277145a4f9480f16dfb117c11b1477e2788aabcf99b29c14

Observation 305a21fa-7344-4800-9742-eb55c66f988b · outbound

This paper cites Can language models learn from explanations in context?.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Can language models learn from explanations in context?

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:24.053343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:9d24ada8707f9fd811131644f638a53b50c933e8774f3d7014a04598156a43cc

Observation 93c3de3b-3f6b-41a5-861f-4c7873294fff · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them The power of scale for parameter-efficient prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.091649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:117a60f2242d1059b5f769f5ead2d31ebda708c71866af9632c9ef5f0d9d7390

Observation 4b1ba2fe-294d-417d-89ea-e4648efad14b · outbound

This paper cites Making Large Language Models Better Reasoners with Step-Aware Verifier.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Making Large Language Models Better Reasoners with Step-Aware Verifier

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:23.855070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:55460c8e006b351b5b3c9172483399aa62cb6f500432ea43645fc16d4776e1d7

Observation 6fbdbbb5-6289-4254-b8bc-7fd1f4492dbb · outbound

This paper cites Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:51:47.316034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:9282e6ffe5215e259c98b19b7f08d1da85e5e6720450c1b3d219979cc5f072cd

Observation d5a89317-0e3c-4829-95fd-1395f931bd28 · outbound

This paper cites AmbiPun: Generating Humorous Puns with Ambiguous Context.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them AmbiPun: Generating Humorous Puns with Ambiguous Context

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:23.877092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:ba9854450253de2f18ee6a55bc8de218feda366623d3f5382051033ab7c5c5f3

Observation bd7ee3ad-eaaa-4649-9e01-5e711d4c4ba8 · outbound

This paper cites Show Your Work: Scratchpads for Intermediate Computation with Language Models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Show Your Work: Scratchpads for Intermediate Computation with Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T00:31:41.641039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:1361b7c54d34de364fcfc8f749f1dc313c513651a2e122ce31e947c896b198bb

Observation f7d7a848-228f-4eed-9d67-9df47b46083c · outbound

This paper cites Training language models to follow instructions with human feedback.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Training language models to follow instructions with human feedback

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.895169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:ce68836fde8bac3b4c795c30a8635dfdaa44b1d39bb2081706b4256f12580b7c

Observation f1cacc23-6704-4283-9fb3-2605388a6c9e · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:13:40.454049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:cab617c944ecddbdc12a7831998c0a0ba59090d6afee80b2b46841bda020d78f

Observation 2beb95dc-4a3d-4c72-9b2e-a45488c0319c · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Language Models are Multilingual Chain-of-Thought Reasoners

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:06:10.854128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:67118c3420d02d000adf4e04e5cc1328a0c52c77dd37a26fb71526afda8c266f

Observation ca686e16-f7fd-4a9b-9aa8-209b6866ec99 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.956384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:12b7aaf3da67aea2a42ce2fb8885e9b877898775b1fb5989e21bc157f5f24035

Observation 44f3be94-b7ff-4f14-bad4-572625b01e0e · outbound

This paper cites Supervising Model Attention with Human Explanations for Robust Natural Language Inference.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Supervising Model Attention with Human Explanations for Robust Natural Language Inference

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:23.966839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:736c1a75b38a3e73875cc389a93714fdf8c6c84a3a6a52cd58f7e967489463e9

Observation 0761b1ec-59cb-4d05-a049-06bc526ccb92 · outbound

This paper cites Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:23.975130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:b88a487e72d21b0fb60467199f8131189aa8a421e2e970efa413e4abf2afde07

Observation 03afd620-e358-4239-917f-2c503614cdc8 · outbound

This paper cites On the machine learning of ethical judgments from natural language.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them On the machine learning of ethical judgments from natural language

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.082993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:da761e29607492874115ef0fc23a161ed4ebbbb7a28d018eab6a19d735bd486b

Observation 531a4e0d-1024-4c0e-9a46-b30274c20eab · outbound

This paper cites Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Scaling Laws vs Model Architectures: How does Inductive Bias Influence Scaling?

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:23.984689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:a8ab17692823e64abfce928a5f3c0f7e56c09bd8f263e1262e0de4b473210eec

Observation b8965c91-0608-4935-b298-9c5ca332b766 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:15:23.995615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:edfbeb2236faccefdc25b7afa21fa69e7683f3925b27eac98223b5746d70df3c

Observation 468806fe-57e1-4417-9b95-ec668a8d4753 · outbound

This paper cites Do Prompt-Based Models Really Understand the Meaning of their Prompts?.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Do Prompt-Based Models Really Understand the Meaning of their Prompts?

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:24.008493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:2b6bda45d0735fcc4c4f4096b8cedf41a8e991c91f834885c416095c1070b31c

Observation 5c90b0d1-3d6a-495e-bcc8-4df51c5342b9 · outbound

This paper cites Finetuned language models are zero-shot learners.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Finetuned language models are zero-shot learners

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.109528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:0c062a2784425835d586ef0c45fa32541591e4c86330fc2d4ffaaf35fbcee214

Observation 480e90e9-1442-463f-8a75-140e69f51cf3 · outbound

This paper cites Emergent abilities of large language models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Emergent abilities of large language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.116398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:a0c8eb6338148f0d66ce9d73c2d85abd6540c1b40cbf893e78b6002526fbb16c

Observation d9c1d26c-5fa7-4781-868c-27109588a307 · outbound

This paper cites In: Zong, C., Xia, F., Li, W., Navigli, R.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them In: Zong, C., Xia, F., Li, W., Navigli, R

Reference 32

Resolution
malformed identifier
doi_truncated, observed 2026-05-11T07:15:23.789348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:4264f462ad2318c00604d93e0e4ba6aa8fd34dc4d25bdaa4f63cbc6b941569e2

Observation 9c32e6b1-ad75-46d8-8758-c7cbc61597ac · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:22:08.467971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:a6bd93f2920fc547f23e1e96c4af6aea55826e4167b88009f9e09cf8e0b67c93

Observation 8149b2b4-0d23-40a7-a5b1-41952995e9f6 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:44:37.129111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:15fcb980eee61905df72ccf696b2cca4446fe12c68ee870c5f6dceccd23c7227

Observation c31f716e-99c9-4921-9a0c-e2f375813af6 · outbound

This paper cites The concert was scheduled to be on 06/01/1943, but was delayed by one day to today. What is the date yesterday in MM/DD/YYYY?.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them The concert was scheduled to be on 06/01/1943, but was delayed by one day to today. What is the date yesterday in MM/DD/YYYY?

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.131928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:edeed79f9fd96e926c5259a24c620050e9f8e32b6bb2ad8430a74734eae04c94

Observation 11a821fa-b545-45ed-b203-483b607a5982 · outbound

This paper cites If today is Christmas Eve of 1937, then today's date is December 24.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them If today is Christmas Eve of 1937, then today's date is December 24

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.071693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:3e359b404a26f2f69e5b50256c7c4de95bbee08defd68e2eb2245cd8e7dd766a

Observation b4c151a8-9de5-422f-8989-f97b03f14ef9 · outbound

This paper cites So the answer is (D).

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them So the answer is (D)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.078440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:13840732429621bfd806ae0c1fd7307277b249982b57dd90c4411fa1b1302ed3

Observation 2286532c-d893-4374-bff4-4bdadecce391 · outbound

This paper cites What is the date tomorrow in MM/DD/YYYY? Options: (A) 01/11/1961 (B) 01/03/1963 (C) 01/18/1961 (D) 10/14/1960 (E) 01/03/1982 (F) 12/03/1960 A: Let's think step by step.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them What is the date tomorrow in MM/DD/YYYY? Options: (A) 01/11/1961 (B) 01/03/1963 (C) 01/18/1961 (D) 10/14/1960 (E) 01/03/1982 (F) 12/03/1960 A: Let's think step by step

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.087626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:6d014254e5f36298e4b941b87a997fd77323240fe19d78ad5357bc9c59a6eb66

Observation c0534341-9d86-413c-b147-6a4f60eef4a8 · outbound

This paper cites [ { [". We will need to pop out.

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them [ { [". We will need to pop out

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.097936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:282cb271a46e70cb60fafd785cf782acd21fd71c301070d96f4c6a21857454a8

Observation 2c00ffa0-59d2-42f5-bc6f-31b70274fa5f · outbound

This paper cites So the answer is (C).

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them So the answer is (C)

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.102188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:6f932118583fee68f211c3dc950c0719541f87344773564acf5a17f2987a4cf5

Observation 348249cf-404c-401a-ac50-e7d37b76fc31 · outbound

This paper cites Amongst all the options, the only movie similar to these ones seems to be Forrest Gump (comedy, drama, romance; 1994).

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them Amongst all the options, the only movie similar to these ones seems to be Forrest Gump (comedy, drama, romance; 1994)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.124001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:8e8ac9dae955b5e9aa70449ee7d65025f723641ed98f958ff8ddf3955bf3bd0c

Observation e5f91c67-0523-489a-9ca1-6afff9eb5583 · outbound

This paper cites So the answer is (D).

Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them So the answer is (D)

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T07:15:24.127300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T07:15:23.725397Z digest=sha256:b24d67c232afbe80445996d4b7eeda687e01826b44ad6e7ec7f550297a802eed

Pith citing papers

Observation f0333b25-5c63-40e7-8461-841883c1209d · inbound

Emergent Abilities of Large Language Models cites this paper.

Emergent Abilities of Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:38:38.197917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T07:38:37.734402Z digest=sha256:db821d48cb7e27766fcb9db18527939c8a5eadac64d175ea050b390d563473f2

Observation 3b7da897-ea02-44ca-960c-b833576979fc · inbound

Large Language Models Are Human-Level Prompt Engineers cites this paper.

Large Language Models Are Human-Level Prompt Engineers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:43:26.350456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T09:43:26.288866Z digest=sha256:dd94fe5961a21e211a2a2e73e85644099f3c05691900d18d6339b79441a0d0fc

Observation 7b0cca8d-9e7c-4ed4-bb13-95e217394241 · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:53:22.040938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:b424ab2fe4d023c7c84d536d277fc693d82f336d9b76d04a39202248d36d3dd5

Observation bddf11af-ec0a-4edc-b1b5-be632b0b97dc · inbound

Galactica: A Large Language Model for Science cites this paper.

Galactica: A Large Language Model for Science Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 239

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:53:22.264120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T05:53:21.810346Z digest=sha256:8e9e851a658a3d3b439dd09cbb000eef4848cdaf84ba0141fd6a4bc688e067f2

Observation b6d7bc77-56a8-441a-870b-61d95b38c45f · inbound

PAL: Program-aided Language Models cites this paper.

PAL: Program-aided Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T05:02:50.153759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-15T05:02:50.117120Z digest=sha256:ad8dd882ac2ea7c98b9b18e46c01dd6aa4406017bd92683140fdb412f1b56db6

Observation 3df3f9de-a657-4dc7-bb3f-ff7ba0676cc3 · inbound

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks cites this paper.

Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T16:48:27.967874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T16:48:27.918334Z digest=sha256:4320737c0043d584eb377d2ac39901579c1d2f32a59c082d5ba9d2d72a5fd4eb

Observation 2676db0a-d9b8-45b7-ba3b-7aca5b8e5b5f · inbound

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning cites this paper.

The Flan Collection: Designing Data and Methods for Effective Instruction Tuning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-24T09:14:16.498887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T09:13:30.054153Z digest=sha256:e15670dbc0a7119779bd9680eb9aaec921b9126018b8faa80466dbfc6cf0a204

Observation c73acf03-6702-4408-bbb9-ace0c7aded1d · inbound

ART: Automatic multi-step reasoning and tool-use for large language models cites this paper.

ART: Automatic multi-step reasoning and tool-use for large language models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 139

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T19:03:06.201861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T19:03:05.597295Z digest=sha256:5448007f5f68479abe0e5c4a0b32b912af98473872e1f44001b582ac9cdeaf43

Observation 68cb8867-7fa5-4d80-b8cf-4384ca190bef · inbound

BloombergGPT: A Large Language Model for Finance cites this paper.

BloombergGPT: A Large Language Model for Finance Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 111

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:19:46.371352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T23:19:46.231145Z digest=sha256:fe8932baa284048d9788a2487956de11730e490e25f5033b55d8038c04127b9e

Observation c91e47a0-7d56-4fc0-a359-af894c635074 · inbound

Teaching Large Language Models to Self-Debug cites this paper.

Teaching Large Language Models to Self-Debug Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 126

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:24:24.878306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T06:24:24.607354Z digest=sha256:3f6839d4a3ccbf83c2d1a6666cef08edb24862c224fdc3d6bdbd6b584eec1f47

Observation e1bbe217-6aa7-4138-a51f-bbc01a072d8a · inbound

PaLM 2 Technical Report cites this paper.

PaLM 2 Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 141

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T11:59:27.410500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-12T11:59:25.813128Z digest=sha256:10c61a6a7440a8e37fd24fac3a7420e9eb810653ba9d3b9cd9f4fd3dd5ef3204

Observation f6b9d02f-a752-4d8e-907f-78e4dad68e3e · inbound

Orca: Progressive Learning from Complex Explanation Traces of GPT-4 cites this paper.

Orca: Progressive Learning from Complex Explanation Traces of GPT-4 Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T09:42:05.301378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T09:42:05.268897Z digest=sha256:f3f0f3b53c71175bc4caec81f9dd3470ab4184ee7aec6bed4cd0a74d446d4aa4

Observation 6a15a149-943a-447b-8420-a36dedb74cb0 · inbound

Simple synthetic data reduces sycophancy in large language models cites this paper.

Simple synthetic data reduces sycophancy in large language models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:48:08.691124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T14:48:08.508109Z digest=sha256:24024e1d39062ac7a02acbe94693bbd8690d6346e45adff70a4aa719e20c66a5

Observation b085600b-ffa1-4f92-9e5a-ddb97706d2fa · inbound

Large Language Models as Optimizers cites this paper.

Large Language Models as Optimizers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:04:31.361212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T00:04:31.212102Z digest=sha256:72b125618fe15bb797e6813cad20f078552d32ebcfe96818e843d2371e4d9c38

Observation 2a8951ca-0d3b-4519-92c1-079ad2a6ba44 · inbound

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning cites this paper.

MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-17T23:46:39.456545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-17T23:46:39.330438Z digest=sha256:17057e40ee9e192b85c90085020d34fe5b48fb0b8122887dac7dbe1bf6badd8d

Observation 2452ba05-f82e-449f-849f-ab857b1dfd5b · inbound

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers cites this paper.

EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 174

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:11:49.617688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T06:11:49.475825Z digest=sha256:6e5ffd473a7ad5a75d9dbfc3f98c4b74cac304dde8d7b5984a769361c1f9b42a

Observation 58747a67-a160-4111-b430-6ef360966dbd · inbound

Baichuan 2: Open Large-scale Language Models cites this paper.

Baichuan 2: Open Large-scale Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:54:03.539145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T06:51:02.531751Z digest=sha256:fca4d1829574f92362996f996dafb0b519b1d9675cbfa86214b687e4b7c8c7aa

Observation 6c697736-8f0d-4ef0-ad6e-9d587565a606 · inbound

Mistral 7B cites this paper.

Mistral 7B Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-24T06:14:00.034133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T06:11:38.406350Z digest=sha256:b7b1add3e06cc940e16dd8eea149dd66fe635128dc46cf54077ee824162c3706

Observation 8c92b47f-a279-4e35-9317-1968abfce620 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-18T03:18:51.753187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:09019e32d5ad9f97c47c85f161dd96c5f65fdd77a24184de05326665cc0414af

Observation c4ff6870-4368-4756-89e3-0147b47cd79b · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:03:55.523806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:8922bea7ce74111905d495f2eee28909489f539eae8d4b0923fcb6ffafbc4075

Observation faedab8f-a203-430e-adf0-ff7ba29349b1 · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:ea9256ef8688a75e985ceb10da69487798c4de30afb8329ada7d75d50f2a45cb

Observation e9d23976-1bd5-4960-8377-340f7f31e2bf · inbound

Mixtral of Experts cites this paper.

Mixtral of Experts Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-24T04:13:53.877887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T04:09:15.921778Z digest=sha256:c09499735e40b6fdc7369c5f84b5f62ce73dfed926d3e829a85bab5331399c48

Observation 37cde923-24da-4194-b6d7-50ef693b4a29 · inbound

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models cites this paper.

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 96

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T22:50:13.705634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T22:50:06.399707Z digest=sha256:ad10712d63cd63ddbba97d92bf1dcb4384340fbb700badca828b866a36a020af

Observation cf57d103-882b-41a8-b1b5-8d95b2eea44f · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 105

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T10:36:18.171383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:c47ebb816a3dbf68cc45ede1b5c45d0e9b4a3a4089e6a6326cc9d4ad1e9c1a78

Observation adba7707-e189-428a-8f4c-33132d408ad4 · inbound

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence cites this paper.

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T17:26:38.812485Z digest=sha256:6975ef5829e1e788e87b5369ca4a653dff34ea062e19a8ac378e04abbf8d5212

Observation a434e9a6-928c-4296-aba4-a533e097827c · inbound

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models cites this paper.

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-24T03:23:49.665779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-24T03:23:18.827351Z digest=sha256:c21be715355202862a083b1e5a8b69316e87d0748e899f883ff270df7447211c

Observation 8072bf01-20ef-4a54-8a5f-ed70093276c1 · inbound

Yi: Open Foundation Models by 01.AI cites this paper.

Yi: Open Foundation Models by 01.AI Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:47:27.911955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T05:47:27.775529Z digest=sha256:98ebe30e5d6fdaf4bad051d0ccb43ca45cdf957a4079e998451d8c8975a0e7f3

Observation 0b330107-2b0b-4a25-9099-5b686843acc6 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:de3fd4d07c2d0326bf9cc0565aa2eb0fedc14f88d2193ff776048cf756122463

Observation 63d22337-1b03-41f1-8275-a38837e26b52 · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 100

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:ec91c4c4b1096f04cba33fc63a7d9b9e0074792252d777d873a6eddaa96e0c85

Observation 049ef5fc-793f-4d5f-b474-063a1f4ffff8 · inbound

Lessons from the Trenches on Reproducible Evaluation of Language Models cites this paper.

Lessons from the Trenches on Reproducible Evaluation of Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 122

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T18:44:49.770643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T18:44:49.519995Z digest=sha256:546d6e61b923f3f6393712329330a8846c8435d953187485b595931b285a17d4

Observation 825355f2-d67c-41ec-a68a-825ff5a528a9 · inbound

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark cites this paper.

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T15:51:05.307004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T15:51:04.674346Z digest=sha256:0d5b0fbd72d8f3aefe21634ee5f4ae0a1f2960d2f4eebb2a02ad8d2a243d1e9e

Observation 0d258557-47dd-4de1-85a6-5c5b69fa9df4 · inbound

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence cites this paper.

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-16T01:06:07.910156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T01:06:07.787696Z digest=sha256:581a294689aa5c2ac24b27168791714a12f19f02ad46ceebdf17df185677cf9d

Observation de55283b-4306-4471-b9bd-87a58a89dd0a · inbound

LiveBench: A Challenging, Contamination-Limited LLM Benchmark cites this paper.

LiveBench: A Challenging, Contamination-Limited LLM Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-15T04:48:26.500818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T04:48:26.303240Z digest=sha256:6107cfb3f51e4b25fec56372b2617379b75fb381aa9cfa48472bc872cf3a0885

Observation 07bff1ea-da89-472d-ad0b-613b07be387b · inbound

Reinforcement Learning for LLM Post-Training: A Survey cites this paper.

Reinforcement Learning for LLM Post-Training: A Survey Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-23T22:35:50.299816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T22:35:35.287039Z digest=sha256:1c8ad7b46bfa0f058370a57b88716de6c67b730486d2f72c2829988002accb57

Observation c50530f3-288b-4c2e-a742-ac28c1f143b2 · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 271

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:38:37.108197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:e0c0f4ce8f811f27cdc5989f1fa604b5e6fbbadd16f91ca46623a1608645ddd2

Observation 60b36141-770e-44b0-8135-0e137ef45a1e · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T06:20:36.319194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:8c6b89572f31efa2e342da57868269459f77475bd5094bb0064de4bd1e997e15

Observation 42c61861-9c70-4723-982d-d25e9f19a944 · inbound

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model cites this paper.

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:14:04.976702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:14:04.976702Z digest=sha256:1f944cdefe54b4e83669215ab6a2816b1cea306ae05af72a989187d4de9f5185

Observation 6cfb8def-264d-444d-84a0-a691b3a8f32a · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:16:17.448881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:e1d1506aab698f0ce23a3c3e126fc1ab414d951a50798b0b22df41fe0f771fd2

Observation 2798ba0c-1687-4752-802c-a3c2dcf5fba7 · inbound

Ultra-Sparse Memory Network cites this paper.

Ultra-Sparse Memory Network Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T17:40:11.858088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:40:11.858088Z digest=sha256:dab90d56c7c6359e25a8630b02fc8a66b93c7446156d7dd2422e26dfa8711f2e

Observation 42aa4d11-1f1f-458a-b147-2bdf34e8ff59 · inbound

SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference cites this paper.

SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T17:20:31.998935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:20:31.998935Z digest=sha256:fa03a8e164c5b3be64907a689959fb6078daf387928b42cb5f1a7ed103e85b4c

Observation abd24fd7-f66e-41d7-8f77-b9a025491119 · inbound

A Survey on Human-Centric LLMs cites this paper.

A Survey on Human-Centric LLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:42:07.803258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:42:07.803258Z digest=sha256:5b2db4f6cfe20ba47215673b1bcf4c80900facf6b30494898e8fb0eba1443364

Observation 4e2ee037-24a5-499c-8b7d-106542d005ee · inbound

INTELLECT-1 Technical Report cites this paper.

INTELLECT-1 Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:44:26.595005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:44:26.595005Z digest=sha256:973af0457e7b661fe2007619d8727bbeb890f83e1be8f8655b9db502d9d8b27e

Observation d27fd8c6-5f55-4951-986c-35a6f51d64c6 · inbound

REVOLVE: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization cites this paper.

REVOLVE: Optimizing AI Systems by Tracking Response Evolution in Textual Optimization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:52:15.171771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:52:15.171771Z digest=sha256:1c698fc646c8b89ddaeffc22829ad85cd960450ea7d5fafe7f996c4eee3d4f5f

Observation 4a1352d5-c874-4066-9051-63d8d304021d · inbound

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation cites this paper.

Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T22:37:56.294266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:37:56.294266Z digest=sha256:417c2eaf506592e12614bd9ca54cd06ca08405866c6c2e97421ee887125f5b4e

Observation da0a4435-febc-4a1d-9d56-ecac1858699a · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 225

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:13f9d6998d740f02b5b6813c72e2f84f095550f960ff2ae8854b11ce97816d75

Observation c18562d5-1b2c-48ca-959c-f8b2d384699c · inbound

Multi-Party Supervised Fine-tuning of Language Models for Multi-Party Dialogue Generation cites this paper.

Multi-Party Supervised Fine-tuning of Language Models for Multi-Party Dialogue Generation Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:14:43.304595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:14:43.304595Z digest=sha256:04e1a5265fc56d453ff96912abc72def95c3468a08affb361b4155847576ee91

Observation 26c89f66-151c-430b-b17b-827b4403e3d3 · inbound

Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need cites this paper.

Predictable Emergent Abilities of LLMs: Proxy Tasks Are All You Need Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:11:48.507975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:11:48.507975Z digest=sha256:62e6a079226485305639cfa167bd3ac2a691c70906001852bbc4b508c48dcef9

Observation 8447ce42-b0ea-44f2-adf6-2e3dd18b9bb8 · inbound

From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents cites this paper.

From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T18:27:26.435247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:27:26.435247Z digest=sha256:76b7b3935bf1efcf679368512fde42bc499fb7ee5c1dd64b351daa68693f5910

Observation 4c037a59-b852-4e65-be2d-3d774f1c1c0f · inbound

GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers cites this paper.

GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:52:28.882333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:52:28.882333Z digest=sha256:9756029754c70ceb6cf938d77a53e549730369f581e211aa2ff3f57c7e92010a

Observation 5b93d489-5b7d-40f2-9c71-449c20c378d5 · inbound

Codenames as a Benchmark for Large Language Models cites this paper.

Codenames as a Benchmark for Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:04:39.150171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:04:39.150171Z digest=sha256:fb5e110bde23c86b00151364536a4f59476b1127b17726a30fe309a594798a3d

Observation 21adfeca-1a8f-4787-92bc-2f1998ab195b · inbound

C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness cites this paper.

C3oT: Generating Shorter Chain-of-Thought without Compromising Effectiveness Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T14:47:03.996961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:47:03.996961Z digest=sha256:5577ef63ab4efbe9a4cd00d98086fa9b5fd2b38becd604ed0bb958d1fe5eb45c

Observation 63fe2f30-9075-4174-899b-f79055ea0f04 · inbound

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge cites this paper.

Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:26.362213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:04:26.362213Z digest=sha256:a38a0bb9e4b4c76bcf865849eb87802913d1c830726ab41bfd8f2abd92cdab47

Observation bb619b33-7315-48c9-baeb-8f9a99d883b3 · inbound

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference cites this paper.

Activating Distributed Visual Region within LLMs for Efficient and Effective Vision-Language Training and Inference Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T13:49:12.674400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:49:12.674400Z digest=sha256:8571a0bf0c4efd68aaa6cb6b6e25b94c6d1e5c935abfd7ac89714fcbade2f71e

Observation d04ef4d7-7c82-4939-a46a-3e5571cc9b40 · inbound

Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs cites this paper.

Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:19:19.021871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:19:19.021871Z digest=sha256:be1f351a0e0f4246f9906c33e8fa5d7bb627af355b9c785602237d79939a0d81

Observation 38d0ece0-5b9d-4acd-89ef-12140dc30d40 · inbound

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment cites this paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.980152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.980152Z digest=sha256:5e38d8fb5a4a96b918a10a0caca6aa54b89dca1c57e5cb1bd8098ba5f0d323a9

Observation 99eb43db-ffe6-4276-9c02-268a8ee5af72 · inbound

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis cites this paper.

ResoFilter: Fine-grained Synthetic Data Filtering for Large Language Models through Data-Parameter Resonance Analysis Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:59:28.415460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:59:28.415460Z digest=sha256:1de9cfc6af1bd20617f7800023ae9a47d988296c9484a0124b3f7b9b7e6cfb41

Observation cd0be1f3-7eab-4981-b200-e6298c3800e8 · inbound

Language Models as Continuous Self-Evolving Data Engineers cites this paper.

Language Models as Continuous Self-Evolving Data Engineers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:52.554445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:52.554445Z digest=sha256:9b0f253b8e8557b52e2b0f8fd01aa19a884beea49a6f9499169ac488eac2c074

Observation e92359ad-a340-49ab-85d3-99b5a7882fcc · inbound

Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying cites this paper.

Critical-Questions-of-Thought: Steering LLM reasoning with Argumentative Querying Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:23.437515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:23.437515Z digest=sha256:a981c842949473bcba344f848a6ce0c38a66222a469d8a3e7a1120625d9fe841

Observation 36bc1ab2-a1db-49a9-a90f-bcb34ab95f26 · inbound

Multi-matrix Factorization Attention cites this paper.

Multi-matrix Factorization Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T00:55:04.003033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:55:04.003033Z digest=sha256:7c85e02663624db85ecfd25616fc82242fff99f872353491e7c22eb8714afac9

Observation 88b4f048-4e97-4b13-b09a-0352066c6f09 · inbound

Aligning Large Language Models for Faithful Integrity Against Opposing Argument cites this paper.

Aligning Large Language Models for Faithful Integrity Against Opposing Argument Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T22:35:43.939107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:35:43.939107Z digest=sha256:1c844aba1777b95fa9899c0d71a331b7d0d588340f5c4eac32b88a15043ca415

Observation ae3590d3-ca72-456d-945f-488df1a5c8f7 · inbound

InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion cites this paper.

InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:35.194067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:08:35.194067Z digest=sha256:d5148b111d65d8b39d542a5c2d1f2797a9a22c83552c0b2ec2b62a78bd4599f5

Observation a7f41a0a-449a-4f73-bbb6-ff013a3286e5 · inbound

A Survey on Large Language Models with some Insights on their Capabilities and Limitations cites this paper.

A Survey on Large Language Models with some Insights on their Capabilities and Limitations Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 218

Resolution
unresolved
no resolver link, observed 2026-08-10T22:17:56.251776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:17:56.251776Z digest=sha256:77668ca0698678296893b2d60508c596077645aedbed182dee5447ab8d1b350a

Observation 2c19a6f9-f116-45e4-8275-9764fc426a7e · inbound

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark cites this paper.

Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T21:18:59.717663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:18:59.717663Z digest=sha256:b67894f4719f4c9e3f30146d1f49c971092a891c40a1a94bd377194c0f5b7155

Observation 0df76838-29dd-462a-9686-812e9eef82e6 · inbound

TAPO: Task-Referenced Adaptation for Prompt Optimization cites this paper.

TAPO: Task-Referenced Adaptation for Prompt Optimization Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T20:57:57.261868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:57:57.261868Z digest=sha256:cf2a43d2d7cc7ee2079556c8297c75f80f9c885fa5519d1466122e8328badcab

Observation 5461f440-0748-4495-9662-88d161f2d8f0 · inbound

ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair cites this paper.

ViBidirectionMT-Eval: Machine Translation for Vietnamese-Chinese and Vietnamese-Lao language pair Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:25:06.274531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:25:06.274531Z digest=sha256:b157301d33dda960a60d4f25eb61270f83173b662af8e8adf28471383076394f

Observation cfcacde6-05fc-4e24-8076-b7417dd6e3bf · inbound

Domain Adaptation of Foundation LLMs for e-Commerce cites this paper.

Domain Adaptation of Foundation LLMs for e-Commerce Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:29.704958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:48:29.704958Z digest=sha256:d9fd1018e7f39fd60f71d294fc5e4301279c7bd2865acdf52aedb8fa1a0ddef4

Observation 208f4aa8-0c4b-45f6-8b11-a4772b66bdd9 · inbound

DNA 1.0 Technical Report cites this paper.

DNA 1.0 Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:04:35.011026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:04:35.011026Z digest=sha256:5f296a5682063af06f785f7d078d0f812b1839c790a2e4ddadf3d2e3bb5665bd

Observation dbd223b4-5fec-4f2a-a188-f0e9c7955c2d · inbound

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? cites this paper.

RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems? Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T18:30:56.673473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:30:56.673473Z digest=sha256:9d99b2fa8917e2c76bcf3f634f860403b428a4816e385aa6f486226bbad3d3cc

Observation 17e1237a-0aaf-425e-955c-4d8a3b1f9292 · inbound

Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities cites this paper.

Improving Influence-based Instruction Tuning Data Selection for Balanced Learning of Diverse Capabilities Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T17:34:52.494040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:34:52.494040Z digest=sha256:c242f933e0e11c812475735e98e7edd8780e6a5d2dfe57a03cdb28ad106f5582

Observation 89aae155-c70d-49c6-a29d-b73f2fcc3f50 · inbound

Spurious Forgetting in Continual Learning of Language Models cites this paper.

Spurious Forgetting in Continual Learning of Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T16:13:49.393292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:13:49.393292Z digest=sha256:b890908a2a91dc44e4cb2136c3c9e5765ae28601a7c5f1d8f063ecbe09ad7f0c

Observation be5bb4b6-fd4f-4e69-9f89-70d457fa39e0 · inbound

Adaptive Testing for LLM-Based Applications: A Diversity-based Approach cites this paper.

Adaptive Testing for LLM-Based Applications: A Diversity-based Approach Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T15:57:28.884092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:57:28.884092Z digest=sha256:595502a037470644888b003d48d265caa3a38df08228a9763fd4b70a76380c07

Observation 8fb2f21f-e0ef-42ee-8ea8-e223838bbc8b · inbound

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models cites this paper.

Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T15:50:21.983664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:50:21.983664Z digest=sha256:1896303c753662bafd0fdd2942d149b7cad3fb6f46a7888ab6c46b25658a59fd

Observation cae2543c-5c9c-464f-9500-08e524451f5c · inbound

On the Reasoning Capacity of AI Models and How to Quantify It cites this paper.

On the Reasoning Capacity of AI Models and How to Quantify It Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T15:38:37.480762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:38:37.480762Z digest=sha256:24fbe88b1d7358102ee546384980338a0b5ec3883796ae2d7bed0c22e19d1580

Observation 6f12650d-6bb5-41b9-b0a4-f2d72c38eb7e · inbound

StaICC: Standardized Evaluation for Classification Task in In-context Learning cites this paper.

StaICC: Standardized Evaluation for Classification Task in In-context Learning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-10T14:06:14.467857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:06:14.467857Z digest=sha256:d60edd28c0a328476a1cc91dfe47de6d2cb4f3f67cdd345a0f22159b54087073

Observation 91f144c8-4bea-4f08-9802-e3d79dc3f769 · inbound

LLM-AutoDiff: Auto-Differentiate Any LLM Workflow cites this paper.

LLM-AutoDiff: Auto-Differentiate Any LLM Workflow Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T11:33:24.669461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T11:33:24.669461Z digest=sha256:c6dbdcc74895b49745d9d1040518faa98fa643a909640e9b68a874c2a104c91a

Observation 45c70adb-9721-44a4-a224-80e0b12dad0b · inbound

BTS: Harmonizing Specialized Experts into a Generalist LLM cites this paper.

BTS: Harmonizing Specialized Experts into a Generalist LLM Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T21:59:01.305357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T21:59:01.305357Z digest=sha256:385c85ebb818f69b2fdc916ebb3c2bd0aac8408af9bb5b7d7c8d93e8bf30faff

Observation 13ff5323-6dbe-4dc2-9d6f-bc7aaffc4c21 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 131

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:23:31.037818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:4729400dedccd13cdb089d03ee45d44cb9c1717452e676a19c3b65e8512ddf48

Observation 1568f826-7b6a-4846-b1cf-8bc21689a904 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-23T04:17:31.061104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:78f7352cc7e877bb7377fa7ad93940f439f9b2cc6384ac8564f104ee22195fe9

Observation 2ff57f1f-cd73-4af4-902b-9fd4fde3f128 · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 110

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T17:30:02.885190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:07bf9373e175b44c2a4e5ce473337688da12d10d57b1dd586dc34821fbe90c91

Observation 5c3ee4f6-2317-4594-92e0-52fca4eafe53 · inbound

Minerva: A Programmable Memory Test Benchmark for Language Models cites this paper.

Minerva: A Programmable Memory Test Benchmark for Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T05:02:35.797402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:02:35.797402Z digest=sha256:e0b6029fac7685c428a625e9b03e4710946434004cc702f518894823474e95c3

Observation cecd3773-fd12-44e2-90f6-cbc869d0af81 · inbound

It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers cites this paper.

It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T00:49:12.213813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:49:12.213813Z digest=sha256:37bb5b16244ad71d2d5c19fe0f453ee4d794548ad45bd74c99e267ae9ca32553

Observation b9b77044-99b1-45fe-9c5d-727777069825 · inbound

CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance cites this paper.

CodeSteer: Symbolic-Augmented Language Models via Code/Text Guidance Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:16:27.416734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:16:27.416734Z digest=sha256:fb6d5830f39274ee0814396040fd6c9007f6848656c73b8eeba5c1f21c3162ae

Observation 00f716b3-e7bf-4484-b15f-2057fbb985da · inbound

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering cites this paper.

Automatic Evaluation of Healthcare LLMs Beyond Question-Answering Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T14:46:29.688673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:46:29.688673Z digest=sha256:33cbdf2f634fd21d3c81de6b09227720dc2093d7b05f429e16c1da4d6c99cf80

Observation a928af93-6277-4e58-a27c-acd6cafd6f1f · inbound

KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems cites this paper.

KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:08:22.864918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:08:22.864918Z digest=sha256:10a6633680b0e0cedae2f53fbe1cbbca1e0385235d2877ffad66b770924fc44f

Observation f2c5ac7a-a6e2-4931-a46e-a312c619e32d · inbound

CryptoX : Compositional Reasoning Evaluation of Large Language Models cites this paper.

CryptoX : Compositional Reasoning Evaluation of Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T18:34:26.786749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:34:26.786749Z digest=sha256:fea51d7680c060180ee71dab39487c9611ca5f528f8567a270d42e7b1f85580e

Observation 83636a88-52f1-4891-b66d-98e554e2fb56 · inbound

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters cites this paper.

SHARP: Accelerating Language Model Inference by SHaring Adjacent layers with Recovery Parameters Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-08T13:44:00.641561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:44:00.641561Z digest=sha256:2a66da02f75a1136356c11896a71cba85439592579eeedb68ba978ec283bb09a

Observation 6921b2ba-1521-4e3d-8b4b-623878249861 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:34.389582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:34.389582Z digest=sha256:9b862a4a81c82ca804dfa885d6fe863b686bbc258a43480c7990971b6cd82adc

Observation 45500c75-d4ab-429d-81d9-7d08780ac842 · inbound

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models cites this paper.

RoSTE: An Efficient Quantization-Aware Supervised Fine-Tuning Approach for Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:03:44.662978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:03:44.662978Z digest=sha256:b56dd06df79cdcbcf16f04d6fe5ecb305927c5022e84cc43d2f016b773056a9b

Observation a4e11429-2c0e-4206-8140-eadeb4db496b · inbound

Large Language Diffusion Models cites this paper.

Large Language Diffusion Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T01:42:54.279353Z digest=sha256:5395a406e837d0b77d0068fceffdf2a585a1fcb5e934c7d04056a146e68e8c88

Observation f88203ad-f0c2-4061-94a4-2b92f6933b7b · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 239

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.630557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:34fd6daf308adad450bc09f2789d7bc2fc57683b841d55f90045e0dd44fd37fb

Observation a083efec-094a-4cb1-925f-85ccbfbf1517 · inbound

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention cites this paper.

Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T23:46:30.126324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-16T23:46:29.975858Z digest=sha256:821ac8b3638ea4626740bcc0c602fac3122f46b789008b170b1209a1c22f3c49

Observation cfedbd3e-8392-472e-8469-63f2cc848781 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:15:24.133279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:297c647f81c0194980ee985e88ab86b569f2d9f5a3b54eda9302fd54e0c60c5e

Observation 7867c7a8-2bc5-40d9-92e9-293fb68fe583 · inbound

MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space cites this paper.

MIG: Automatic Data Selection for Instruction Tuning by Maximizing Information Gain in Semantic Space Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T12:06:23.497019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:06:23.497019Z digest=sha256:3a8549de4b4e5095f4432b75310fab7e4d3273cf92aa7ce313c394a1db72b936

Observation 0961d143-6f64-4c50-985e-aec82a16c75b · inbound

FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models cites this paper.

FarsEval-PKBETS: A new diverse benchmark for evaluating Persian large language models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T11:47:09.996193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:47:09.996193Z digest=sha256:eaae886ff4e3070bf5e8347e44ddefe24605fbea3e84178d85ee6ee593361319

Observation 1652c657-4c63-43ea-8e8d-05ef74895200 · inbound

Trillion 7B Technical Report cites this paper.

Trillion 7B Technical Report Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:06.079010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:32:06.079010Z digest=sha256:f8356f3930d980c39fdc35eea5489085a55fee9f906642aa5bf9bbf28390f5e7

Observation 7d90395e-ccc4-4a76-8e28-8fe87d62527f · inbound

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism cites this paper.

Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:17:29.443099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:17:29.443099Z digest=sha256:73c3ad556177d21391a6a928dee14906e5d6ba08d95ecf2a9f8572f6fbac5a67

Observation 187cc37e-634a-4aa6-92e8-8c403a0c25f0 · inbound

CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges cites this paper.

CipherBank: Exploring the Boundary of LLM Reasoning Capabilities through Cryptography Challenges Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T06:05:12.114299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T06:05:12.114299Z digest=sha256:1550a0983f88d7ff87d6f4ef58842e0342a8b4b385a62a02eda979f0ef881ee6

Observation 6b2eb093-5e1e-4e85-aee7-f00ab4adabc8 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 131

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:38.501204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:7b7a36a047ede3b00fd17ff358ac2e64c2830b3da84bbc123f7a19e85b987ad5

Observation 050e536f-97e7-4336-ac09-e6c678f15308 · inbound

Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection cites this paper.

Combatting Dimensional Collapse in LLM Pre-Training Data via Diversified File Selection Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T05:30:51.296395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:30:51.296395Z digest=sha256:dee222935410f4f122e06dbd8fefe75b03493faa50dd05da23db2b7e6a9f007d

Observation 7e6519e7-9c1c-4eed-80db-d4a23fefdf07 · inbound

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models cites this paper.

100 Days After DeepSeek-R1: A Survey on Replication Studies and More Directions for Reasoning Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-16T04:43:44.486624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:43:44.486624Z digest=sha256:3fd968e270277e6aae6b68d3c20483c34b1576a3c6f9882b78882db0afbc9733