Pith. sign in

Paper Citation Record · LEDGER

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

As of 22 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 1 inbound Pith citation observation for arXiv:2607.13248.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13248 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:48:58.571640Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:31:24.935640Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:31:24.994128Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved88
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bce190b-adc0-4c98-a513-006763232a2e · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:47.190993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:47.190993Z digest=sha256:8085e47e46753ed5cab2d8002de684fb7d5126f6416aa5e9ad7d0b9c2cf5463b

Observation 63d763d1-33f7-46c5-843d-8f4f2530a870 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:47.276280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:47.276280Z digest=sha256:53e71427262c9a489356d775f8b2e82f27012ff04aea4892f7bec4d510a3013f

Observation 40f49a4d-4e4f-4ead-97be-60958143c626 · outbound

This paper cites Language models are few-shot learners.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:47.439348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:47.439348Z digest=sha256:fcd6468030fb3cb4e16ca00c41b21b610824d2b9febb5e63424c2fe446d097ce

Observation ab9e9de5-1d85-482b-be45-783fb80fc3e1 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models PaLM: Scaling Language Modeling with Pathways

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:47.604661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:47.604661Z digest=sha256:998aea8b1ab26a120d78ff9dc0096067044f7bacdeff6d448efaec63182ff637

Observation 4fbfb4ab-2f67-4b35-8a1e-54b9f7795e5b · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:47.775905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:47.775905Z digest=sha256:4edcd1adfd348b158342a2f60731f38b818f643b09c39f17adb0011e9f783191

Observation 2be025c6-5f38-4f05-9030-348c7d9a48ce · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Chain-of-thought prompting elicits reasoning in large language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:47.909087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:47.909087Z digest=sha256:bb58563691b9e5d1f0bd85ef5406b04e25e9f1403dc7cd93baddc2ecd799e9be

Observation e2530934-3214-4a28-ac35-0782901d0a2d · outbound

This paper cites Lampinen, Ishita Dasgupta, Stephanie C.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Lampinen, Ishita Dasgupta, Stephanie C

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:48.095647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:48.095647Z digest=sha256:1ec4556b6f6bd3738cb142333748079fef6772dfc678a58022a4995c793cba90

Observation 9a0cef49-b2c0-4af0-b46a-7ce90c5869e8 · outbound

This paper cites Evaluating mathematical reasoning of large language models: A focus on error identification and correction.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Evaluating mathematical reasoning of large language models: A focus on error identification and correction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:48.305628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:48.305628Z digest=sha256:5210bac5044336e0b46f306cf318019c5b5e74ecd71f7f523a8f393f75146ab4

Observation 4592e1e0-2d16-49d4-abe5-723a21cc7b1e · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:48.409446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:48.409446Z digest=sha256:1a8f32e92af88920c4e40e47e492b8c383979c79e2e0808ae370be43543e9243

Observation a6df44a5-3b0d-4201-96e5-ec60be9f6360 · outbound

This paper cites Mathify: Evaluating large language models on mathematical problem solving tasks, 2024.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Mathify: Evaluating large language models on mathematical problem solving tasks, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:48.610512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:48.610512Z digest=sha256:373fb03716ea81470a4d2ab483f353485c44d026ffcc215c4217cecafa091ef0

Observation 1c9328e9-4903-4205-b70b-77bec0f5b6ea · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:48.762081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:48.762081Z digest=sha256:6e38f4c434236d9e020ee75be501ed4ac89a5cea2be0a1937e05c5dfb0ee6629

Observation 83a612ed-45d0-4109-9621-4a37230ec241 · outbound

This paper cites Benchmarking in-context learning strategies of large lan- guage models for math reasoning tasks.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking in-context learning strategies of large lan- guage models for math reasoning tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:48.942078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:48.942078Z digest=sha256:3bd4a00f9e691a9b258c5cd79a38c30c8bb0d9b382615186f03aaf265662c7ea

Observation 3e52db02-361b-4445-a234-4321c9799b42 · outbound

This paper cites Omni-math: A universal olympiad level mathematic benchmark for large language models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Omni-math: A universal olympiad level mathematic benchmark for large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.107551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.107551Z digest=sha256:d0c54d7fc559907d14df5a723f02c82d3f2eedf5c15196edc24eba23961380c1

Observation 6036ef5c-4d9f-4230-892a-109d6ee7d315 · outbound

This paper cites Benchmarking Large Language Models for Math Reasoning Tasks.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking Large Language Models for Math Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.307485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.307485Z digest=sha256:f24be62861e20a7a45caae4bf4612a063d97acc7e2a7e14bb9ff81de51e46731

Observation 4756a493-e589-482f-bdb1-fd70f81a5178 · outbound

This paper cites Lila: A unified benchmark for mathematical reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Lila: A unified benchmark for mathematical reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.448507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.448507Z digest=sha256:293e4207760562325bbca4d7da02cea74b55393a1a08ea4e758759a0d11df48b

Observation eaf2dbb9-3575-4377-a92a-3ce96e527b01 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.589809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.589809Z digest=sha256:0fbdd3c0c5da2451158fc3e40b279fe9234fa46a76ea92233517399f637587e0

Observation 8e901685-b5f2-441e-9668-87c3b1e7cc58 · outbound

This paper cites an unresolved cited work.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.748099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.748099Z digest=sha256:80924a4725df55a75c197ec886d84fe1ee5cdb6f854a22342080131aa27516ea

Observation 73f688c5-7310-4a74-a238-526dc0661efe · outbound

This paper cites Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:49.892122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:49.892122Z digest=sha256:d09ddd512ec7787fe9dce6713fa606eca790e720af15b9347d98e74401da1196

Observation 26c5fce6-9cdd-450f-8990-06621c28cd78 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Measuring Massive Multitask Language Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.008979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.008979Z digest=sha256:ceb6021c1d313c8a0a7c495c123691f22517e325b9790c84e484dd0ebd1229f3

Observation c5525a8a-1a30-4f71-aa23-f6427d7e39c3 · outbound

This paper cites Math-perturb: Benchmarking llms’ math reasoning abilities against hard perturbations, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Math-perturb: Benchmarking llms’ math reasoning abilities against hard perturbations, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.166797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.166797Z digest=sha256:8ea6a01fea01223a1fe77554e2051954f60337a9248dddafc95a469113bad5a7

Observation 0a3f7f6f-6947-4bae-ad36-5ddd12773e8d · outbound

This paper cites Dynamath: A dynamic vi- sual benchmark for evaluating mathematical reasoning robustness of vision language models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Dynamath: A dynamic vi- sual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.315864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.315864Z digest=sha256:f0bb47d262234d6e5717eccfaca2740fa0f3f840d031a2a4dc68d07631cff081

Observation 7dee0196-0d50-4c69-8992-8c15a07ff5f6 · outbound

This paper cites GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models GSM-Plus: A Comprehensive Benchmark for Evaluating the Robustness of LLMs as Mathematical Problem Solvers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.502614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.502614Z digest=sha256:edb02133d33c49b6c52f8ef4d7585270450ceed407f29f687a9ec28d80081be7

Observation 505620ff-c7c7-4194-8c7d-3e98fbb3765e · outbound

This paper cites Evaluating robustness of llms to numerical variations in mathematical reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Evaluating robustness of llms to numerical variations in mathematical reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.680645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.680645Z digest=sha256:0a273e17dab2c413a639c5f4e98119595ed3448f97d7b2dd32431a2da1f25303

Observation 2999396e-1f59-4047-8fbe-f9e154466d78 · outbound

This paper cites An investigation of robustness of llms in mathematical reasoning: Benchmarking with mathematically-equivalent transformation of advanced mathematical problems, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models An investigation of robustness of llms in mathematical reasoning: Benchmarking with mathematically-equivalent transformation of advanced mathematical problems, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.802411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.802411Z digest=sha256:cb81d6c628ca243d43a4065d8067a9cdf486cba8660b37b0f21fc9daf6f0f164

Observation 7b2decfb-6133-4160-b4af-e2b58ff08b91 · outbound

This paper cites Robustness in large language models: A survey of mitigation strategies and evaluation metrics.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Robustness in large language models: A survey of mitigation strategies and evaluation metrics

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:50.943635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:50.943635Z digest=sha256:bcfcedf68d0516e92ac3b7574a209771bda30fa3cb55e168b24664546889b331

Observation 4f0245c0-cc4b-4706-923f-d4535319a498 · outbound

This paper cites A novel metric for measuring the ro- bustness of large language models in non-adversarial scenarios.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models A novel metric for measuring the ro- bustness of large language models in non-adversarial scenarios

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.056352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.056352Z digest=sha256:f2f38a0b48de246cf63ddb6636766d6af0d70135a9c648393bb1fd111cdc2054

Observation a94dbcb5-9fe2-4e1a-a596-8c2c6fdf9920 · outbound

This paper cites Do large language models understand their knowledge? AIChE Journal , 71(3):e18661, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Do large language models understand their knowledge? AIChE Journal , 71(3):e18661, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.233717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.233717Z digest=sha256:7f6481c4d0d3410b7399336ddf73641700ab5f4c1c603416b170db8543132672

Observation c89bc803-1387-448f-ad87-9b1b577e8b9d · outbound

This paper cites Is your model really a good math reasoner? evaluating mathematical reasoning with checklist.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Is your model really a good math reasoner? evaluating mathematical reasoning with checklist

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.340082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.340082Z digest=sha256:d707079760d905450d95a72b48dcd4a73c9773017d8f8d86c7aa4e879a382b81

Observation 95e1fa83-4c96-45be-980c-41b8312ffafb · outbound

This paper cites Bhaasha, bhas.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Bhaasha, bhas

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.474683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.474683Z digest=sha256:8ec80904892a93c75e18d69603dd632eab6825f5b98e13a95eb5e1a174d8d4c5

Observation abea8591-ebd7-4900-8bf2-9265e2288696 · outbound

This paper cites Llms for low resource lan- guages in multilingual, multimodal and dialectal settings.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Llms for low resource lan- guages in multilingual, multimodal and dialectal settings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.573597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.573597Z digest=sha256:c2e743b235cfdd9b96febd1a31c4669b8c9d05a536715ecaefa204ca1c7b8b48

Observation 31f71879-3b2c-4193-a966-5ad0300d72b5 · outbound

This paper cites Natural language processing applications for low-resource languages.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Natural language processing applications for low-resource languages

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.657980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.657980Z digest=sha256:c0845fb2d5f2700977dcc94f58d0e8b9d7a0a083cc9558655a8d4b0d44a9a04e

Observation 204bd9ac-a1ea-499e-85f5-2e5f9d1a7407 · outbound

This paper cites Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Linguistic Nepotism: Trading-off Quality for Language Preference in Multilingual RAG

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.805431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.805431Z digest=sha256:a979d133dd296d5ece077f691ce73677c9bc362e6631d0f5cf414bd5668fbafe

Observation b7fc5b32-edec-4ab3-8a22-ac84241fb9fe · outbound

This paper cites Bengali language — Wikipedia, The Free Encyclopedia.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Bengali language — Wikipedia, The Free Encyclopedia

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:51.946286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:51.946286Z digest=sha256:f722b3dc4f0ec7aa587d08d78f8ddf7e3f48cf56e3a62527de36222eebf064e5

Observation 25e33802-b0c3-4150-9b8b-0345c0334975 · outbound

This paper cites Geospatial and Temporal Trends in Urban Transportation: A Study of NYC Taxis and Pathao Food Deliveries.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Geospatial and Temporal Trends in Urban Transportation: A Study of NYC Taxis and Pathao Food Deliveries

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.098465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.098465Z digest=sha256:b1c8afa27a98368964821926dab115a5c048735af034f163947b952d9ec8dab2

Observation 466d3e7a-5760-4936-8948-056b027cddcf · outbound

This paper cites Banglamath: A bangla benchmark dataset for testing llm mathematical reasoning at grades 6, 7, and 8.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Banglamath: A bangla benchmark dataset for testing llm mathematical reasoning at grades 6, 7, and 8

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.247263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.247263Z digest=sha256:238d5d6dcdc43de37d62046cdab8d389b707a1f09f19b5bc693cad720b6d648f

Observation c627aa7e-3546-45aa-901f-aea399da88e1 · outbound

This paper cites Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Leveraging Large Language Models for Bengali Math Word Problem Solving with Chain of Thought Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.372383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.372383Z digest=sha256:18b18f9e7823705a1a0684a28113691d95fdb71e63a0fb9a6aaedb569fe3a3c1

Observation e5717357-f1b7-4d2d-a034-19802bbad25c · outbound

This paper cites Ganitllm: Difficulty-aware bengali mathematical reasoning through curriculum-grpo.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Ganitllm: Difficulty-aware bengali mathematical reasoning through curriculum-grpo

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.494430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.494430Z digest=sha256:b3c44d5d6435eca32ea9bec0e268fccf50361ae6a89aa678c05aa5fe696981b8

Observation 6ff1017c-d4e7-49e5-bc00-a778c80bfa1c · outbound

This paper cites BanglaBERT: A large-scale language model for bangla.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models BanglaBERT: A large-scale language model for bangla

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.601585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.601585Z digest=sha256:a1f41b108f6e085be5c342f78ab8dff19a73d4cd8ec6efd67514e8ad67bc1e21

Observation aec5dda4-e2af-4c5a-adb5-69f01d517f08 · outbound

This paper cites BanglaBERT-Base: A smaller model for bangla NLP.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models BanglaBERT-Base: A smaller model for bangla NLP

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.691619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.691619Z digest=sha256:e62215d9e3d754a48ee3edc4000d7be57f51293ff47fdbbfde25267ee2ee9a88

Observation f6fd5a36-933e-419a-acb4-4ebadd4d0d1b · outbound

This paper cites Xtreme: A mas- sively multilingual multi-task benchmark for evaluating cross-lingual generalisation.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Xtreme: A mas- sively multilingual multi-task benchmark for evaluating cross-lingual generalisation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.815611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.815611Z digest=sha256:fd782cee530aaa231c38e7afe1cec1b404027380adb0cc168b1e9c3bc0cd0ee0

Observation c8780b91-41bd-4b4e-9fac-39d5bbb6ec40 · outbound

This paper cites XGLUE: A new benchmark dataset for cross-lingual knowledge graph construction.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models XGLUE: A new benchmark dataset for cross-lingual knowledge graph construction

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:52.927498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:52.927498Z digest=sha256:3142622a9c44a2fa0e7e4c33f9732e9fe9a2e6ccf96e2400865d29387da9ecd2

Observation c0b752c0-e67a-495f-9ef4-7e9d299b519e · outbound

This paper cites Breaking language barriers in mul- tilingual mathematical reasoning: Insights and observations.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Breaking language barriers in mul- tilingual mathematical reasoning: Insights and observations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.039838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.039838Z digest=sha256:f9a69f19d00ae1f0f2274ef69fc8ca112ebe901367655d20b1d8b64d3d6b04dc

Observation 1db08267-cdb3-492d-946b-0edbf9ccef9c · outbound

This paper cites BanglaNER: A named entity recognition dataset for bangla.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models BanglaNER: A named entity recognition dataset for bangla

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.140170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.140170Z digest=sha256:0ab87211261d5ecca08b9463f5c77211b1d16054a8fe213ba8a5feec4b8e1baa

Observation 5a3a1306-8684-4f24-a26b-39bdd8275269 · outbound

This paper cites Bnmmlu: Measuring massive multitask language understanding in bengali.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Bnmmlu: Measuring massive multitask language understanding in bengali

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.271273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.271273Z digest=sha256:8f41cdb71703627d8f49315345e7c897983e42bac1a5ff7f34089fbcc97977a7

Observation c7b9c5bf-13fa-4790-9cc5-bfd9034e4696 · outbound

This paper cites MathQA: A dataset for mathematical question answering.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models MathQA: A dataset for mathematical question answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.433383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.433383Z digest=sha256:ff98f25fa5da03dc901a67532d17b5a2e17869467015bef3e196f09efcc64ca5

Observation 1833ce5b-fdc2-49d4-afc1-004c6d66a705 · outbound

This paper cites Qwen3 Technical Report.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Qwen3 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.614538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.614538Z digest=sha256:eca1b06a5d2423ee643f7bfbfd94714cf30d09144da33e67a72d5536cccac1f6

Observation 89f4b833-b3dc-45e4-a54f-21399982ed24 · outbound

This paper cites Qwen Technical Report.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Qwen Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.804306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.804306Z digest=sha256:bbc0233ee26036ecd3eda7d0b5cff8fdfd33f7ac54126a751e5b949f35546f3e

Observation aa08cc3f-3280-48d9-92f1-a344a9f85664 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:53.921488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:53.921488Z digest=sha256:0916a253bf10801ff395371b061938041909eb20d830f19713f3bcdce7ac501d

Observation 8fc7011d-9030-4487-95c9-fda1fd4036b3 · outbound

This paper cites The Llama 3 Herd of Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models The Llama 3 Herd of Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.030331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.030331Z digest=sha256:f1dc049ebfd6cb987f3ca69bae51467bab4eefac7f56524f05a277c7d6ea60d4

Observation 5c2d37a6-77b7-4512-a6d7-854d1ce5014d · outbound

This paper cites Evolution of meta’s llama models and parameter- efficient fine-tuning of large language models: a survey.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Evolution of meta’s llama models and parameter- efficient fine-tuning of large language models: a survey

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.156413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.156413Z digest=sha256:c781f6917951c4056f416f0216d9085fe5bb95df628ab895235290462b2929b9

Observation 29b098e4-9e7f-409a-848b-759e17fa518c · outbound

This paper cites gpt-oss-120b & gpt-oss-20b Model Card.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models gpt-oss-120b & gpt-oss-20b Model Card

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.250857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.250857Z digest=sha256:570f9af9e451d0d92a1f6c844a2423d6dea5fadc45ef818c18788fd4e073b6ad

Observation dcbbdc53-86c7-47ba-b0fe-0d96a133baac · outbound

This paper cites Large language models are zero-shot reasoners, 2023.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Large language models are zero-shot reasoners, 2023

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.358479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.358479Z digest=sha256:ee60562c203bda35ddf07a27bf7fa325aa404d4085764c2b873d61180acea23d

Observation e7e60d2b-25e7-4dac-bffe-9c7334449653 · outbound

This paper cites Towards understanding chain-of-thought prompting: An empirical study of what matters, 2023.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Towards understanding chain-of-thought prompting: An empirical study of what matters, 2023

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.413479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.413479Z digest=sha256:f3e0fc778ac7d0c1b614f902b1bc3921cd4c8d0e6d68149dbcd14ab835279e69

Observation d4794586-3015-46eb-a7de-b8c97662d436 · outbound

This paper cites Large language models are zero-shot reasoners.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Large language models are zero-shot reasoners

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.514102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.514102Z digest=sha256:210c7f0f643eaad2e387c5828d5a4d16ae0917f03c1031d376b744c7b22be752

Observation 9a0da260-5d7c-42e6-80ad-974eadc0fe48 · outbound

This paper cites MathPrompter: Mathematical Reasoning using Large Language Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models MathPrompter: Mathematical Reasoning using Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.586744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.586744Z digest=sha256:6c82d114c216a274cb29a718713938431a781b9209941f47be5fece29d645fec

Observation 0663d788-710c-4807-b43b-8b3a30ab33d1 · outbound

This paper cites A survey on large language models for mathematical reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models A survey on large language models for mathematical reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.655414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.655414Z digest=sha256:39b2689775edb291baee9b91644d7ca18633cbcc488e37ad8f499b902fcfae1a

Observation d8f98cd7-2f4f-425f-a7e5-a2c0c4012ec7 · outbound

This paper cites A survey on large language models for mathematical reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models A survey on large language models for mathematical reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.739341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.739341Z digest=sha256:33f498cd51ae7b9037cd359713b495dc5596997a5643e222e179ea781d3df927

Observation dac39ad0-9a77-4e66-9591-7c61ce3a4cb3 · outbound

This paper cites Bennumeval: A benchmark to assess llms numerical reasoning capabilities in bengali.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Bennumeval: A benchmark to assess llms numerical reasoning capabilities in bengali

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.827163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.827163Z digest=sha256:8f0577eb3d546dfe1c77265ad21a1c0a7f6ed564cc43bc0fdbb521d4fb52abfb

Observation 954cd7f3-e902-4cef-8814-c3afe9bdf5a0 · outbound

This paper cites Empowering bengali education with ai: Solving bengali math word problems through transformer models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Empowering bengali education with ai: Solving bengali math word problems through transformer models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.884850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.884850Z digest=sha256:3d89af15ca1872ab3e9da7485c1c5b20bc7b9b3b0623d550f126b13ca4ec130c

Observation f7beae9b-ee7b-449e-8594-0779c3c0d513 · outbound

This paper cites Bmwp: the first bengali math word problems dataset for operation prediction and solving.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Bmwp: the first bengali math word problems dataset for operation prediction and solving

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:54.948646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:54.948646Z digest=sha256:752fd8dacca720a0587e6ac71a216a310220152a04e9588298a876e2fde04a09

Observation 89d8a220-55fa-4335-bcaf-b76e9ae94316 · outbound

This paper cites Mathmist: A parallel multilingual benchmark dataset for mathematical problem solving and reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Mathmist: A parallel multilingual benchmark dataset for mathematical problem solving and reasoning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.047689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.047689Z digest=sha256:23f8e4a56dc889e8e561d6993f98f17ca91367ee345fda57afdfd47cf7aac537

Observation 928c4702-3eef-4b53-ad6b-39053625fffe · outbound

This paper cites Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Bangla-Bayanno: A 52K-Pair Bengali Visual Question Answering Dataset with LLM-Assisted Translation Refinement

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.126526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.126526Z digest=sha256:860f4e7166886b1652cb236a0fcd282e7a6f541f1d5e7f71cf29ea3d50e695f5

Observation acdf606c-3787-41c2-819a-35d0beddf705 · outbound

This paper cites Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Mathify: Evaluating Large Language Models on Mathematical Problem Solving Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.261125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.261125Z digest=sha256:1fd0677169c4453eaf227dc84120d0b97e10f672086a893971d9d2b6b79fc435

Observation 827760a3-26cf-493a-a2ca-eebe1f86bf14 · outbound

This paper cites Benchmarking Reasoning Robustness in Large Language Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.417781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.417781Z digest=sha256:37150b361161da5483da0cde46ec5a206961b5d9aba9930ec93a9a24d6f9aa2d

Observation 82436922-1446-4743-a5ef-c6f0eaed9ab3 · outbound

This paper cites MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.512315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.512315Z digest=sha256:bf97a72c57f1205eae2b2e50bbdcd6dd2be1f87cce784d7d2fb0e7f9e7fe47d2

Observation 39a75135-f0dc-4add-9c6c-3f4da563e6f2 · outbound

This paper cites FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.669532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.669532Z digest=sha256:77a753caab40a7e7dbe03803cf958c21bd0c9161c6da80e371482179fc31e2f3

Observation 58c4303a-6199-4dd9-b2a3-cbeaea850670 · outbound

This paper cites On memorization of large language models in logical reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models On memorization of large language models in logical reasoning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.820215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.820215Z digest=sha256:cea205d014a896982a8cd3c6c4423e4eb9a7a6148aa8758b8993c1b3657de702

Observation 9b4d95d0-78fe-4469-8085-f4d0a7bc3e70 · outbound

This paper cites Polymath: Evaluating mathematical reasoning in multilingual contexts.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Polymath: Evaluating mathematical reasoning in multilingual contexts

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.930098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.930098Z digest=sha256:26cea1114c556cfdc241a73f24db1df3f26ad1d70c993e3a420c1fa2055ef4e6

Observation 5531e0cf-a9bf-490c-af53-0bd0060f6645 · outbound

This paper cites Mmath: A multilingual benchmark for mathematical reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Mmath: A multilingual benchmark for mathematical reasoning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.046656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.046656Z digest=sha256:2f52ca6b0b9b9d041e3bf7a92019cb47c17d0b0e3bc2e097bf6bf6820534a898

Observation 392d09f2-57d7-4d8c-b073-7ba207256783 · outbound

This paper cites Matheval: A comprehensive benchmark for evaluating large language models on mathematical reasoning capabilities.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Matheval: A comprehensive benchmark for evaluating large language models on mathematical reasoning capabilities

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.207068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.207068Z digest=sha256:6b3e21d423dffef4338adc6cd2bc98194f8b042d98d7471ab4656baa9f63b481

Observation b0a54985-5382-4159-81f8-96873230ee51 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.362570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.362570Z digest=sha256:ff3dee68a8d3baa3dcccf8abb24614f9c8b48bba7adc9bf6ef566d474bca3032

Observation 6120cf10-79bc-42ed-a82a-d685be410383 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Measuring multimodal mathematical reasoning with math-vision dataset

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.517056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.517056Z digest=sha256:9711752ef495463cfd12ba044bf0cbb383d34a50c591751e22935c923645168c

Observation 04a3dc8d-ae8b-4a4c-9032-c7bd20109f4a · outbound

This paper cites MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.635190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.635190Z digest=sha256:a91450bde5704d24e3308983db8526f0300def0444a0d870330365f800d86d33

Observation 1b4baf3b-32aa-4fc1-9b59-7f2ed226b745 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.792230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.792230Z digest=sha256:d2265cad6108e3231073fbeeecf3657488e155898cd21259553b9a7e0e69110f

Observation 3c88d0d2-06e0-49ff-a296-1b950b37acc7 · outbound

This paper cites Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:56.951428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:56.951428Z digest=sha256:f2541889fcbbb8e2eae513f4885b3d5c13349c3468e2d83e8e7d65cb8b0e1b0e

Observation 37fe66f5-d49a-4d2f-aa9e-afa511215aa8 · outbound

This paper cites Evaluating and improving tool-augmented computation-intensive math reasoning.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Evaluating and improving tool-augmented computation-intensive math reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.058728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.058728Z digest=sha256:8922a75070dff3102d21bb71d8b664585982a8b2bb39310a7b9de034d6b0de7b

Observation be6990d5-da67-42aa-a201-a94e14b909ca · outbound

This paper cites MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models MARIO: MAth Reasoning with code Interpreter Output -- A Reproducible Pipeline

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.154852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.154852Z digest=sha256:27ee4b7c7cbb85d061410504b9729c4ac783b9427821841a244a239e4eb45fc8

Observation 866f18a4-6413-4b82-9d1e-6782299b57bd · outbound

This paper cites Malt: Improving reasoning with multi-agent llm training.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Malt: Improving reasoning with multi-agent llm training

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.314261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.314261Z digest=sha256:8659353c75c60c1c800aba52cb545048257f2f0f007a0847ed07bc2d96afa931

Observation 0da25313-98b7-4fd1-9cd9-10ac9341f7a0 · outbound

This paper cites {\dag} dagger: Distractor-aware graph generation for executable reasoning in math problems.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models {\dag} dagger: Distractor-aware graph generation for executable reasoning in math problems

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.466725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.466725Z digest=sha256:8e66ed2ec4f02a110effc2ff6acab3e1215bb78a615ce614c020f6e5b6ae14e4

Observation 14ff5c91-a02b-4f94-82c7-30371db05577 · outbound

This paper cites Structured reasoning with tree-of-thoughts for bengali math word problems.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Structured reasoning with tree-of-thoughts for bengali math word problems

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.580171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.580171Z digest=sha256:25bb012711336fd23ffe1caaa88dc9cab757beb6bef79930d3d547b4d963b019

Observation 8ddffcde-488c-4ede-9011-ed8ecae53410 · outbound

This paper cites Groqcloud api documentation, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Groqcloud api documentation, 2025

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.709445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.709445Z digest=sha256:e5e1865e921bdf54b9d7a7b6e75abe78e0c70cc1de56220a91897647da6ce9c7

Observation cb858195-0fc2-457c-a813-ee356ab80ed2 · outbound

This paper cites Qwen3-32b technical report, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Qwen3-32b technical report, 2025

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.813164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.813164Z digest=sha256:a4e2228a3d29f5daaee58a0fe3bb32baf26e036a5a53462b89363c6238c6b1fa

Observation f7a35da5-55d1-4830-a73b-975de06302ca · outbound

This paper cites Llama 3.1 model card, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Llama 3.1 model card, 2025

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.890633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.890633Z digest=sha256:342401fb31fc111f75d7f21d20f458f3ec11d8ccb081cfc9f521fdd2daa8ccdd

Observation cb55f4e3-0bf3-42e1-af53-9f1b8a10819b · outbound

This paper cites Llama 3.3 model card.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Llama 3.3 model card

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:57.991945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:57.991945Z digest=sha256:e05d0f05b85e5eccd108843f5e7c35a57aa8088660c77e8443fc1b49eacc9d2f

Observation 53336ff1-0632-4844-a00d-64909720bd8e · outbound

This paper cites Llama 4 scout model card, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Llama 4 scout model card, 2025

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:58.188112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:58.188112Z digest=sha256:083de5814730c814270aed49d9a9b48594712deb64ad948d9edace0584ed4ead

Observation 50abc97c-1f00-4f60-9002-ee88ef87e893 · outbound

This paper cites Gpt-oss: Open weight language models, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Gpt-oss: Open weight language models, 2025

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:58.345895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:58.345895Z digest=sha256:bae154e43a267e9a9a1fad929f574e9c6e753de10ad71c40114c0a7590fdbd66

Observation 741d1a19-425b-4c7a-afd1-76fec1261d6f · outbound

This paper cites Is gpt-oss good? a comprehensive evaluation of openai’s latest open source models.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Is gpt-oss good? a comprehensive evaluation of openai’s latest open source models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:58.453639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:58.453639Z digest=sha256:bd96547d8a8d40778f33e99968ac6ff77870135dbd790b11f905522ec0e77147

Observation 7568bea2-b988-48c7-8a19-901d1765342b · outbound

This paper cites Gpt-oss 120b technical report, 2025.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Gpt-oss 120b technical report, 2025

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:58.571640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:58.571640Z digest=sha256:480eb3bb5d83632c182f24c75e3312221f70379d9cd825cba201aea1a7f91bba

Observation a08f0297-3fe5-4313-80ae-5a42044ba600 · outbound

This paper cites an unresolved cited work.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Unresolved cited work

Reference 2025

Resolution
parse uncertain
no resolver link, observed 2026-08-02T05:48:58.091669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:58.091669Z digest=sha256:630d6ab7f975bf5a38a92469a49000b8fd8761a3ee63381c4a03467df205b48c

Pith citing papers

Observation d4906a80-fc8d-498c-a0e3-e0aed751b58a · inbound

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs cites this paper.

PatiGonit22K: A Comprehensive Dataset for Solving Complex Bengali MWPs GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:31:25.001382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T15:31:24.935640Z digest=sha256:526a0bc3f51b53e3ad0998a5e1693367402d7feb855d7f46eafbe58a45a46393