Pith. sign in

Paper Citation Record · LEDGER

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2605.24703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24703 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T13:11:48.095713Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact8
  • verified fuzzy35
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a0ad9340-ee7d-4fc4-b817-7f98c17cd78a · outbound

This paper cites ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large- Scale Multitask Dataset, 2025.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering ITFormer: Bridging Time Series and Natural Language for Multi-Modal QA with Large- Scale Multitask Dataset, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.597128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:6447ac555745fa45f16f2674812bb34697d9928a0c97950e7d4cb37729baa959

Observation 00cf001c-46c4-4862-a3af-fa5d5cad38b7 · outbound

This paper cites Chatts: Aligning time series with llms via synthetic data for enhanced understanding and reasoning.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Chatts: Aligning time series with llms via synthetic data for enhanced understanding and reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.710518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:1396076d95c52945730b428fe0fae93ac1335a5f6b635cde290a99475e70db50

Observation 71351116-385a-40cf-978c-67ce30e8e8ed · outbound

This paper cites ECG-QA: A Comprehen- sive Question Answering Dataset Combined with Electrocardiogram.Conference on Neural Information Processing Systems, 36:66277–66288, 2023.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering ECG-QA: A Comprehen- sive Question Answering Dataset Combined with Electrocardiogram.Conference on Neural Information Processing Systems, 36:66277–66288, 2023

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.606442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:6d06b46cba1d98946301f95927266762b919228790db7f81c804dacd3e3ada16

Observation 0717984d-7c3b-4f83-ae66-e1ec84585ff3 · outbound

This paper cites SensorQA: A Question Answering Benchmark for Daily-Life Monitoring, 2025.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering SensorQA: A Question Answering Benchmark for Daily-Life Monitoring, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.608255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:e32fc6500c3853fa4cefb45819f08bdbb7da004158b1cc94cce0b5fc77d6dc1e

Observation 807842a6-1c2c-4799-99b5-1719940df86e · outbound

This paper cites TimeSerie- sExamAgent: Creating Time Series Reasoning Benchmarks at Scale, 2026.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering TimeSerie- sExamAgent: Creating Time Series Reasoning Benchmarks at Scale, 2026

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.602696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:03c75fc7b46a2096df24371fd9036e053c47b55320c1d9fd03ded248bfba46bc

Observation 6a148433-c3f9-42ad-a5c0-8f77a31ea651 · outbound

This paper cites Towards Time Series Reasoning with LLMs.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Towards Time Series Reasoning with LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.698832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:faf3a375e7d207af7a3976bc8ae993e1af065aa944153261cea003890a94fe91

Observation c2f56070-dba6-406b-9c8e-d8fbea122e77 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.600886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:2d446e11086958b669726dd5d1b625b82a9a496a83e7810a328a97309aed60a1

Observation 3d5fef1d-d225-4094-8a02-39ef14fd7674 · outbound

This paper cites Natural Questions: A Benchmark for Question Answering Research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Natural Questions: A Benchmark for Question Answering Research.Transactions of the Association for Computational Linguistics, 7:453–466, 2019

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.604557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:70d47324858acd0a580db14b80cd67c525464fa873e21ed3cd012cbf5b3a52cc

Observation b96cf92e-ae68-4aa1-8dc2-d4d881e0b97c · outbound

This paper cites TGIF-QA: Toward Spatio- Temporal Reasoning in Visual Question Answering.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering TGIF-QA: Toward Spatio- Temporal Reasoning in Visual Question Answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.610018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:7733fd993214ae06048c15024b0185cbd6332ad9701cad676a40b27e7db5c0d2

Observation 8b063946-4595-4472-8558-92012ae37ea7 · outbound

This paper cites AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering AGQA: A Benchmark for Compositional Spatio-Temporal Reasoning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.613435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:9c449427ae585a132e3a2567ba6c563769819f5e058b49b4633950fa7a418e03

Observation 7c6b8883-8f99-41b1-bb7d-981222e1ec12 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Video Question Answering: Datasets, Algorithms and Challenges

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.599128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:12fe9b2a89e38446d21d70bc266b8316461f78570092fc2c5fd4e7a082f9ee2d

Observation 120eff2b-9930-4bea-a27a-4e234c15e2dc · outbound

This paper cites Deep Learning for Time Series Anomaly Detection: A Survey.ACM Computing Surveys, 57(1):1–42, 2024.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Deep Learning for Time Series Anomaly Detection: A Survey.ACM Computing Surveys, 57(1):1–42, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.629611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:df8b9035de5262ab3a104fc90d1948a0fa2cdb54d664ee55ecac462ab14e615b

Observation 9f82b974-cce2-4264-8d79-e23aeb0d1975 · outbound

This paper cites Anomaly Detection in Time Series: A Comprehensive Evaluation.Proceedings of the VLDB Endowment, 15(9):1779–1797, 2022.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Anomaly Detection in Time Series: A Comprehensive Evaluation.Proceedings of the VLDB Endowment, 15(9):1779–1797, 2022

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.589692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:00cc7faae8f8f036667542e3bca99656a4f0604d304898715c23c920811a3eb0

Observation ef020db6-dee6-4683-bae4-ab9cf6965876 · outbound

This paper cites TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-08-04T02:23:40.720787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:8702f6de581d6815afcd7fbe695974261f0313fb821d9ef5bb78a01b9830c06f

Observation 690752b1-30ae-408c-b510-56c051b13bc6 · outbound

This paper cites STL: A Seasonal-Trend Decomposition Procedure Based on Loess.Journal of Official Statistics, 6:3–73, 1990.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering STL: A Seasonal-Trend Decomposition Procedure Based on Loess.Journal of Official Statistics, 6:3–73, 1990

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.587855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:116d63ace3a7b987a061422ba129997a889509ac1d5f419fbe61b42048f753f4

Observation 4562aeb8-466c-4ebb-b737-06457250adff · outbound

This paper cites Time-Series Anomaly Detection: Overview and New Trends.Proceedings of the VLDB Endowment, 17(12):4229–4232, 2024.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Time-Series Anomaly Detection: Overview and New Trends.Proceedings of the VLDB Endowment, 17(12):4229–4232, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.593371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:cd9d62b7be5bbb2cae2babc6e05b1a80ecc02177cbc271bfbdabb8ee878cf25d

Observation 78d45ba8-b21c-4da7-b4a1-72fc5062cd3b · outbound

This paper cites A Survey of Methods for Time Series Change Point Detection.International Conference on Information and Knowledge Systems, 51(2):339–367, 2017.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering A Survey of Methods for Time Series Change Point Detection.International Conference on Information and Knowledge Systems, 51(2):339–367, 2017

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.615222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:f022cde453bd60e08c13e8235a8121547b349d5e2a811c79a663c37fb9d65ef3

Observation a1790ebf-b40a-474a-a8ac-0834c3989f04 · outbound

This paper cites Selective Review of Offline Change Point Detection Methods.IEEE Transactions on Signal Processing, 167:107299, 2020.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Selective Review of Offline Change Point Detection Methods.IEEE Transactions on Signal Processing, 167:107299, 2020

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.586042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:0330075e966ef348927750548c479b19df469300a2a85ae031f211f107f040c1

Observation e06b038e-d0dc-49dc-80b7-ee38ae31f570 · outbound

This paper cites Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement, 2025.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement, 2025

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.595239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:47411ada9c0f5a1874aab65c2babfc5f481ecf4662cfea6ba108274065270969

Observation 4d0920f5-3632-47d7-a4d8-98a4f6d21cff · outbound

This paper cites Monash Time Series Forecasting Archive.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Monash Time Series Forecasting Archive

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.704363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:a684d56da4d7f8d079201e8749caa0dbd4b7b154ca05f35ee91f6c9caf62a5a9

Observation 81e7de9d-4145-4390-961b-265f6cdae667 · outbound

This paper cites Toward Reasoning-Centric Time-Series Analysis, 2025.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Toward Reasoning-Centric Time-Series Analysis, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.582487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:5f68f4572330a0a1c12e8d444e339c1ce1b07ed8ac77ba7a1b90b3a5e67343c0

Observation bed7c868-cc59-42c3-a280-31f9c22d7b07 · outbound

This paper cites Xu, Harish Haresamudram, Catherine W.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Xu, Harish Haresamudram, Catherine W

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.579020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:66a45854c74d021ceceabcfb58a36023c03fcc035e33a4120ee070358d2dc6d0

Observation 1c816215-23fe-46f8-9090-b66a7724eebe · outbound

This paper cites Maddix, Abdul Fatir Ansari, Akash Chandrayan, Abhinav Pradhan, Bernie Wang, and Matthew Reimherr.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Maddix, Abdul Fatir Ansari, Akash Chandrayan, Abhinav Pradhan, Bernie Wang, and Matthew Reimherr

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.584243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:befc04d411e62f9e64572bc84bfd6a56146028032db611186c4938f4c0fd9547

Observation 79980840-38cf-4d4c-9577-9733c7d54611 · outbound

This paper cites Holistic Evaluation of Language Models.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Holistic Evaluation of Language Models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.580755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:282f898d4ab8a458045b31bdb8d6ec18f356f28295b542be18e4914a3e9ea695

Observation 7ed0d358-19c7-4219-8146-caa3cda5e89b · outbound

This paper cites MMTS-BENCH: A Comprehensive Benchmark for Time Series Under- standing and Reasoning.arXiv preprint arXiv:2602.08588, 2026.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering MMTS-BENCH: A Comprehensive Benchmark for Time Series Under- standing and Reasoning.arXiv preprint arXiv:2602.08588, 2026

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.690597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:df57058d54e95d9d1837286f11dd3aaa3bb89e87420aa35116ba34abbe696049

Observation fc6781fe-e191-4d7b-93af-18fc7107ef40 · outbound

This paper cites QuAnTS: Question Answering on Time Series.arXiv preprint arXiv:2511.05124, 2025.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering QuAnTS: Question Answering on Time Series.arXiv preprint arXiv:2511.05124, 2025

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.693381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:23a1a4c33b669617e80ab577c95c46d72ce41cb75d22f7bd249f5aa5413ff510

Observation e192a688-fd4f-4108-bebe-f45a597d31b5 · outbound

This paper cites Evaluating Large Language Models on Time Series Feature Understanding: A Comprehen- sive Taxonomy and Benchmark.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Evaluating Large Language Models on Time Series Feature Understanding: A Comprehen- sive Taxonomy and Benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.591579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:fdbc1ce7de3f746da40367e0402f16e4680cc97b1152e08be2bea891df0cc16f

Observation ea28cc74-1676-4909-8629-a3a8f7e3edb1 · outbound

This paper cites WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions, 2025.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex Instructions, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.611752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:fb08fab141ecd11e956b770e8f51f633b8ebf79eef8539df8c91572afbb2be39

Observation 9b7dea95-7c2d-493c-8535-fcfe2027ae03 · outbound

This paper cites Best Practices for the Human Evaluation of Automatically Generated Text.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Best Practices for the Human Evaluation of Automatically Generated Text

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.635428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:a4ba1ffe9362a35d63e7ac40877ffa361d40796c7b8dad6d37c6dcd8547bf3ad

Observation 17d11f36-e677-4bfb-96d4-ae0fc5342f7d · outbound

This paper cites Measuring Massive Multitask Language Understanding.International Conference on Learning Representations, 2021.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Measuring Massive Multitask Language Understanding.International Conference on Learning Representations, 2021

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.637251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:f3595c9f20665e891812a6e5d864159ec71675b7c3ff25448d4ba996ebb142d4

Observation 38e969ce-5f9b-407e-8071-8b444cd41759 · outbound

This paper cites TSAQA: Time Series Analysis Question And Answering Benchmark.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering TSAQA: Time Series Analysis Question And Answering Benchmark

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-06-30T13:14:40.695790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:2841e86b88b7397c312814017543e86a7736de41956929451d95cbe7a2ca136f

Observation 2bf0f460-a9f1-4103-9af8-10d2066558d4 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.640840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:f2e72d4602e874983e6d02f16dcc63ece21e0e2e07ee3bdbdfe4c784bd11b68c

Observation e49c375a-0d3b-4c1b-8c97-9b8b7d492bb1 · outbound

This paper cites G-Eval: NLG Evaluation Using Gpt-4 with Better Human Alignment.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering G-Eval: NLG Evaluation Using Gpt-4 with Better Human Alignment

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.639084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:6f359e8650ccffd00c6969cfa66e89519ef44f1cced2ebfeb5975cfb6cf1b1fb

Observation 1a3ee8e3-00a0-4ec3-b8b0-768c8765710b · outbound

This paper cites Mtbench: A multimodal time series benchmark for temporal reasoning and question answering.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Mtbench: A multimodal time series benchmark for temporal reasoning and question answering

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.707394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:da8d7fcc546242d54b7b91f025598b3043382decf26a20649a31f4a37ee7724b

Observation 4ca17c17-f459-49f5-b567-bfa8ca620ec2 · outbound

This paper cites How to Do Human Evaluation: A Brief Introduction to User Studies in NLP.Natural Language Engineering, 29(5):1199–1222, 2023.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering How to Do Human Evaluation: A Brief Introduction to User Studies in NLP.Natural Language Engineering, 29(5):1199–1222, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.617060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:5d01a80250f8c475172b891c8693951a925866c3f59e4c78472d921db117ac4e

Observation 472f5faa-a2d9-4c43-a46e-7a598d68f7d6 · outbound

This paper cites High Agreement but Low Kappa: I.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering High Agreement but Low Kappa: I

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.626117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:a551aec73427f3fd542c82aad20a91e7e4cbd2a5485fa585a66413875888077b

Observation 00bea1ff-5ab4-4d17-8f75-45ce5d20bab7 · outbound

This paper cites High Agreement but Low Kappa: II.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering High Agreement but Low Kappa: II

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.627793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:e45114158ad6412705c89314ffcfcf0b6edb3b1d80c9518278dda9116688f473

Observation 1f545b2a-640b-49d7-b6bc-89bc3f268591 · outbound

This paper cites Leveraging Large Language Models for Multiple Choice Question Answering.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Leveraging Large Language Models for Multiple Choice Question Answering

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.618844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:7131b8e4412d81e715dfb6a5cd7e267ca29ab8554c5d7d8671456dc0e5cd0d91

Observation 38218e3b-2f72-4fed-a3f5-aa921f9a8503 · outbound

This paper cites Is Your Large Language Model Knowledgeable or a Choices-Only Cheater? InWorkshop on towards Knowledgeable Language Models (KnowLLM), 2024.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Is Your Large Language Model Knowledgeable or a Choices-Only Cheater? InWorkshop on towards Knowledgeable Language Models (KnowLLM), 2024

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.622481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:78e0c66759542cebdb9596a0f583907f1885a0b1a84bbce4f0322a8ec0b5fde8

Observation 2ce5e5e0-e909-4b00-aab5-b15ebc1e298d · outbound

This paper cites Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Which of These Best Describes Multiple Choice Evaluation with LLMs? A) Forced B) Flawed C) Fixable D) All of the Above

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.620699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:0bbec1005633b4441b7efcc23c0cb13358b4ef4e80babc310772a55c55850c7a

Observation ebc8e271-6bfb-4a2b-828f-7f6a8e62626f · outbound

This paper cites Large Language Models Sensitivity to the Order of Op- tions in Multiple-Choice Questions.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Large Language Models Sensitivity to the Order of Op- tions in Multiple-Choice Questions

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.624447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:9ba39b9d0f2cf12a9fdf63b7d3ba08ec93ce58ab405508c5bbb3c1be2c8b6988

Observation 2ba9e0f0-df25-49d2-b676-2a2ded139ba4 · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Large Language Models Are Not Robust Multiple Choice Selectors

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:14:40.681857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:17065d52acbff439de5f034f14ef323c84c60a65e081884cb23458100d4b65ab

Observation 28b33d7b-898c-4a4d-aea5-426162ec297e · outbound

This paper cites Datasheets for Datasets.Communications of the ACM, 64(12):86–92, 2021.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Datasheets for Datasets.Communications of the ACM, 64(12):86–92, 2021

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.633535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:a9cec30976d5ae123f814a005c320c24d3d41406328baeeac6ac25e961034e53

Observation 0c5886ba-2870-44d7-913a-7d5e2e6b4957 · outbound

This paper cites Data Cards: Purposeful and Transpar- ent Dataset Documentation for Responsible Ai.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Data Cards: Purposeful and Transpar- ent Dataset Documentation for Responsible Ai

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T03:15:57.631456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:b3bb520e89303d4e841ca5d82134d1bbf3610bdf036fc90781a2f8846899a6fd

Observation 307ff7b8-9087-4d53-916f-0fc4f0d5a29c · outbound

This paper cites Synthetic Data -- what, why and how?.

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Synthetic Data -- what, why and how?

Reference 45

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T13:14:40.688030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T13:11:48.095713Z digest=sha256:2bce60ebf62b0961d3e6a66bdb59ac807e14e5c122e01ac80acfdb35fd555d72

Pith citing papers

No inbound Pith citation observations are available.