Pith. sign in

Paper Citation Record · LEDGER

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

As of 6 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2508.18756.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18756 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:17:42.154186Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:13:54.031036Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:43.816739Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06d29732-220e-4398-b214-e33e734496f2 · outbound

This paper cites Program Synthesis with Large Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.862060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.862060Z digest=sha256:77d7b370bd852075509b5eb4d82e99aaca08aee3944f22bfd325d7c05bde26fb

Observation bf4eed67-8daf-4c43-8846-cac6c833c7be · outbound

This paper cites Memory Layers at Scale.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Memory Layers at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.871650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.871650Z digest=sha256:ddc7c0d7b5510b4cb8b4bc657794e35b34fbd41746c6b4357c66affbbd9a4597

Observation 8ec2a355-90da-4f99-b3b6-d2b9780417c3 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Piqa: Reasoning about physical commonsense in natural language

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:43.095234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:41.882241Z digest=sha256:74a7666c005eea3c11bfd468723e9989f690901c2829e40c68ec2b947143cda2

Observation 7176ecf7-38b0-4c4f-8966-32b3acc9f0f5 · outbound

This paper cites an unresolved cited work.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.891161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.891161Z digest=sha256:0bb2d170b9d2136f8425cba43849e1cdddd3be16f08b09ee9f93261295b3bb7a

Observation 4ec88c91-b50e-4e60-9e65-e55879060f8b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.898090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.898090Z digest=sha256:096aa4623e374a01ffaf943e7ce495a7df221394ef27cd0b1be56fd8384788e2

Observation 6e86179a-0fa9-41f1-8fdf-703e4db0a865 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.903827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.903827Z digest=sha256:1eb9f3320f4e8d25258e431cdb53b565dd2dca4ce3cf34164284da64334ff005

Observation f1c6fb34-e731-41cc-aa3f-ff880a8298ca · outbound

This paper cites Approximating Two-Layer Feedforward Networks for Efficient Transformers.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.909750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.909750Z digest=sha256:af147f389a1ec37bae9e93ce822dff8709949cf0b7341655065ce8d78b4a7cfb

Observation bd542213-d772-41c0-9435-d9a454db8c4f · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.916736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.916736Z digest=sha256:b4fa2ae91eeeeeae78af5532a8529a88aeaad8b119144af8bf452963d390b4a5

Observation 21d0a2ff-763f-45c3-83b6-db892710129a · outbound

This paper cites DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:43.057371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:41.924101Z digest=sha256:c4a8cf917ac48a9f559117469c2532f652863d2940db85788590fbd0807d3dd2

Observation b28f6439-323b-4962-b986-44410a875be0 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.932009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.932009Z digest=sha256:054dc1361637f620d6e9bf2c37fcffe66a788ddfdf5f02e3ac6c705c912c3183

Observation 90764af7-1ac7-4b06-8082-f40da840398e · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning FastMoE: A Fast Mixture-of-Expert Training System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.941220Z digest=sha256:cf2a0f1f74659e260862e8aec27040ee232567c093c3dfa3184dae0d1cba8093

Observation 7b5d0f0f-6f70-4f1f-9b93-74b345fadb35 · outbound

This paper cites Mixture of A Million Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Mixture of A Million Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.950238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.950238Z digest=sha256:c6eebd3718e3c104f1dc3ff6b69f891a86a545517212aef96389d43137338e8a

Observation e8412630-7a60-4108-8d70-e5f05c4506fb · outbound

This paper cites Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.956630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.956630Z digest=sha256:a258a81c0076cd1c41c8a31ed3684d39c5f756a16a0a42242d71458fe5b9f17f

Observation 74431a28-f424-4255-9527-53e7dcb99fd1 · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.963854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.963854Z digest=sha256:49237cf883708b08cf53f50e43a9b1b690d15f484bb4a26a07c8da65e46ccce4

Observation 552c83df-6bf9-4d15-91a2-b1eebbd700d8 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.NeurIPS, 2021.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Measuring mathematical problem solving with the math dataset.NeurIPS, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.973737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.973737Z digest=sha256:2b4a83320fc38544d600216f27f48505f4254d8d8cd3f014e2166ebd5a115071

Observation 2dd5a2a9-6fac-496a-b3a1-432231d05997 · outbound

This paper cites Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.981651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.981651Z digest=sha256:3de54f8d448fff765d3ea25bd51385ed155c7c9898d0a2967e965c6c387de2ec

Observation d9fc7fbc-9f52-419b-9d0c-6cbae2d58457 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.964852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:41.989216Z digest=sha256:68f03718ae9296d56f9495d7c67f8bfa9aac8f8af965649e688d31dfe008f732

Observation e8ecbb92-aa3c-432f-9a7c-a9f8c04ef171 · outbound

This paper cites Ultra-Sparse Memory Network.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Ultra-Sparse Memory Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.998354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.998354Z digest=sha256:7f48c2ef51b6cf853923cc3bedf511f51239f08b2061e192a96714fd652e1e22

Observation f339dd7c-1b5f-4c73-b747-c9875d60c6f5 · outbound

This paper cites dots.llm1 Technical Report.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning dots.llm1 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.004470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.004470Z digest=sha256:a415c45ce6404511779a297c4ac9f868ab6ba616342c735a40347c5ceb0134ef

Observation aa57d30e-03f4-47d6-8ff9-e67b51d831d7 · outbound

This paper cites Product quantization for nearest neighbor search.IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128, 2010.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Product quantization for nearest neighbor search.IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128, 2010

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.010279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.010279Z digest=sha256:1383d21e380c5fef42ea2be9b8c34e619662460516e5c1cf34dd34c7a4a64334

Observation ac4e10c0-960f-4123-be6b-4c0c545c5a69 · outbound

This paper cites Mixtral of Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Mixtral of Experts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.015268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.015268Z digest=sha256:4d0193c8664cd1d24ce0b82e1605df00906d024af5cd40228eec2062f30ceef5

Observation 5448c1fb-c3ec-4903-9c3c-64346f545032 · outbound

This paper cites an unresolved cited work.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.020585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.020585Z digest=sha256:f6b54663e95f54bd2a78d05af6355b24a5e13c1f58edccfcde92e35e03694167

Observation c4b2915f-07ec-48e6-9918-706388389735 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.025904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.025904Z digest=sha256:823f4397416d37c781eabf044ed4d260bdbda32d61bb0a2bccc7da1e89952b4f

Observation e005e7a5-2a00-4426-85f7-543c19ac5dec · outbound

This paper cites Large Product Key Memory for Pretrained Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Large Product Key Memory for Pretrained Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.032313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.032313Z digest=sha256:a6202bafe4c1dc812fd781e5ea7d4aab17b2817bbff710a25a38fca5191713fd

Observation e766e8a8-41d6-4b09-a7bd-4f0b18b05e96 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Scaling Laws for Fine-Grained Mixture of Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.037489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.037489Z digest=sha256:728879e66b466262104d3e2b04b07b8346ae7e7081be40abb3914c77d14f65ef

Observation 33ab54b5-1975-4c31-8a67-f73c40d16c14 · outbound

This paper cites Large memory layers with product keys.Advances in Neural Information Processing Systems, 32, 2019.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Large memory layers with product keys.Advances in Neural Information Processing Systems, 32, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.918347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:42.043484Z digest=sha256:d5e881fff5ec31a557fcd6f3f8c103a50b51954b75f6c7de4185bab493e3e521

Observation f8dc512d-f6d8-4d4e-8e75-efa4d47529ac · outbound

This paper cites DeepSeek-V3 Technical Report.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.050746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.050746Z digest=sha256:b2209a6141e9537edbd67e5ffedb53f54ebd8611ea75a198755fdc954ca4b91d

Observation 21d9ec13-c829-41cc-9880-81543435c15d · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.058600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.058600Z digest=sha256:ba99b655e2802f112acf4e70c19771421b3fe2e9fd292b4bba259702ce886fa7

Observation d9140e8c-5cb1-454c-9934-2386a73cfa2f · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning OLMoE: Open Mixture-of-Experts Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.066072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.066072Z digest=sha256:fcad46027658c51d29ed8ef2fb0b833d63a3b3b34b94204e354be93b158cf220

Observation 0b698c83-5376-4f15-832e-f29ac974cc19 · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Transformers without Tears: Improving the Normalization of Self-Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.074129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.074129Z digest=sha256:5e2b4238a7c5278bd89e3c20418525756d835511a58cf15419e466973a8173a1

Observation 673c66de-f8de-4e25-af8b-aa7ff99e8d7e · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.081799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.081799Z digest=sha256:0bf36d771597a95ae4fc9b2bacb145353d83914283e92092320c5d4a0e3df464

Observation 7e3617df-af46-4e77-bc6f-63f84a873655 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.088345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.088345Z digest=sha256:da8b564c85ad52b6795dd4c049efce9a7fd3e444a9aa8c8127b4898c40e1b51f

Observation 04d84cf5-9f83-46ab-8cbd-00cb00d46f2f · outbound

This paper cites GLU Variants Improve Transformer.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning GLU Variants Improve Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.097567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.097567Z digest=sha256:e949f47c5be25249e14f097dbed2ba742b645f5db1dd23d079ddb18745e556af

Observation d62372b7-4833-4869-904d-1f40b18da95c · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.104052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.104052Z digest=sha256:4c2be0eebb6827f5fb34d905a9002ab7aedafc0a7b833ec4a6c42a2d7f1e099f

Observation ce438d7d-5e54-4a3c-b4a0-78845ca3dc63 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.109834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.109834Z digest=sha256:6d493ebffa6be6c9092425f1082f563c53764f0765b955678b274f46a8c0b918

Observation e9893b79-0301-4daa-88ca-f6b7c2aa2e11 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.115682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.115682Z digest=sha256:8bb4b2d51c6c7c12ef313cff39a48ba0e5aa3e9dd7eb3ef35d202a7e743da358

Observation 7a0b0e3b-efbc-4799-98ed-7533ba8244f1 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.858993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:42.120614Z digest=sha256:bed608c8818a996b3c0299dd927e84b803761bd02e65af631b09868cb577ced1

Observation 55063835-0b19-4ea0-8d12-83f01057f58c · outbound

This paper cites Moec: Mixture of expert clusters.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Moec: Mixture of expert clusters

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.834215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:42.125390Z digest=sha256:3722f4970501cde4ea08aab4476581060cdcb443a1fd683dd17cab128f3b4c67

Observation 5dcddd93-72b7-4b4a-bf0d-487940f0af07 · outbound

This paper cites Qwen3 Technical Report.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Qwen3 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.130675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.130675Z digest=sha256:55ee765cd9899e06c9da1c9c76fe01ca08dcab0a6c8e1b16aecb3b5882bce8f8

Observation 277491ed-0343-46cb-b108-e7c227a182f7 · outbound

This paper cites Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.136085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.136085Z digest=sha256:161bca0bdc42c338961640171b001655fca6ca39802a5c63fe7a0988432ae4dc

Observation 0fd5e50b-593b-4303-934a-0ef98b620337 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.142224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.142224Z digest=sha256:36dac35dbc5a0366b1e4e823de83673be3e2e3b77167345be01ffabb580b8f20

Observation f386348f-bd3d-400b-8408-925a8d03cd26 · outbound

This paper cites Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.148465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.148465Z digest=sha256:5c0489caa0ece824aeddc294c2518f143e4ba308b6143442ea92109fe82cf1cf

Observation d97cc08f-2759-43f4-9634-36c8a1599b67 · outbound

This paper cites Pre-values.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Pre-values

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.787037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T16:17:42.154186Z digest=sha256:f5c837d30e955e588b9877ec21869fec5ab1960edcb0e8c656d748f44b6be788

Pith citing papers

Observation 9c5b9c36-33bd-4b84-8514-b1935b865e6d · inbound

MIDUS: Memory-Infused Depth Up-Scaling cites this paper.

MIDUS: Memory-Infused Depth Up-Scaling UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:03:36.159265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T22:02:42.297041Z digest=sha256:fdbb056071b6e8606fcd064cdbcf6b65a1b8896eddcf00246694857ef6fa4e1e

Observation 18678fdf-ed1f-4050-b3ff-cb9b215df4f0 · inbound

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks cites this paper.

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:43.818439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T07:13:54.031036Z digest=sha256:4a3deafcaed916cf8d7a27d1dcb5f6fdf85cb10504413c94af27b58bd78c1479