Pith. sign in

REVIEW 4 cited by

Oracle-MNIST: a Dataset of Oracle Characters for Benchmarking Machine Learning Algorithms

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.09442 v2 pith:2XONVLZN submitted 2022-05-19 cs.CV

classification cs.CV
keywords datasetimagesoracle-mnistancientcharactersbenchmarkingclassificationlearning
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

We introduce the Oracle-MNIST dataset, comprising of 28$\times $28 grayscale images of 30,222 ancient characters from 10 categories, for benchmarking pattern classification, with particular challenges on image noise and distortion. The training set totally consists of 27,222 images, and the test set contains 300 images per class. Oracle-MNIST shares the same data format with the original MNIST dataset, allowing for direct compatibility with all existing classifiers and systems, but it constitutes a more challenging classification task than MNIST. The images of ancient characters suffer from 1) extremely serious and unique noises caused by three-thousand years of burial and aging and 2) dramatically variant writing styles by ancient Chinese, which all make them realistic for machine learning research. The dataset is freely available at https://github.com/wm-bupt/oracle-mnist.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. JieZi: A Large-Scale Expert-Audited Dataset and Benchmark for Ancient Chinese Character Exegesis

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A new expert-audited dataset and benchmark for ancient Chinese character exegesis shows that multimodal LLMs improve substantially when fine-tuned on domain-specific VQA data.

  2. Enhancing Environmental Robustness in Few-shot Learning via Conditional Representation Learning

    cs.CV 2025-02 reject novelty 5.0 of 10

    A conditional representation network with cross-attention and 4D convolution is claimed to improve few-shot classification on a new hard-query benchmark by 6.83% to 16.98%.

  3. Federated Testing (FedTest): A New Scheme to Enhance Convergence and Mitigate Adversarial Attacks in Federating Learning

    cs.LG 2025-01 conditional novelty 5.0 of 10

    FedTest lets users evaluate each other's models with local data and aggregates models by these scores, claiming faster convergence and better robustness to malicious users.

  4. Vision Eagle Attention: a new lens for advancing image classification

    cs.CV 2024-11 conditional novelty 3.0 of 10

    Adding three small convolutional attention blocks to ResNet-18 improves classification accuracy by up to 1.5 percentage points on three benchmark image datasets.

Pith tools