Pith. sign in

REVIEW 1 cited by

Mathematical Language Processing Project

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1407.0167 v1 pith:NHOHDPSU submitted 2014-07-01 cs.DL cs.CLcs.IR

classification cs.DLcs.CLcs.IR
keywords approachidentifierslanguagemathematicalmeaningdefinitionsidentifier-definitionprocessing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In natural language, words and phrases themselves imply the semantics. In contrast, the meaning of identifiers in mathematical formulae is undefined. Thus scientists must study the context to decode the meaning. The Mathematical Language Processing (MLP) project aims to support that process. In this paper, we compare two approaches to discover identifier-definition tuples. At first we use a simple pattern matching approach. Second, we present the MLP approach that uses part-of-speech tag based distances as well as sentence positions to calculate identifier-definition probabilities. The evaluation of our prototypical system, applied on the Wikipedia text corpus, shows that our approach augments the user experience substantially. While hovering the identifiers in the formula, tool-tips with the most probable definitions occur. Tests with random samples show that the displayed definitions provide a good match with the actual meaning of the identifiers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. E-Gen: Leveraging E-Graphs to Improve Continuous Representations of Symbolic Expressions

    cs.LG 2025-01 conditional novelty 6.0 of 10

    An e-graph-based data generator produces 55 million equivalent-expression training pairs, and embeddings trained on them beat prior math-embedding models and GPT-4o on several symbolic math tasks.

Pith tools