REVIEW 5 cited by
On Characterizations for Language Generation: Interplay of Hallucinations, Breadth, and Stability
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
We study language generation in the limit - introduced by Kleinberg and Mullainathan [KM24] - building on classical works of Gold [Gol67] and Angluin [Ang79]. [KM24]'s main result is an algorithm for generating from any countable language collection in the limit. While their algorithm eventually generates unseen strings from the target language $K$, it sacrifices coverage or breadth, i.e., its ability to generate a rich set of strings. Recent work introduces different notions of breadth and explores when generation with breadth is possible, leaving a full characterization of these notions open. Our first set of results settles this by characterizing generation for existing notions of breadth and their natural extensions. Interestingly, our lower bounds are very flexible and hold for many performance metrics beyond breadth - for instance, showing that, in general, it is impossible to train generators which achieve a higher perplexity or lower hallucination rate for $K$ compared to other languages. Next, we study language generation with breadth and stable generators - algorithms that eventually stop changing after seeing an arbitrary but finite number of strings - and prove unconditional lower bounds for such generators, strengthening the results of [KMV25] and demonstrating that generation with many existing notions of breadth becomes equally hard, when stability is required. This gives a separation for generation with approximate breadth, between stable and unstable generators, highlighting the rich interplay between breadth, stability, and consistency in language generation.
Forward citations
Cited by 5 Pith papers
-
On Computational Hardness of Mistake-Bounded Language Generation: A Random-Oracle Query Separation
Under a random oracle, a countable family of infinite languages has closure dimension zero and admits zero-mistake unbounded generation, yet every polynomial-query generator incurs an exponential expected-mistake lowe...
-
Validity, Sparse Holes, and Breadth in Language Generation: Banach Density, Topology, and Geometry
Under the stricter Banach-density measure, valid generation in the limit guarantees the optimal 1/2 coverage exactly when the language collection has finite Cantor-Bendixson rank; other collections force arbitrarily l...
-
Representative Language Generation
A new 'representative generation' requirement is formalized, characterized by a group closure dimension, with a feasibility result under finite support and a membership-query impossibility.
-
Hallucination Rates in Language Generation
Allowing infinitely many but rare hallucinations strictly enlarges the class of languages generatable in the limit, and the allowed hallucination rate orders these classes into a strict hierarchy.
-
Characterizing the Effect of Noise in Language Generation in the Limit
In uniform and non-uniform language generation in the limit, noise level 1 and any finite noise level are equivalent, and the first noisy string strictly reduces the family of generatable collections.
Discussion (0). Continue with ORCID to comment.