Pith. sign in

hub

104.https://eleosai.org/

13 Pith papers cite this work, alongside 22 external citations. Polarity classification is still indexing.

13 Pith papers citing it
22 external citations · external index

hub tools

years

2026 12 2025 1

representative citing papers

Probing Persona-Dependent Preferences in Language Models

cs.CL · 2026-05-13 · unverdicted · novelty 6.0

Linear probes on residual-stream activations identify a shared preference vector in LLMs that tracks choices across prompts and causally steers decisions even for anti-correlated personas.

Initial results of the Digital Consciousness Model

cs.CY · 2026-01-22 · unverdicted · novelty 6.0

A new probabilistic model integrates leading consciousness theories to assess AI, finding moderate evidence against 2024 LLMs being conscious but weaker evidence than for simpler AI systems.

Consistency Training Along the Transformer Stack

cs.LG · 2026-06-04 · unverdicted · novelty 5.0

The paper introduces MLPCT and AttCT internal consistency targets and applies consistency training to four new safety threats, reporting reduced misalignment and some cross-threat generalization.

AI Consciousness and Existential Risk

cs.AI · 2025-11-24 · unverdicted · novelty 2.0

Consciousness does not directly predict AI existential risk unlike intelligence, though it may indirectly affect risk through alignment or capability requirements.

citing papers explorer

Showing 13 of 13 citing papers.