A diagonal attack for LLM truth probes shows why no probe on a language model’s embedding space can pin down truth.
A linear dream
... continue reading
This article highlights the limitations of using embedding space directions to determine truth in large language models (LLMs). It demonstrates that the idea of a single 'truth' direction is fundamentally flawed, echoing Gödel's findings that truth cannot be fully captured by formal systems, which has significant implications for AI safety and the development of trustworthy AI systems.
A diagonal attack for LLM truth probes shows why no probe on a language model’s embedding space can pin down truth.
A linear dream
... continue reading