STORY · FORSKNING_

«Truth is not a direction»: criticism of LLM probing methods

The article analyzes Tarski's theorems and argues against the idea that truth in language models can be represented as a single linear direction in the model's vector space. This challenges popular probing techniques that assume truth representations are unambiguously identifiable.

WHY IT MATTERS

If truth is not unidirectional as earlier research has assumed, better methods are needed to understand and demonstrate what language models actually "know" or "believe" – essential for trust and safety.

SOURCES

MACHINE-GENERATED SUMMARY This summary is written by machine from the sources below. We sort and explain — but we are a way into the field, not the final word. Check the source when something matters to you.