Position: Semantic Uncertainty Measures Disagreement, Not Reliability
At a glance
- Citations
- 0
- References
- 0
- Comments
- 0
Abstract
<div> Large Language Models and their multi-modal variants see increasingly rapid adoption and deployment, including in settings where reliability matters. However their uncertainty remains difficult to assess as classic or early token-level approaches cannot deal with the open-ended outputs of such models. Semantic uncertainty arises as a promising solution to the limits of token-level confidence by measuring disagreement among multiple generated responses. This position paper argues that this framing is incomplete: semantic uncertainty measures semantic disagreement, not actual reliability. A model can be uncertain while producing several valid answers, or certain while repeatedly producing the same wrong answer. We introduce a semantic bias-uncertainty decomposition to show that reliability depends both on variability across meanings and systematic deviation from correct or grounded meanings. This perspective reveals that common hallucination-detection evaluations conflate uncertainty with error. We argue for reliability-centered evaluation that separates semantic disagreement, correctness, and enables a more fine-grained characterization of uncertainty beyond the level of the full answer. </div>
Publication details
- OpenAlex
- W7163602350
- Document type
- preprint
- Language
- EN
- Source
- HAL (Le Centre pour la Communication Scientifique Directe)
- Last metadata update
Comments
Log in to join the discussion.