Back to the shelf
Read
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
This study evaluates whether large language model explanations reliably reflect their decision behavior using intervention-based necessity and sufficiency scores. Testing eight models across advisor recommendation and prompt monitoring tasks, the authors find that cited factors show only partial agreement with measured influence. Uncited factors often score higher than cited ones, indicating that self-reported rankings do not consistently identify the strongest drivers of model outputs.
Key takeaways
- The study tests necessity, where changing a factor alters the output, and sufficiency, where retaining it preserves the output.
- Eight models from Claude, GPT, and Gemini families were evaluated using controlled black-box interventions.
- Mean Spearman correlations for advisor recommendation are 0.349 for necessity and 0.354 for sufficiency.
- Mean Spearman correlations for prompt monitoring are 0.431 for necessity and 0.580 for sufficiency.
- Uncited factors score above the lowest cited factor in 57.6 percent of advisor responses under necessity.
- Uncited factors score above the lowest cited factor in 58.1 percent of advisor responses under sufficiency.
Reflections
No notes on this one yet. I add reflections as I finish or revisit a book.