Back to the shelf
Read
Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence
This study evaluates whether LLM-generated explanations reliably reflect the factors driving their decisions. Using intervention-based necessity and sufficiency scores across eight models in advisor recommendation and prompt monitoring, the authors find that cited features show only partial agreement with measured influence. Uncited factors often score higher than cited ones, indicating that self-reported rankings do not consistently identify the strongest decision drivers.
Key takeaways
- The study tests necessity, where changing a factor alters the output, and sufficiency, where retaining it preserves the output.
- Eight models from Claude, GPT, and Gemini families were evaluated using controlled black-box interventions.
- Mean Spearman correlations between cited rankings and necessity scores range from 0.349 to 0.431 across use cases.
- Mean Spearman correlations between cited rankings and sufficiency scores range from 0.354 to 0.580 across use cases.
- Uncited factors score above the lowest cited factor in 57.6 percent of advisor responses under necessity.
- Uncited factors score above the lowest cited factor in 58.1 percent of advisor responses under sufficiency.
Reflections
No notes on this one yet. I add reflections as I finish or revisit a book.