Confidence and Calibration of Activation Oracles for Reliable Interpretation of Language Model Internals
This paper compares five confidence methods across four Qwen and Gemma activation oracles. Forced choice is most accurate when possible answers are known. Bootstrap agreement is calibrated for free text without annotated data.