Saliency maps have become the public face of artificial intelligence transparency. When a convolutional neural network declares that a photograph contains a tiger shark, a colorful heatmap shows which pixels supposedly drove that decision, offering engineers, regulators, and end users a reassuring window into the model’s reasoning. But a new study from the University of Virginia, published in the journal Machine Learning, suggests that this window may be far murkier than it appears. The research demonstrates that many of the most widely used explanation methods produce nearly identical heatmaps regardless of which class label they are asked to explain, raising uncomfortable questions about whether these visual justifications actually explain anything at all.
The study, led by doctoral researcher Dane Williamson together with Yangfeng Ji and Matthew Dwyer, begins with a deceptively simple observation. When the popular Grad-CAM method is asked to explain why a network chose its top prediction for an image, and then asked again to explain the second-ranked prediction, the highlighted regions barely change. This happens even when the two candidate labels describe semantically unrelated objects, such as an airplane and a barbell. If an explanation method highlights the same pixels whether the model claims to see a bird or an airplane, the researchers argue, that explanation is not telling anyone what distinguishes one hypothesis from another. It is, in effect, a plausible-looking picture with little diagnostic content.
The stakes are higher than they might seem. Prior human-centered research cited in the paper, including the HIVE evaluation study, found that visual explanations can increase user trust even when the underlying model is wrong. In other words, a beautiful heatmap can manufacture confidence in a decision that the heatmap does not actually illuminate. The problem becomes acute in domains such as medical imaging, where documented failures have shown networks latching onto dataset artifacts like text tokens or positional cues instead of actual pathology. If the saliency maps generated in those settings are insensitive to the class label, clinicians auditing a diagnosis could be misled by an explanation that looks informative but is essentially identical to the explanation the model would have produced for any other condition.
To move beyond anecdotes, the researchers formalized a diagnostic test for what they call class sensitivity. For each validation image, they generated saliency maps for the top-1 and top-2 predicted classes and measured the overlap between the top 5 percent of pixels highlighted in each map. A method that genuinely distinguishes between competing labels should produce substantially different maps. They then applied a one-sided Wilcoxon signed-rank test, asking whether the median agreement between maps fell below 50 percent. If a method cannot statistically beat a coin-flip level of overlap, it fails the test of distinguishing one class explanation from another. This framework is deliberately method-agnostic and model-agnostic, meaning it can be applied to any saliency technique and any convolutional backbone.
The results of that diagnostic were sobering. Across five widely used class activation map methods, Grad-CAM, Grad-CAM++, ScoreCAM, AblationCAM, and LayerCAM, evaluated on four architectures spanning ResNet-50, VGG-19, DenseNet-201, and ConvNeXt-Large, and on two datasets, ImageNet and CIFAR-100, most baseline methods failed to distinguish between competing classes in a large fraction of settings. Grad-CAM and AblationCAM in particular frequently produced explanations for the top and second predictions that were statistically indistinguishable. The failure was not confined to one model family or benchmark, suggesting a systemic weakness in how activation-based attribution is computed rather than an artifact of any single network.
Motivated by these findings, the team introduced CASE, short for Contrastive Activation for Class-Sensitive Explanations. The method works by explicitly removing the attribution that a target class shares with its most confusable competitors. Concretely, CASE first computes the gradient of the target class score with respect to the final convolutional activation maps. It then identifies a contrast set consisting of the classes the model most frequently confuses with the target, using the model’s confusion matrix computed on a held-out validation set. The gradients for those contrast classes are averaged into a single vector, and the component of the target gradient that points along this shared direction is subtracted away through an orthogonal projection. What remains is a residual gradient encoding only evidence that is uniquely discriminative for the target class, which is then used to weight the activation maps and produce the final heatmap after a ReLU and upsampling step.
The design choices behind CASE are notable for what they avoid. Unlike prior contrastive explanation approaches such as CWOX, which relies on learned cluster structures, or Contrastive LRP, which is tied to backpropagation-based relevance propagation, CASE requires no architectural modifications, no external baselines, and no auxiliary models. It operates directly on the internal activations of any already-trained convolutional network. The confusion-based contrast selection is also behaviorally grounded: because the contrast set is drawn from classes the model empirically mixes up, the method suppresses exactly the attribution directions that compete with the target in practice, rather than those suggested by semantic similarity alone.
In evaluations across all four architectures and both datasets, CASE was the only method that consistently rejected the null hypothesis of the class-sensitivity test with high confidence, achieving p-values below 0.0001 in every tested setting. ScoreCAM also performed well on most models but stumbled on VGG, while traditional methods failed more often. The study also uncovered a surprising architectural insight: DenseNet proved the most naturally class-sensitive network, and a channel-activity analysis revealed why. Its final convolutional layer activates fewer than 20 channels per class on average, an unusually sparse but highly selective representation, whereas VGG and ConvNeXt activate hundreds of channels that dilute class-specific signals. When the researchers swapped DenseNet’s attribution layer to the final normalization layer, which activates nearly all channels uniformly, most methods lost their class sensitivity entirely. Even retraining across five random seeds confirmed the effect was a stable property of the architecture, and CASE retained its sensitivity even in the diffuse regime.
Class sensitivity alone is not sufficient, of course. A good explanation should also be faithful, meaning the highlighted regions should genuinely matter to the model’s output. The team tested this by ablating the top 5 percent of salient pixels and measuring the drop in the model’s confidence. Here the picture was more nuanced. CASE matched the strongest baselines on DenseNet, ResNet, and ConvNeXt, and significantly outperformed AblationCAM and LayerCAM on ConvNeXt, but on VGG three methods, Grad-CAM++, LayerCAM, and ScoreCAM, produced larger confidence drops. The authors are candid about these failure cases, noting that uniquely discriminative regions are not always the regions whose removal maximally reduces confidence. When class evidence is diffuse and shared across labels, perturbing broadly activated regions can produce bigger confidence swings even if those regions are less informative for separating one class from another.
The paper also probes the robustness of the method itself. A pre-registered degenerate-residual check searched for cases where the target and contrast gradients were nearly collinear, which would collapse the orthogonal projection toward zero; across 4,774 generated maps, not a single near-zero residual was found. Ablations over the contrast set size showed that three confused classes are sufficient, with larger sets yielding no consistent benefit. The authors are explicit about the method’s boundaries: CASE applies to single-label classification with convolutional networks, and extending it to transformers, multi-label tasks, or multimodal models would require rethinking the contrast formulation. Those limitations aside, the study delivers a message that should resonate far beyond the machine learning community. Explanation tools are trusted as safeguards, and this work shows that trust must be earned through formal diagnostics rather than visual plausibility. As deep learning spreads into medicine, safety monitoring, and high-stakes auditing, the researchers argue, we need explanations that reveal what truly sets a model’s decision apart, not just what makes a convincing picture.
Subject of Research: Class sensitivity and contrastive saliency methods for explaining convolutional neural network image classifiers
Article Title: CASE: Contrastive Activation for Class-Sensitive Explanations
Article References: Williamson, D., Ji, Y., & Dwyer, M. (2026). CASE: Contrastive Activation for Class-Sensitive Explanations. Machine Learning, 115(10), Article 230. https://doi.org/10.1007/s10994-026-07169-w
Image Credits: AI Generated
DOI: 10.1007/s10994-026-07169-w
Keywords: explainable AI, saliency maps, Grad-CAM, contrastive explanations, convolutional neural networks, class activation mapping, machine learning, attribution fidelity, ImageNet, CIFAR-100, DenseNet, model interpretability
Cite Scienmag News
Blake Davidson. (October 1, 2026). AI Heatmaps May Be Lying: New Test Exposes Flawed Explanations in Image-Recognition Networks. Scienmag. https://scienmag.com/ai-heatmaps-may-be-lying-new-test-exposes-flawed-explanations-in-image-recognition-networks/
Blake Davidson. "AI Heatmaps May Be Lying: New Test Exposes Flawed Explanations in Image-Recognition Networks." Scienmag, 1 October 2026, https://scienmag.com/ai-heatmaps-may-be-lying-new-test-exposes-flawed-explanations-in-image-recognition-networks/. Accessed 1 October 2026.
Blake Davidson. "AI Heatmaps May Be Lying: New Test Exposes Flawed Explanations in Image-Recognition Networks." Scienmag. October 1, 2026. https://scienmag.com/ai-heatmaps-may-be-lying-new-test-exposes-flawed-explanations-in-image-recognition-networks/

