Modular post-hoc enhancements for zero-shot histopathology classification using vision-language models
Résumé
Abstract Zero-shot learning with vision-language models (VLMs) has shown promising results in histopathological image classification. However, existing approaches often under-utilize domain-specific adaptations that could enhance performance in specialized medical tasks. In this study, we systematically evaluate three modular, post-hoc enhancements to standard zero-shot VLM pipelines: (1) cosine similarity calibration for affinity matrix construction, (2) prompt template enhancement via diverse clinical prompts, and (3) adaptive hyperparameter tuning using a Gaussian mixture model-based method for adjusting pseudo-labeling weights (λ) and support set contributions (γ), following Zanella et al We conduct extensive zero-shot experiments on the NCT-CRC-HE-100 K dataset, using 60 000 patches for validation of inference-time hyperparameters and reserving 40 000 patches as an independent test set. No training or fine-tuning was performed. The evaluation was carried out across five VLMs: CLIP, CONCH, PLIP, Quilt-B16, and Quilt-B32. Results show that each enhancement module yields significant accuracy gains. Cosine calibration improves CLIP (+7.8%), CONCH (+1.53%), PLIP (+14.3%), Quilt-B16 (+16.1%), and Quilt-B32 (+11.4%). Prompt enhancement yields accuracy gains in PLIP (+16.8%) and Quilt-B32 (+29.9%), with limited effect on CLIP or CONCH. Adaptive tuning yields large improvements in PLIP (+28.4%). The best performance is achieved with the full combination of enhancements, with PLIP reaching 87.88% (+27.88%, p < 0.001) and Quilt-B32 reaching 88.38% (+36.1%) over their baselines. CLIP showed significance only with cosine calibration (p < 0.001), reflecting its contrastive learning backbone. Although CONCH reached the highest raw accuracy (92.63%) under the full enhancement setting, this performance appears primarily driven by its histopathology-specific pretraining; statistically significant gains emerged only in cosine-based combinations, indicating limited responsiveness to prompt or tuning. Our findings highlight the critical role of structured modular enhancements in optimizing VLM performance for specialized clinical domains.
Citer ce document
Accès au document
Texte intégral en lecture en ligne, réservé aux abonnés SPHAERO et aux membres de l'institution. Se connecter
Voir l'article sur le site de la revueAuteur(s)
Statistiques
Consultations : 1
Téléchargements : 0