Architectural Enhancements to HoVer-Net for Automated Nuclear Segmentation and Classification

Authors

DOI:

https://doi.org/10.62978/2601qsz120

Keywords:

Nuclear Instance Segmentation, Histopathology Image Analysis, Deep Learning, HoVer-Net

Abstract

Nuclear instance segmentation and classification support quantitative computational pathology by connecting cell boundaries with biologically meaningful labels. This exploratory study implemented five architectural variants of HoVer-Net on PanNuke, spanning channel recalibration, encoder attention, decoder refinement, and two transformer encoder integrations. The variants were evaluated using a shared reported image-patch partition and core metric pipeline, with one reported training run per configuration. The composite CNN configuration, MSDHV-Net, achieved inflammatory type-map Dice of 0.7162 versus 0.6728 for the reproduced HoVer-Net baseline, and merged Other type-map Dice of 0.7241 versus 0.6939. The SE-only variant yielded slightly higher neoplastic and inflammatory type-map Dice than the composite, while MSDHV-Net's neoplastic instance PQ was 0.3825 versus the baseline's 0.4005. In the reported setup, both transformer encoder substitutions had lower nuclear-pixel Dice than the reproduced baseline; HoverViTNet also showed a higher inflammatory type-map Dice than that baseline. Together, these point estimates describe how five implementation choices distribute performance across pixel overlap, cell type, and the available instance metrics. The study contributes an organized experimental comparison that identifies class-specific directions and concrete priorities for repeated, source-aware evaluation.

References

1. Sirinukunwattana, K., Snead, D., Epstein, D., Aftab, Z., Mujeeb, I., Tsang, Y. W., Cree, I., & Rajpoot, N. (2018). Novel digital signatures of tissue phenotypes for predicting distant metastasis in colorectal cancer. Scientific Reports, 8, Article 13692.

2. Javed, S., Fraz, M. M., Epstein, D., Snead, D., & Rajpoot, N. M. (2018). Cellular community detection for tissue phenotyping in histology images. In Computational Pathology and Ophthalmic Medical Image Analysis (pp. 120–129). Springer.

3. Ronneberger, O., Fischer, P., & Brox, T. (2015). U-Net: Convolutional networks for biomedical image segmentation. In Lecture Notes in Computer Science (pp. 234–241). Springer.

4. Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J., & Maier-Hein, K. H. (2021). nnU-Net: A self-configuring method for deep learning-based biomedical image segmentation. Nature Methods, 18, 203–211.

5. Graham, S., Vu, Q. D., Raza, S. E., Azam, A., Tsang, Y. W., Kwak, J. T., & Rajpoot, N. (2019). HoVer-Net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images. Medical Image Analysis, 58, Article 101563.

6. Gamper, J., Koohbanani, N. A., Graham, S., Jahanifar, M., Khurram, S. A., Azam, A., Hewitt, K., & Rajpoot, N. (2020). PanNuke dataset extension, insights and baselines. arXiv. https://arxiv.org/abs/2003.10778

7. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., et al. (2020). An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv. https://arxiv.org/abs/2010.11929

8. Liu, Z., Lin, Y., Cao, Y., Hu, H., Wei, Y., Zhang, Z., et al. (2021). Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 10012–10022).

9. Hörst, F., Rempe, M., Heine, L., Seibold, C., Keyl, J., Baldini, G., et al. (2024). CellViT: Vision transformers for precise cell segmentation and classification. Medical Image Analysis, 94, Article 103143.

10. Xu, H., Xu, Q., Cong, F., Kang, J., Han, C., Liu, Z., Madabhushi, A., & Lu, C. (2023). Vision Transformers for Computational Histopathology. IEEE Reviews in Biomedical Engineering, 17, 63-79.

11. Hu, J., Shen, L., & Sun, G. (2018). Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 7132–7141).

12. Zhang, Z. (2025). Enhancing distributed machine learning through data shuffling: Techniques, challenges, and implications. ITM Web of Conferences, 73, Article 03018.

13. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.

14. Voita, E., Talbot, D., Moiseev, F., Sennrich, R., & Titov, I. (2019). Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. arXiv. https://arxiv.org/abs/1905.09418

15. Mao, X., Shen, C., & Yang, Y. B. (2016). Image restoration using very deep convolutional encoder-decoder networks with symmetric skip connections. Advances in Neural Information Processing Systems, 29, 2802–2810.

16. Wang, Z., Li, T., Zheng, J. Q., & Huang, B. (2022). When CNN meets ViT: Towards semi-supervised learning for multi-class medical image semantic segmentation. In Proceedings of the European Conference on Computer Vision (pp. 424–441). Springer.

17. Yang, J., Li, C., Zhang, P., Dai, X., Xiao, B., Yuan, L., & Gao, J. (2021). Focal self-attention for local-global interactions in vision transformers. arXiv. https://arxiv.org/abs/2107.00641

18. Hatamizadeh, A., Nath, V., Tang, Y., Yang, D., Roth, H. R., & Xu, D. (2021). Swin UNETR: Swin transformers for semantic segmentation of brain tumors in MRI images. In International MICCAI Brainlesion Workshop (pp. 272–284). Springer.

19. Kingma, D. P., & Ba, J. (2014). Adam: A method for stochastic optimization. arXiv. https://arxiv.org/abs/1412.6980

20. He, K., Zhang, X., Ren, S., & Sun, J. (2016). Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (pp. 770–778).

21. Loshchilov, I., & Hutter, F. (2017). Decoupled weight decay regularization. arXiv. https://arxiv.org/abs/1711.05101

22. Kirillov, A., He, K., Girshick, R., Rother, C., & Dollár, P. (2019). Panoptic segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 9404–9413).

Downloads

Published

2026-09-29

Issue

Section

文章