Stain Normalisation, Foundation Encoders and Centre Shift: A Reproducible Cloud-Based Pipeline for Gigapixel Histopathology Analysis in Google Colab

Authors

  • Puneeta Rosmi , Dr.Ajay Agarwal

Keywords:

Whole slide image; stain normalisation; domain shift; foundation model; feature extraction; reproducibility; Google Colab; computational pathology

Abstract

Background: Deep learning on whole slide images (WSIs) is usually developed on dedicated GPU servers with terabytes of local storage, and the preprocessing choices that shape the data a model sees are rarely reported in detail. Stain variation between laboratories is a known source of domain shift, but its effect on frozen foundation-model features has seldom been measured directly. Methods: We built a notebook-based pipeline that runs entirely in Google Colab with Google Drive as the only persistent store. Raw slides are downloaded in batches, segmented by Otsu thresholding of the saturation channel, tiled at twenty times magnification, filtered for blur by the variance of the Laplacian, stain-normalised and encoded once by a frozen encoder, after which only coordinates and features are kept. The pipeline was applied to 13,615 slides from CAMELYON16, CAMELYON17, TCGA-NSCLC, PANDA and BCNB. Three stain conditions (none, Macenko, Vahadane) and two encoders (CTransPath, UNI) were compared through linear probes and through downstream weakly supervised models under cross-validation and leave-one-centre-out evaluation. Results: Tiling produced 35.66 million patches, of which 4.1 per cent were removed as out of focus, leaving 34.2 million. Validation AUC was flat for minimum tissue fractions between 0.40 and 0.60. Vahadane normalisation raised the leave-one-centre-out AUC on CAMELYON17 from 0.874 to 0.911 and the worst centre from 0.829 to 0.889. A linear probe identified the source centre from unnormalised CTransPath features with 88.6 per cent accuracy, falling to 41.3 per cent after normalisation against a chance level of 20 per cent. UNI features improved downstream performance on all five datasets by 0.009 to 0.017. Features for both encoders occupied 122.6 GB, and the full model trained within the memory of a free T4 runtime. Conclusion: Careful preprocessing and stain normalisation remain necessary with foundation encoders, centre information persists after normalisation, and gigapixel pathology studies can be run reproducibly on hosted notebooks.

References

Bisong, E. (2019). Google Colaboratory. In Building machine learning and deep learning models on Google Cloud Platform (pp. 59–64). Apress. https://doi.org/10.1007/978-1-4842-4470-8_7

Bulten, W., Kartasalo, K., Chen, P.-H. C., Ström, P., Pinckaers, H., Nagpal, K., Cai, Y., Steiner, D. F., van Boven, H., Vink, R., Hulsbergen-van de Kaa, C., van der Laak, J., Amin, M. B., Evans, A. J., van der Kwast, T., Allan, R., Humphrey, P. A., Grönberg, H., Samaratunga, H., … Litjens, G. (2022). Artificial intelligence for diagnosis and Gleason grading of prostate cancer: The PANDA challenge. Nature Medicine, 28(1), 154–163. https://doi.org/10.1038/s41591-021-01620-2

Bándi, P., Geessink, O., Manson, Q., Van Dijk, M., Balkenhol, M., Hermsen, M., Ehteshami Bejnordi, B., Lee, B., Paeng, K., Zhong, A., Li, Q., Zanjani, F. G., Zinger, S., Fukuta, K., Komura, D., Ovtcharov, V., Cheng, S., Zeng, S., Thagaard, J., … Litjens, G. (2019). From detection of individual metastases to classification of lymph node status at the patient level: The CAMELYON17 challenge. IEEE Transactions on Medical Imaging, 38(2), 550–560. https://doi.org/10.1109/TMI.2018.2867350

Campanella, G., Hanna, M. G., Geneslaw, L., Miraflor, A., Werneck Krauss Silva, V., Busam, K. J., Brogi, E., Reuter, V. E., Klimstra, D. S., & Fuchs, T. J. (2019). Clinical-grade computational pathology using weakly supervised deep learning on whole slide images. Nature Medicine, 25(8), 1301–1309. https://doi.org/10.1038/s41591-019-0508-1

Carneiro, T., Medeiros Da Nóbrega, R. V., Nepomuceno, T., Bian, G.-B., De Albuquerque, V. H. C., & Rebouças Filho, P. P. (2018). Performance analysis of Google Colaboratory as a tool for accelerating deep learning applications. IEEE Access, 6, 61677–61685. https://doi.org/10.1109/ACCESS.2018.2874767

Chen, R. J., Ding, T., Lu, M. Y., Williamson, D. F. K., Jaume, G., Song, A. H., Chen, B., Zhang, A., Shao, D., Shaban, M., Williams, M., Oldenburg, L., Weishaupt, L. L., Wang, J. J., Vaidya, A., Le, L. P., Gerber, G., Sahai, S., Williams, W., & Mahmood, F. (2024). Towards a general-purpose foundation model for computational pathology. Nature Medicine, 30(3), 850–862. https://doi.org/10.1038/s41591-024-02857-3

Ciompi, F., Geessink, O., Ehteshami Bejnordi, B., de Souza, G. S., Baidoshvili, A., Litjens, G., van Ginneken, B., Nagtegaal, I., & van der Laak, J. (2017). The importance of stain normalization in colorectal tissue classification with convolutional networks. In 2017 IEEE 14th International Symposium on Biomedical Imaging (pp. 160–163). IEEE. https://doi.org/10.1109/ISBI.2017.7950492

Grossman, R. L., Heath, A. P., Ferretti, V., Varmus, H. E., Lowy, D. R., Kibbe, W. A., & Staudt, L. M. (2016). Toward a shared vision for cancer genomic data. New England Journal of Medicine, 375(12), 1109–1112. https://doi.org/10.1056/NEJMp1607591

Howard, F. M., Dolezal, J., Kochanny, S., Schulte, J., Chen, H., Heij, L., Huo, D., Nanda, R., Olopade, O. I., Kather, J. N., Cipriani, N., Grossman, R. L., & Pearson, A. T. (2021). The impact of site-specific digital histology signatures on deep learning model accuracy and bias. Nature Communications, 12, Article 4423. https://doi.org/10.1038/s41467-021-24698-1

Kleppe, A., Skrede, O.-J., De Raedt, S., Liestøl, K., Kerr, D. J., & Danielsen, H. E. (2021). Designing deep learning studies in cancer diagnostics. Nature Reviews Cancer, 21(3), 199–211. https://doi.org/10.1038/s41568-020-00327-9

Litjens, G., Bándi, P., Ehteshami Bejnordi, B., Geessink, O., Balkenhol, M., Bult, P., Halilovic, A., Hermsen, M., van de Loo, R., Vogels, R., Manson, Q. F., Stathonikos, N., Baidoshvili, A., van Diest, P., Wauters, C., van Dijk, M., & van der Laak, J. (2018). 1399 H&E-stained sentinel lymph node sections of breast cancer patients: The CAMELYON dataset. GigaScience, 7(6), Article giy065. https://doi.org/10.1093/gigascience/giy065

Lu, M. Y., Williamson, D. F. K., Chen, T. Y., Chen, R. J., Barbieri, M., & Mahmood, F. (2021). Data-efficient and weakly supervised computational pathology on whole-slide images. Nature Biomedical Engineering, 5(6), 555–570. https://doi.org/10.1038/s41551-020-00682-w

Macenko, M., Niethammer, M., Marron, J. S., Borland, D., Woosley, J. T., Guan, X., Schmitt, C., & Thomas, N. E. (2009). A method for normalizing histology slides for quantitative analysis. In 2009 IEEE International Symposium on Biomedical Imaging: From Nano to Macro (pp. 1107–1110). IEEE. https://doi.org/10.1109/ISBI.2009.5193250

Otsu, N. (1979). A threshold selection method from gray-level histograms. IEEE Transactions on Systems, Man, and Cybernetics, 9(1), 62–66. https://doi.org/10.1109/TSMC.1979.4310076

Pech-Pacheco, J. L., Cristóbal, G., Chamorro-Martínez, J., & Fernández-Valdivia, J. (2000). Diatom autofocusing in brightfield microscopy: A comparative study. In Proceedings of the 15th International Conference on Pattern Recognition (Vol. 3, pp. 314–317). IEEE. https://doi.org/10.1109/ICPR.2000.903548

Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d'Alché-Buc, F., Fox, E., & Larochelle, H. (2021). Improving reproducibility in machine learning research. Journal of Machine Learning Research, 22(164), 1–20.

Pocock, J., Graham, S., Vu, Q. D., Jahanifar, M., Deshpande, S., Hadjigeorghiou, G., Shephard, A., Bashir, R. M. S., Bilal, M., Lu, W., Epstein, D., Minhas, F., Rajpoot, N. M., & Raza, S. E. A. (2022). TIAToolbox as an end-to-end library for advanced tissue image analytics. Communications Medicine, 2, Article 120. https://doi.org/10.1038/s43856-022-00186-5

Song, A. H., Jaume, G., Williamson, D. F. K., Lu, M. Y., Vaidya, A., Miller, T. R., & Mahmood, F. (2023). Artificial intelligence for digital and computational pathology. Nature Reviews Bioengineering, 1(12), 930–949. https://doi.org/10.1038/s44222-023-00096-8

Stacke, K., Eilertsen, G., Unger, J., & Lundström, C. (2021). Measuring domain shift for deep learning in histopathology. IEEE Journal of Biomedical and Health Informatics, 25(2), 325–336. https://doi.org/10.1109/JBHI.2020.3032060

Tellez, D., Litjens, G., Bándi, P., Bulten, W., Bokhorst, J.-M., Ciompi, F., & van der Laak, J. (2019). Quantifying the effects of data augmentation and stain color normalization in convolutional neural networks for computational pathology. Medical Image Analysis, 58, Article 101544. https://doi.org/10.1016/j.media.2019.101544

Vahadane, A., Peng, T., Sethi, A., Albarqouni, S., Wang, L., Baust, M., Steiger, K., Schlitter, A. M., Esposito, I., & Navab, N. (2016). Structure-preserving color normalization and sparse stain separation for histological images. IEEE Transactions on Medical Imaging, 35(8), 1962–1971. https://doi.org/10.1109/TMI.2016.2529665

van der Maaten, L., & Hinton, G. (2008). Visualizing data using t-SNE. Journal of Machine Learning Research, 9(86), 2579–2605.

Vorontsov, E., Bozkurt, A., Casson, A., Shaikovski, G., Zelechowski, M., Severson, K., Zimmermann, E., Hall, J., Tenenholtz, N., Fusi, N., Yang, E., Mathieu, P., van Eck, A., Lee, D., Viret, J., Robert, E., Wang, Y. K., Kunz, J. D., Lee, M. C. H., … Fuchs, T. J. (2024). A foundation model for clinical-grade computational pathology and rare cancers detection. Nature Medicine, 30(10), 2924–2935. https://doi.org/10.1038/s41591-024-03141-0

Wang, X., Yang, S., Zhang, J., Wang, M., Zhang, J., Yang, W., Huang, J., & Han, X. (2022). Transformer-based unsupervised contrastive learning for histopathological image classification. Medical Image Analysis, 81, Article 102559. https://doi.org/10.1016/j.media.2022.102559

Xu, F., Zhu, C., Tang, W., Wang, Y., Zhang, Y., Li, J., Jiang, H., Shi, Z., Liu, J., & Jin, M. (2021). Predicting axillary lymph node metastasis in early breast cancer using deep learning on primary tumor biopsy slides. Frontiers in Oncology, 11, Article 759007. https://doi.org/10.3389/fonc.2021.759007

Xu, H., Usuyama, N., Bagga, J., Zhang, S., Rao, R., Naumann, T., Wong, C., Gero, Z., González, J., Gu, Y., Xu, Y., Wei, M., Wang, W., Ma, S., Wei, F., Yang, J., Li, C., Gao, J., Rosemon, J., … Poon, H. (2024). A whole-slide foundation model for digital pathology from real-world data. Nature, 630(8015), 181–188. https://doi.org/10.1038/s41586-024-07441-w

Downloads

How to Cite

Puneeta Rosmi , Dr.Ajay Agarwal. (2025). Stain Normalisation, Foundation Encoders and Centre Shift: A Reproducible Cloud-Based Pipeline for Gigapixel Histopathology Analysis in Google Colab. International Journal of Engineering Science & Humanities, 15(3), 663–681. Retrieved from https://www.ijesh.com/j/article/view/1217

Similar Articles

<< < 18 19 20 21 22 23 24 > >> 

You may also start an advanced similarity search for this article.