Exploring Foundation Models for Visual Feature Extraction from Tissue Images
(2026) KIMM05 20261Department of Immunotechnology
- Abstract
- Spatial omics for comprehensive investigation of the tumor immune microenvironment has the potential to guide treatment and decision-making in the immuno-oncology setting. However, their use in the clinical setting remains limited due to high costs, time consumption and complexity. Image2Count, a deep learning model for deconvoluting molecular expression from low-plex (4 to 5-color) immunofluorescence images was developed to address this issue. The model has demonstrated the ability to predict tumor, immune and stromal cell specific marker expression and the ability to provide detailed predicted tumor profiles. For this study we investigated three foundation models that can be used as visual feature backbones inside the Image2Count... (More)
- Spatial omics for comprehensive investigation of the tumor immune microenvironment has the potential to guide treatment and decision-making in the immuno-oncology setting. However, their use in the clinical setting remains limited due to high costs, time consumption and complexity. Image2Count, a deep learning model for deconvoluting molecular expression from low-plex (4 to 5-color) immunofluorescence images was developed to address this issue. The model has demonstrated the ability to predict tumor, immune and stromal cell specific marker expression and the ability to provide detailed predicted tumor profiles. For this study we investigated three foundation models that can be used as visual feature backbones inside the Image2Count workflow: Deepcell, Kronos and Eva. We measured performance using a Mantle cell lymphoma GeoMx dataset with ”bulk” Region of Interest data from an 1811-plex transcript panel and a 18677-plex transcript panel, as well as non-small cell lung cancer CosMx single-cell resolution data based on a 960-plex transcript panel. Image2Count with foundation model visual feature backbones was able to predict distinct spatial expression patterns, and concordance of pathways enriched in true and predicted data indicates ability to capture biologically relevant information. However, while the investigated foundation models demonstrated strong representational capabilities and cross-panel generalizability, their performance did not consistently surpass smaller specialized models optimized for specific datasets and prediction tasks, highlighting the current limitations of multiplex immunofluorescence foundation models. (Less)
- Popular Abstract
- Every tumor contains a complex environment that influences disease development and treatment. Modern technologies can measure this activity, but are still expensive. This study investigates if AI can do this in a faster and more accessible way.
This study builds on previous research at Lund University that developed Image2Count, a method that can estimate gene activity directly from microscope images of tissue. By analyzing the appearance of cells, the method can predict which genes are active and provide insight into biological processes inside the tissue.
The goal of this project was to investigate whether newer artificial intelligence models could improve Image2Count. These models have been trained on very large image datasets... (More) - Every tumor contains a complex environment that influences disease development and treatment. Modern technologies can measure this activity, but are still expensive. This study investigates if AI can do this in a faster and more accessible way.
This study builds on previous research at Lund University that developed Image2Count, a method that can estimate gene activity directly from microscope images of tissue. By analyzing the appearance of cells, the method can predict which genes are active and provide insight into biological processes inside the tissue.
The goal of this project was to investigate whether newer artificial intelligence models could improve Image2Count. These models have been trained on very large image datasets and may be better at recognizing patterns than the smaller model previously used. Because few studies have compared such models in this type of application, it is still unclear which approaches work best.
Three AI models were evaluated: Kronos, Deepcell, and Eva. Deepcell and Eva combine information from biological texts with image information, while Kronos relies only on what it sees in the images.
The models were tested on three datasets containing measurements of gene activity from tissue samples. The AI models first analyzed the images, and their outputs were then used to predict gene activity.
The results showed clear differences between the models. Kronos was particularly good at recognizing tissue structure and was the most reliable when applied to data it had not previously seen. Deepcell performed well for genes involved in immune responses and provided useful biological insights. Eva achieved strong results when tested on the same type of data it had been trained on. However, when the models were tested on a dataset that had previously been used to evaluate Image2Count, none of them outperformed the simpler model used in earlier work.
Overall, the results suggest that models trained only on images may be better at handling new datasets, while models that combine image and biological information can provide a deeper understanding of biological processes. At the same time, the study indicates that current large AI models still have limitations and are not yet able to consistently outperform smaller models designed for specific biological tasks. (Less)
Please use this url to cite or link to this publication:
https://lup.lub.lu.se/student-papers/record/9238592
- author
- Olsson, Jonathan LU
- supervisor
- organization
- alternative title
- Prediction of Single-Cell Expression from Low-Plex Immunofluorescence
- course
- KIMM05 20261
- year
- 2026
- type
- H2 - Master's Degree (Two Years)
- subject
- keywords
- Deep Learning, Multiplex Imaging, Spatial Biology, Computer Vision, Molecular Oncology, GeoMx, CosMx
- language
- English
- id
- 9238592
- date added to LUP
- 2026-06-16 15:19:38
- date last changed
- 2026-06-16 15:19:38
@misc{9238592,
abstract = {{Spatial omics for comprehensive investigation of the tumor immune microenvironment has the potential to guide treatment and decision-making in the immuno-oncology setting. However, their use in the clinical setting remains limited due to high costs, time consumption and complexity. Image2Count, a deep learning model for deconvoluting molecular expression from low-plex (4 to 5-color) immunofluorescence images was developed to address this issue. The model has demonstrated the ability to predict tumor, immune and stromal cell specific marker expression and the ability to provide detailed predicted tumor profiles. For this study we investigated three foundation models that can be used as visual feature backbones inside the Image2Count workflow: Deepcell, Kronos and Eva. We measured performance using a Mantle cell lymphoma GeoMx dataset with ”bulk” Region of Interest data from an 1811-plex transcript panel and a 18677-plex transcript panel, as well as non-small cell lung cancer CosMx single-cell resolution data based on a 960-plex transcript panel. Image2Count with foundation model visual feature backbones was able to predict distinct spatial expression patterns, and concordance of pathways enriched in true and predicted data indicates ability to capture biologically relevant information. However, while the investigated foundation models demonstrated strong representational capabilities and cross-panel generalizability, their performance did not consistently surpass smaller specialized models optimized for specific datasets and prediction tasks, highlighting the current limitations of multiplex immunofluorescence foundation models.}},
author = {{Olsson, Jonathan}},
language = {{eng}},
note = {{Student Paper}},
title = {{Exploring Foundation Models for Visual Feature Extraction from Tissue Images}},
year = {{2026}},
}