Abstract:
Hematoxylin and eosin (H&E) staining is ubiquitous in pathology practice
and research. As digital pathology has evolved, the reliance of
quantitative methods that make use of H&E images has similarly expanded.
For example, cell counting and nuclear morphometry rely on the
accurate demarcation of nuclei from other structures and each other.
One of the major obstacles to quantitative analysis of H&E images is
the high degree of variability observed between different samples and
different laboratories. In an effort to characterize this variability,
as well as to provide a substrate that can potentially mitigate this
factor in quantitative image analysis, we developed a technique to
project H&E images into an optimized space more appropriate for many
image analysis procedures. We used a decision tree-based support vector
machine learning algorithm to classify 44 H&E stained whole slide images
of resected breast tumors according to the histological structures that
are present. This procedure takes an H&E image as an input and produces
a classification map of the image that predicts the likelihood of a pixel
belonging to any one of a set of user-defined structures (e.g.,
cytoplasm, stroma). By reducing these maps into their constituent pixels
in color space, an optimal reference vector is obtained for each
structure, which identifies the color attributes that maximally
distinguish one structure from other elements in the image. We show that
tissue structures can be identified using this semi-automated
technique. By comparing structure centroids across different images,
we obtained a quantitative depiction of H&E variability for each
structure. This measurement can potentially be utilized in the
laboratory to help calibrate daily staining or identify troublesome
slides. Moreover, by aligning reference vectors derived from this
technique, images can be transformed in a way that standardizes their
color properties and makes them more amenable to image processing.