Glossary and References
Glossary
Bayesian Information Criterion (BIC) Score
The BIC score is a measure of the log-likelihood of all of the points in the data set being correctly classified.
Each data point is assigned a Gaussian point probability based on the parameters of the cluster it is assigned to,
and a score for each cluster is then computed. The BIC score of the entire data set is the sum of the BIC scores
for all clusters in that data set. Given two clustering models for the same data set, the model with the higher
BIC score is preferred. [1]
Cluster-level outliers
After the data has been sorted into different clusters, outliers between clusters can be eliminated. The logic
used is the same as that at the set-level where for each cluster the Mahalanobis distances of all the cluster
points from the cluster center are computed. The boundary then set around the cluster is a function of n standard
deviations of those cluster distances away from the mean. Cluster points that lie beyond their cluster boundary
get eliminated at this stage as cluster-level outliers.
Data-set-level outlier
Spike shapes that are drastically different from the majority of the spike shapes present in the data tend to lie
away from the center of the corresponding feature space representation. These spike shapes can be excluded from
being allocated to a cluster by assigning them as outliers prior to the sort process. The Mahalanobis distances of
the data points are computed from the center of the data set. The mean and the standard deviation of these
distances are then used to set a boundary around the dataset, such that any point lying beyond that boundary gets
classified as an outlier. This boundary is set as a function of n standard deviations away from the mean. Data
points classified as set-level outliers are not considered when an automated or semi-automated sorting algorithm
is run.
Euclidean distance
The straight-line distance between two points.
Feature space
An abstract space where each event is represented as a point in n-dimensional space. Each measurement ("feature")
about the event gives the coordinate of the point along one axis of the space. The dimensionality of the feature
space is equal to the number of features used to describe the event.
Isolation Distance
A Mahalanobis distance measure of the nearness of a cluster to the non-cluster points. The greater the value of
the isolation distance, the better the separation. The Isolation distance for the largest cluster is not defined
if the number of points in the largest cluster is greater than the sum of the non-cluster points. [5]
L-ratio
A statistical measure of how separated a given cluster is from other clusters. It is the normalized sum of the
probabilities with which non-cluster points belong to that cluster. Ideally it should have a value of zero. Since
the L-Ratio compares separation between clusters, it cannot be computed for single cluster data sets. [5]
Mahalanobis distance
The Mahalanobis distance is the straight-line distance between two points weighted by the inverse of the variance
in the data set/cluster. Real-world noise that is responsible for the variance in a cluster is typically Gaussian,
and Gaussian distributions form elongated/elliptical clusters. In the feature space if the data has a greater
variance along one axis as compared to the others, weighting the data by the inverse of its variance along each
axis ensures that data points along the direction of the maximum variance axis are not at a disadvantage just
because they are further away from the cluster center as compared to those data points that are closer to the
center but are along a direction with lesser variance. Hence, using Mahalanobis distances ensures an elliptical
boundary around the cluster, and data points further away from the cluster center but along the direction of its
maximum variance get included within that elliptical boundary.
Offline sort
Sort codes calculated and associated with events after data acquisition is complete.
Online sort
Sort codes associated with an event during data collection.
Principal component
Principal components are a multi-dimensional representation of the data set where the first principal component is
that parameter that represents the maximum variance in the spike shapes; the second principal component represents
the next highest variance in the spike data and so on. Transforming spike data in terms of its principal components
allows representing the data in parameters that best describe the differences within the data set while also
reducing its dimensionality. The number of principal components that can be obtained for a data set equals the
number of sample points in each waveform. However, for data sets with well-defined spike shapes, the first
two-three principal components represent about 90-95% of the variance in the data set. All the principal components
taken together represent 100% of the variance. Increasing the number of principal components to represent the data
may lead to a higher computation cost without a commensurate increase in the percentage variance.
Pseudo F-stat (PFS)
A statistical measure that represents a scaled version of the sum of the variance values between clusters over the
sum of the variance values calculated within each cluster. This measure is used to help determine the optimum
number of clusters. The higher the value, the greater the separation between clusters. [4]
Silhouette Index
A statistical measure that indicates how well a data point has been classified to a cluster. The measurement is
made in terms of its average Euclidean distance from the cluster it is assigned to, with respect to the minimum of
the average distances from each of the other clusters in the data set. The Silhouette Indices can range from -1 to
1, where 1 indicates a good classification, -1 indicates a bad classification and 0 indicates that the
classification could go either way.
References
-
Pelleg D., Moore A., "X-means: Extending K-means with efficient estimation of the number of clusters," in ICML 2000.
-
Hamerly G., Elkan C., Learning the k in k-means. In proceedings of the seventeenth annual conference on neural information processing systems (NIPS), 281-288, December 2003. (Older UCSD technical report CS2002-0716).
-
Lewicki M.S., A review of methods for spike sorting: the detection and classification of neural action potentials. Network: Computation in Neural Systems, 9 (4): 53-78, 1998.
-
Devore J., Peck R., Statistics: The exploration and analysis of data.
-
Schmitzer-Torbert N., Jackson J., Henze D., Harris K.D., Redish A.D. (2005) Quantitative measures of cluster quality for use in extracellular recordings. Neuroscience, 131:1-11.
-
Bolshakova N., Azuaje F.: Cluster validation techniques for genome expression data. Signal Processing (2003) 825-833.
-
Wheeler B.C., Automatic Discrimination of Single Units in Methods for Neural Ensemble Recordings, ed. by Nicolelis, M., CRC Press, 1999.