Skip to content

Glossary and References

Glossary

Bayesian Information Criterion (BIC) Score
The BIC score is a measure of the log-likelihood of all of the points in the data set being correctly classified. Each data point is assigned a Gaussian point probability based on the parameters of the cluster it is assigned to, and a score for each cluster is then computed. The BIC score of the entire data set is the sum of the BIC scores for all clusters in that data set. Given two clustering models for the same data set, the model with the higher BIC score is preferred. [1]

Cluster-level outliers
After the data has been sorted into different clusters, outliers between clusters can be eliminated. The logic used is the same as that at the set-level where for each cluster the Mahalanobis distances of all the cluster points from the cluster center are computed. The boundary then set around the cluster is a function of n standard deviations of those cluster distances away from the mean. Cluster points that lie beyond their cluster boundary get eliminated at this stage as cluster-level outliers.

Data-set-level outlier
Spike shapes that are drastically different from the majority of the spike shapes present in the data tend to lie away from the center of the corresponding feature space representation. These spike shapes can be excluded from being allocated to a cluster by assigning them as outliers prior to the sort process. The Mahalanobis distances of the data points are computed from the center of the data set. The mean and the standard deviation of these distances are then used to set a boundary around the dataset, such that any point lying beyond that boundary gets classified as an outlier. This boundary is set as a function of n standard deviations away from the mean. Data points classified as set-level outliers are not considered when an automated or semi-automated sorting algorithm is run.

Euclidean distance
The straight-line distance between two points.

Feature space
An abstract space where each event is represented as a point in n-dimensional space. Each measurement ("feature") about the event gives the coordinate of the point along one axis of the space. The dimensionality of the feature space is equal to the number of features used to describe the event.

Isolation Distance
A Mahalanobis distance measure of the nearness of a cluster to the non-cluster points. The greater the value of the isolation distance, the better the separation. The Isolation distance for the largest cluster is not defined if the number of points in the largest cluster is greater than the sum of the non-cluster points. [5]

L-ratio
A statistical measure of how separated a given cluster is from other clusters. It is the normalized sum of the probabilities with which non-cluster points belong to that cluster. Ideally it should have a value of zero. Since the L-Ratio compares separation between clusters, it cannot be computed for single cluster data sets. [5]

Mahalanobis distance
The Mahalanobis distance is the straight-line distance between two points weighted by the inverse of the variance in the data set/cluster. Real-world noise that is responsible for the variance in a cluster is typically Gaussian, and Gaussian distributions form elongated/elliptical clusters. In the feature space if the data has a greater variance along one axis as compared to the others, weighting the data by the inverse of its variance along each axis ensures that data points along the direction of the maximum variance axis are not at a disadvantage just because they are further away from the cluster center as compared to those data points that are closer to the center but are along a direction with lesser variance. Hence, using Mahalanobis distances ensures an elliptical boundary around the cluster, and data points further away from the cluster center but along the direction of its maximum variance get included within that elliptical boundary.

Offline sort
Sort codes calculated and associated with events after data acquisition is complete.

Online sort
Sort codes associated with an event during data collection.

Principal component
Principal components are a multi-dimensional representation of the data set where the first principal component is that parameter that represents the maximum variance in the spike shapes; the second principal component represents the next highest variance in the spike data and so on. Transforming spike data in terms of its principal components allows representing the data in parameters that best describe the differences within the data set while also reducing its dimensionality. The number of principal components that can be obtained for a data set equals the number of sample points in each waveform. However, for data sets with well-defined spike shapes, the first two-three principal components represent about 90-95% of the variance in the data set. All the principal components taken together represent 100% of the variance. Increasing the number of principal components to represent the data may lead to a higher computation cost without a commensurate increase in the percentage variance.

Pseudo F-stat (PFS)
A statistical measure that represents a scaled version of the sum of the variance values between clusters over the sum of the variance values calculated within each cluster. This measure is used to help determine the optimum number of clusters. The higher the value, the greater the separation between clusters. [4]

Silhouette Index
A statistical measure that indicates how well a data point has been classified to a cluster. The measurement is made in terms of its average Euclidean distance from the cluster it is assigned to, with respect to the minimum of the average distances from each of the other clusters in the data set. The Silhouette Indices can range from -1 to 1, where 1 indicates a good classification, -1 indicates a bad classification and 0 indicates that the classification could go either way.

References

  1. Pelleg D., Moore A., "X-means: Extending K-means with efficient estimation of the number of clusters," in ICML 2000.

  2. Hamerly G., Elkan C., Learning the k in k-means. In proceedings of the seventeenth annual conference on neural information processing systems (NIPS), 281-288, December 2003. (Older UCSD technical report CS2002-0716).

  3. Lewicki M.S., A review of methods for spike sorting: the detection and classification of neural action potentials. Network: Computation in Neural Systems, 9 (4): 53-78, 1998.

  4. Devore J., Peck R., Statistics: The exploration and analysis of data.

  5. Schmitzer-Torbert N., Jackson J., Henze D., Harris K.D., Redish A.D. (2005) Quantitative measures of cluster quality for use in extracellular recordings. Neuroscience, 131:1-11.

  6. Bolshakova N., Azuaje F.: Cluster validation techniques for genome expression data. Signal Processing (2003) 825-833.

  7. Wheeler B.C., Automatic Discrimination of Single Units in Methods for Neural Ensemble Recordings, ed. by Nicolelis, M., CRC Press, 1999.