Glossary

Cluster analysis

Definition

Cluster analysis divides observations into groups (clusters) so that units within a cluster are as similar as possible and clusters differ from each other — without predefined classes. Hierarchical clustering builds a tree (dendrogram); k-means assigns observations to a given number of centres. Typical uses are customer segmentation, product typologies or pattern discovery in measurements.

Multivariate methods

The result depends on the distance measure (Euclidean, Manhattan, correlation), the linkage method (Ward, average) and on variable scaling — standardize, otherwise wide-range variables dominate. For k-means, choose the number of clusters via the within-cluster variability plot or silhouette.

Clusters are a descriptive construct, not proof: validate them by stability on another sample and by meaningful profiles (variable means within clusters). For large datasets k-means or EM clustering are preferable to hierarchical methods.

In Statistica

Statistics → Multivariate Exploratory Techniques → Cluster Analysis contains hierarchical clustering with a dendrogram, k-means with a plot of means and clustering of variables; for big data and automatic choice of cluster count use Generalized EM & k-Means Clustering in Data Mining with cross-validation.

Related terms

Knowledgebase guides

FAQ

How many clusters is right?
There is no single correct answer — combine statistical criteria (silhouette, the elbow in variability) with whether the clusters are interpretable and usable in practice.
What is the difference between cluster and discriminant analysis?
Cluster analysis searches for groups (unsupervised learning); discriminant analysis separates and predicts groups known in advance (supervised learning).

Try it on your own data

Statistica free for 30 days

Full version, no credit card. Or get a pricing estimate in a minute.

We'll prepare a tailored quote. Free and with no obligation.

Tell us how many users and what analyses you need — we'll get back to you with a concrete license quote, usually within a few business days.

Try for free first
Statistica.pro

Official Statistica partner for the European Union, based in Prague, Czech Republic. Operated by DataBon s.r.o.

On the market since 1995, formerly as StatSoft CR s.r.o.

Contact

DataBon s.r.o.
Korunní 2569/108, 101 00 Prague 10, Czech Republic
+420 602 284 038
Company ID 09743804 · VAT ID CZ09743804

University licence users: please contact the licence administrator at your institution first.

© 2026 DataBon · statistica.pro · Personal data processing (GDPR) ·