Glossary

Text mining

Definition

Text mining converts unstructured text — complaints, service notes, open survey answers, reports — into structured data that statistics can work with. Documents are broken into words and word stems, frequencies and weighted matrices (TF-IDF) are computed, and clustering, classification or topic discovery run on top of them.

Data mining & machine learning

Typical tasks: automatic sorting of complaints by cause, finding recurring problems in service records, sentiment analysis of customer reviews or linking text with numeric data (how much do defects described by certain words cost).

Preprocessing is key: stop words, lemmatization or stemming, synonyms and a domain dictionary. The resulting document × word matrix is usually reduced by singular value decomposition (SVD) before modelling — otherwise thousands of columns overwhelm even robust methods.

In Statistica

The Text Mining module in the Data Mining menu loads documents from files, databases or the web, applies stemming for many languages, builds a frequency matrix with TF-IDF weighting, reduces it via SVD and hands the results to clustering, trees or neural networks in the Workspace environment; word frequencies and concepts are shown in interactive graphs.

Related terms

Knowledgebase guides

FAQ

Does text mining handle Czech and other languages?
Yes — the basis is stemming and a custom stop-word list; for technical texts it pays to add synonyms and abbreviations from your field.
How many documents are needed?
Hundreds suffice for descriptive analysis, thousands of labelled documents for reliable classification.

Try it on your own data

Statistica free for 30 days

Full version, no credit card. Or get a pricing estimate in a minute.

We'll prepare a tailored quote. Free and with no obligation.

Tell us how many users and what analyses you need — we'll get back to you with a concrete license quote, usually within a few business days.

Try for free first
Statistica.pro

Official Statistica partner for the European Union, based in Prague, Czech Republic. Operated by DataBon s.r.o.

On the market since 1995, formerly as StatSoft CR s.r.o.

Contact

DataBon s.r.o.
Korunní 2569/108, 101 00 Prague 10, Czech Republic
+420 602 284 038
Company ID 09743804 · VAT ID CZ09743804

University licence users: please contact the licence administrator at your institution first.

© 2026 DataBon · statistica.pro · Personal data processing (GDPR) ·