Icon

Sentiment_​Text_​Mining_​Audit_​Ready

1. Import, label and combine the three sources

Three tab-separated files, Source added before concatenation, binary label mapped to Sentiment.

2. Counts, balance and duplicate sensitivity

Primary corpus keeps all 3,000 rows. Duplicate diagnostics and 2,983-row sensitivity branch are separate.

3. Text normalization and cleaning

Control characters and contractions first, then punctuation, numbers, lowercase, minimum length 2, and a custom stop-word list that preserves negators.

4. Unigrams, TF-IDF, source/sentiment comparisons and bigrams

Inspect the ranked GroupBy/Sorter outputs and both visualization nodes after execution.

Replace C1 controls; expand can't, won't and n't before tokenization
String Manipulation
Constant Value Column Appender
Text = NormalizedText | Source metadata = Source | Category = Sentiment
Strings to Document
Cleaning 1: remove punctuation
Punctuation Erasure
Cleaning 4: remove one-character terms; minimum length = 2
N Chars Filter
CSV Reader (deprecated)
Cleaning 2: remove terms representing numbers
Number Filter
Cleaning 3: lowercase
Case Converter
Absolute term frequency by document
TF
Convert Term to string for grouping and display
Term to String
Cleaning 5: custom list only; preserve not, no, never, nor, neither, without, hardly, barely
Stop Word Filter
Create unigram bag of words from cleaned Document
Bag Of Words Creator
Overall cleaned unigram frequency
GroupBy
Sorter
Recover Source and Category metadata after BoW
Document Data Extractor
Sorter
Unigrams by Source
GroupBy
Visualization 2: overall cleaned unigram tag cloud
Tag Cloud (JavaScript) (legacy)
Unigrams by Sentiment/Category
GroupBy
Sorter
Sorter
Unigrams by Source x Sentiment
GroupBy
Stack IMDb and Amazon; retain common columns
Concatenate
TF-IDF = TF rel * IDF
Math Formula
Convert TF-IDF branch Term to string
Term to String
Relative TF = term count divided by document token count
TF
Smooth IDF = log(1 + N/df)
IDF
Sorter
Create word bigrams; preserves negation context
NGram Creator
Recover Source and Category metadata
Document Data Extractor
Aggregate TF-IDF by Category and Term
GroupBy
Absolute bigram frequency by document
TF
Convert bigram Term to string
Term to String
Recover Source and Category metadata
Document Data Extractor
Bigrams by Sentiment/Category
GroupBy
Q2: count Positive and Negative reviews
GroupBy
Sorter
Overall bigram frequency
GroupBy
Sorter
Pivot sentiment into Positive and Negative count columns
Pivot
Line Reader
Map label 1 to Positive; label 0 to Negative
Rule Engine
Visualization 1: stacked counts by source and sentiment
Bar Chart (JavaScript) (legacy)
Constant Value Column Appender
Q3: count reviews per source
GroupBy
Line Reader
Source x Sentiment counts | expected six cells of 500
GroupBy
Line Reader
Optional sensitivity branch | keep first exact observation
Duplicate Row Filter
Sensitivity counts by Source x Sentiment | expected 2,983 rows total
GroupBy
Flag exact duplicate observations; retain all rows in primary analysis
Duplicate Row Filter
Constant Value Column Appender
Duplicate audit counts | expected 17 extra observations
GroupBy

Nodes

Extensions

Links