Outlier Removal

The distributions of each parameter will be searched for outliers according to a method of choice and rows containing outliers will be removed.

Options

Method
The method to determine the lower and upper bounds of the data (outlier limits). Mean +- SD: This assumes a normal distribution. outliers are defined to be greater than "Mean + Factor*SD" or smaller than "Mean - Factor*SD". Factor is 3 per default. Boxplot: outliers are defined to be greater than "Q85+Factor*IQR" or smaller than "Q25 - Factor*IQR", where Q are the quantiles and IQR is the inter quantile range. Be careful the default 3 goes with the default method (Mean +- SD). A standard value for this method would be 1.5.
Factor
The factor multiplies the value describing the spread of the distribution.
Subsets
Select the columns by which the measurements should be grouped (example: plates, batches, runs...)
Constraints
The first column filter allows to select multiple columns to define groups. Rows that share the same values for all constraint columns belong to one and the same group.
Parameter
The second column filter is to select the paramters where the values have to be ckecked for outliers.
All Parameter
Per default unchecked. In this case an row is removed if it has an outlier value for at least one parameter. If you tick this checkbox, the row is only removed if all the values for all parameter are outliers.

Input Ports

Icon
data for outlier analysis

Output Ports

Icon
table with outlier rows removed.
Icon
No description for this port available.

Popular Predecessors

  • CSV Reader15 %
  • Number To String9 %
  • Color Manager5 %
  • String To Number4 %
  • Table Reader3 %
  • Nominal Value Row Filter3 %
  • Row Splitter2 %
  • Column List Loop Start2 %
  • Joiner2 %
  • Number To String1 %
  • Scatter Plot (JavaScript)1 %
  • Column Combiner1 %
  • Sorter1 %
  • Column Splitter1 %
  • Excel Reader (XLS)1 %
  • Linear Regression Learner1 %
  • RowID1 %
  • String Replacer1 %
  • Excel Reader (XLS)1 %
  • Partitioning< 1 %
  • Transpose< 1 %
  • XLS Reader< 1 %
  • Java Snippet< 1 %
  • Row Sampling< 1 %
  • String Manipulation< 1 %
  • Value Counter< 1 %
  • Rule-based Row Splitter< 1 %
  • Line Plot (JavaScript)< 1 %
  • R Snippet< 1 %
  • Normalize Plates (POC)< 1 %
  • Table Creator< 1 %
  • Auto-Binner< 1 %
  • Domain Calculator< 1 %
  • GroupBy< 1 %
  • Denormalizer< 1 %
  • Pivoting< 1 %
  • Group Loop Start< 1 %
  • Column Appender< 1 %
  • Column Filter< 1 %
  • Missing Value Column Filter< 1 %
  • Row Filter< 1 %
  • Missing Value< 1 %
  • Column Rename< 1 %
  • Numeric Outliers< 1 %
  • SDF Reader< 1 %
  • Math Formula< 1 %
  • RDKit Interactive Table< 1 %
  • Normalize Plates (Z-Score)< 1 %
  • Duplicate Row Filter< 1 %
  • String To Number< 1 %
  • Z-Primes (PC x NC)< 1 %
  • Category To Number< 1 %
  • Column Aggregator< 1 %
  • Rule-based Row Filter< 1 %
  • Statistics< 1 %
  • Data Generator< 1 %
  • Box Plot (local)< 1 %
  • R Source (Table)< 1 %
  • Group Mutual Information< 1 %
  • Math Formula (Multi Column)< 1 %
  • Excel Reader (XLS)< 1 %
  • Date&Time Difference< 1 %
  • IF Switch< 1 %
  • Range Filter< 1 %
  • Outlier Removal< 1 %
  • Loop End< 1 %
  • k-Means< 1 %
  • Concatenate< 1 %
  • Column Auto Type Cast< 1 %
  • Linear Correlation< 1 %
  • Low Variance Filter< 1 %
  • Rank< 1 %
  • Date&Time to String< 1 %
  • Time Difference< 1 %
  • Counting Loop Start< 1 %
  • Number To String (PMML)< 1 %
  • Database Connection Table Reader< 1 %
  • Database Reader< 1 %
  • Chunk Loop Start< 1 %
  • Decision Tree Predictor< 1 %
  • SMOTE< 1 %
  • Numeric Binner< 1 %
  • Cell Replacer< 1 %
  • Column Merger< 1 %
  • Column Rename (Regex)< 1 %
  • Column Resorter< 1 %
  • Equal Size Sampling< 1 %
  • Normalizer< 1 %
  • String To Number (PMML)< 1 %
  • Extract Table Spec< 1 %
  • Conditional Box Plot (local)< 1 %
  • Java Snippet (simple)< 1 %
  • Recursive Loop Start< 1 %
  • PCA< 1 %
  • Merge Variables< 1 %
  • Column Expressions< 1 %
  • Column Filter< 1 %
  • Splitter By Type< 1 %
  • Create Collection Column< 1 %
  • Table Column to Variable< 1 %
  • File Reader< 1 %
  • X-Aggregator< 1 %
  • Column Comparator< 1 %
  • Constant Value Column< 1 %
  • Correlation Filter< 1 %
  • Double To Int< 1 %
  • Reference Row Filter< 1 %
  • Normalizer (Apply)< 1 %
  • Missing Value (Apply)< 1 %
  • Numeric Row Splitter< 1 %
  • Rule Engine< 1 %
  • Data Explorer (JavaScript)< 1 %
  • CASE Switch Data (End)< 1 %
  • Crosstab< 1 %
  • Integer Input< 1 %
  • Parameter Optimization Loop Start< 1 %
  • Python Script (2⇒1)< 1 %
  • Extract Time Window< 1 %
  • Moving Aggregation< 1 %
  • WAV Reader< 1 %
  • Random Numbers Generator< 1 %
  • Join Layout< 1 %
  • Multivariate Z-Primes< 1 %
  • R Snippet< 1 %
  • File Meta Info< 1 %
  • Table Row To Variable Loop Start< 1 %
  • Cluster Assigner< 1 %
  • Regression Predictor< 1 %
  • Bootstrap Sampling< 1 %
  • Cell Splitter< 1 %
  • CAIM Binner< 1 %
  • Reference Column Filter< 1 %
  • Constant Value Column Filter< 1 %
  • Nominal Value Row Splitter< 1 %
  • Normalizer (PMML)< 1 %
  • Round Double< 1 %
  • String Replace (Dictionary)< 1 %
  • Renderer to Image< 1 %
  • Bland-Altman Plot< 1 %
  • Numeric Outliers (Apply)< 1 %
  • Cache< 1 %
  • DB Column Filter< 1 %
  • Number Filter< 1 %
  • Value Filter< 1 %
  • Double Input< 1 %
  • Moving Average< 1 %
  • Time to String< 1 %
  • XGBoost Predictor (Regression)< 1 %
  • MapQuest Geocoder< 1 %
  • Table Row to Variable< 1 %
  • CSV Reader< 1 %
  • File Reader< 1 %
  • PCA< 1 %
  • Counter Generation< 1 %

Popular Successors

Views

This node has no views

Workflows

Links

Developers

You want to see the source code for this node? Click the following button and we’ll use our super-powers to find it for you.