Free reference · 596 entries
The words, before the course.
Short, plain definitions of the terms that turn up in statistics and data science — what they mean, and how they relate to each other. Nothing to sign in for.
70 of 596 entries under S
- SampleA sample is a portion of the elements of a population. A sample is chosen to make inferences about the population by examining or measuring the elements in the sample.
- Sample Size CalculationsSample size calculations typically arise in significance testing, in the following context: how big a sample size do I need to identify a significant difference of a certain size?
- Sample SpaceThe set of all possible outcomes of a particular experiment is called the sample space for the experiment.
- Sample SurveyIn a sample survey, a sample of units drawn from the population of interest is analyzed. A related concept is the census survey.
- SamplingSampling is a process of drawing a sample from a population. Sampling may be performed from both real and hypothetical populations.
- Sampling DistributionWhen a sample is drawn, some summary value (called a statistic) is usually computed. For example, the sample mean and the sample variance are two statistics.
- Sampling FrameSampling frame (synonyms: “sample frame”, “survey frame”) is the actual set of units from which a sample has been drawn: in the case of a simple random sample, all units from the sampling frame have an…
- Scale Invariance (of Measures)Scale invariance is a property of descriptive statistics. If a statistic is scale-invariant, it has the following property for any sample and any non-negative value : (1) or, in mathematically equivalent…
- Scatter GraphsA scatter graph shows the joint distribution of observed values of two variables. Each pair of values is shown as a point on X-Y plane with coordinates (Xi,Yi), where Xi and Yi are the values of the first…
- Seasonal AdjustmentThe seasonal adjustment is used in time series analysis to remove a periodic component with the known period from the observed time series.
- Seasonal DecompositionThe seasonal decomposition is a method used in time series analysis to represent a time series as a sum (or, sometimes, a product) of three components – the linear trend, the periodic (seasonal) component,…
- Seemingly Unrelated Regressions (SUR)Seemingly unrelated regressions (SUR) is a class of multivariate regression ( multiple regression) models, normally belonging to the sub-class of linear regression models.
- Self-Controlled DesignIn randomized trials, a self-controlled design is one in which results are measured in each subject before and after treatment.
- SensitivitySensitivity (of a medical diagnostic test for a disease) is the probability that the test is positive for a person with the disease.
- Sequential Icon PlotsSequential icon plots (or column icon plots) are a category of icon plots. For each unit, variables are represented as a sequence of bars with the height reflecting the value of the corresponding variable.
- Serial CorrelationIn analysis of time series, the Nth order serial correlation is the correlation between the current value and the Nth previous value of the same time series.
- Shift Invariance (of Measures)Shift invariance is a property of descriptive statistics. If a statistic is shift-invariant, it possesses the following property for any data set : or, in equivalent form In other words, if a statistic is…
- Sign TestThe sign test is a nonparametric test used with paired replicates to test for the difference between the 1st and the 2nd measurement in a group of “subjects”.
- SignalThe signal is the component of the observed data (e.g. of a time series) that carries useful information.
- Signal ProcessingSignal processing is a branch of applied statistics concerned with analysis of functions of time that take on scalar or vector values.
- Significance TestingSee Hypothesis Testing
- Similarity MatrixSimilarity matrix is the opposite concept to the distance matrix. The elements of a similarity matrix measure pairwise similarities of objects – the greater similarity of two objects, the greater the value…
- Simple Linear RegressionThe simple linear regression is aimed at finding the “best-fit” values of two parameters – A and B in the following regression equation: Yi = A Xi + B + Ei, i=1,¼,N where Yi, Xi, and Ei are the values of…
- Simple Linear Regression (Graphical)The simple linear regression is aimed at finding the “best-fit” values of two parameters – A and B in the following regression equation: where Yi, Xi, and Ei are the values of the dependent variable, of the…
- SimulationIn general, simulation is modelling of a process or phenomenon. In statistics, Monte Carlo simulation is often used to model outcomes of a random experiment.
- Single Linkage ClusteringThe single linkage clustering method (or the nearest neighbor method) is a method of calculating distance between clusters in hierarchical cluster analysis.
- SingularityIn regression analysis, singularity is the extreme form of multicollinearity – when a perfect linear relationship exists between variables or, in other terms, when the correlation coefficient is equal to…
- Six-SigmaSix sigma means literally six standard deviations. The phrase refers to the limits drawn on statistical process control charts used to plot statistics from samples taken regularly from a production process.
- SkewnessSkewness measures the lack of symmetry of a probability distribution. A curve is said to be skewed to the right (or positively skewed) if it tails off toward the high end of the scale (right tail longer…
- Smoother (Example)A simple example of a smoother is the moving average procedure. It is based on averaging elements closest in time to the current time.
- Smoother (Smoothing Filter)Smoothers, or smoothing filters, are algorithms for time-series processing that reduce abrupt changes in the time-series and make it look smoother.
- SmoothingSmoothing is a class of time series processing which is intended to reduce noise and to preserve the signal itself.
- Social Network AnalyticsNetwork analytics applied to connections among humans. Recently it has come also to encompass the analysis of web sites and internet services like Facebook.
- SparkSpark is a second generation computing environment that sits on top of a Hadoop system, supporting the workflows that leverage a distributed file system.
- Spatial FieldA spatial field is a function of spatial variables , or in 3D cases. A spatial field is named a “scalar field” if the function takes on scalar values.
- SpecificitySpecificity (of a medical diagnostic test for a disease) is the probability that the test will come out negative for a person without the disease.
- Spectral AnalysisSpectral analysis is concerned with estimation of the spectrum of a stationary random process or a stationary time series from the observed realization(s) of the process (or series).
- SpectrumSee Fourier spectrum and power spectrum.
- SplineA spline is a continuous function which coincides with a polynomial on every subinterval of the whole interval on which is defined.
- Split-Halves MethodIn psychometric surveys, the split-halves method is used to measure the internal consistency reliability of survey instruments, e.g.
- SQLSQL stands for structured query language, a high level language for querying relational databases, extracting information.
- Standard DeviationThe standard deviation is a measure of dispersion. It is the positive square root of the variance.
- Standard errorThe standard error measures the variability of an estimator (or sample statistic) from sample to sample.
- Standard Normal DistributionThe standard normal distribution is the normal distribution where the mean is zero and the standard deviation is one.
- Standard ScoreThe standard score of an observation is the number of standard deviation units it is above or below the mean.
- Standardized Mean DifferenceThe standardized mean difference is the difference between two normalized means – i.e. the mean values divided by an estimate of the within-group standard deviation.
- StanineA stanine is a “standard ninth,” an interval used in dividing school test results into (more or less) ninths.
- Star Icon PlotsStar icon plots are a subclass of circular icon plots in which the rays tend to form a star.
- State SpaceState space is an abstract space representing possible states of a system. A point in the state space is a vector of the values of all relevant parameters of the system.
- Stationary time seriesA time series x(t); t=1,… is called to be stationary if its statistical properties do not depend on time t .
- Statistic1. A number measuring something 2. A measure calculated from a sample of data. Contrast “statistic” (drawn from a sample) with “parameter,” which is a characteristic of a population.
- Statistical SignificanceOutcomes to an experiment or repeated events are statistically significant if they differ from what chance variation might produce.
- Statistical TestA statistical test is a procedure for statistical hypothesis testing. The outcome of a statistical test is a decision to reject or accept the null hypothesis for given probability of type I error.
- Statistics1. A collection of numerical data that measure something. 2. The science of recording, organizing, analyzing and reporting quantitative information.
- StemmingIn processing unstructured text, stemming is the process of converting multiple forms of the same word into one stem, to simplify the task of analyzing the processed text.
- Step-wise RegressionStep-wise regression is one of several computer-based iterative variable-selection procedures.
- Stochastic ProcessStochastic process is a synonym for random process.
- Stratified SamplingStratified sampling is a method of random sampling. In stratified sampling, the population is first divided into homogeneous groups, also called strata.
- Strip transectA strip transect is a small subsection of a geographically-defined study area, typically chosen randomly.
- Structural Equation ModelingStructural equation modeling includes a broad range of multivariate analysis methods aimed at finding interrelations among the variables in linear models by examining variances and covariances of the variables.
- Structured vs. unstructured dataStructured data is data that is in a form that can be used to develop statistical or machine learning models (typically a matrix where rows are records and columns are variables or features).
- Sufficient StatisticSuppose X is a random vector with probability distribution (or density) P(X | V), where V is a vector of parameters, and Xo is a realization of X.
- Sufficient Statistic (Graphical)Suppose X is a random vector with probability distribution (or density) P(X | V), where V is a vector of parameters, and Xo is a realization of X.
- Sun Ray PlotsSun ray plots are a subclass of circular icon plots in which the rays tend to form a circle.
- Support Vector MachinesSupport vector machines are used in data mining (predictive modeling, to be specific) for classification of records, by learning from training data.
- SurveyStatistical surveys are general methods to gather quantitative information about a particular population.
- Survival AnalysisSurvival analysis is concerned with “time-to-event” data. In medical statistics, the data are often in the form of “time-to-death”.
- Survival FunctionIn medical statistics, the survival function is a relationship between a proportion and time.
- Systematic ErrorSystematic error is the error that is constant in a series of repetitions of the same experiment or observation.
- Systematic SamplingSystematic sampling is a method of random sampling. The elements to be sampled are selected at a uniform interval that is measured in time, order, or space.
Ready to do it rather than read it?
Every course runs on a fixed start date with an instructor who marks your work, and selected ones carry a credit recommendation from the American Council on Education.