Free reference · 596 entries
The words, before the course.
Short, plain definitions of the terms that turn up in statistics and data science — what they mean, and how they relate to each other. Nothing to sign in for.
30 of 596 entries under D
- DataData are recorded observations made on people, objects, or other things that can be counted, measured, or quantified in some way.
- Data MiningData mining is concerned with finding latent patterns in large data bases. The goal is to discover unsuspected relationships that are of practical importance, e.g., in business.
- Data PartitionData partitioning in data mining is the division of the whole data available into two or three non-overlapping sets: the training set, the validation set, and the test set.
- Data ProductA data product is a product or service whose value is derived from using algorithmic methods on data, and which in turn produces data to be used in the same product, or tangential data products.
- DecileDeciles are percentiles taken in tens. The first decile is the 10th percentile, the second decile is the 20th percentile, etc.
- Decile LiftIn predictive modeling, the goal is to make predictions about outcomes on a case-by-case basis: an insurance claim will be fraudulent or not, a tax return will be correct or in error, a subscriber will…
- Decision TreesIn the machine learning community, a decision tree is a branching set of rules used to classify a record, or predict a continuous value for a record.
- Deep LearningDeep Learning refers to complex multi-layer neural nets. They are especially suitable for image and voice recognition, and for unsupervised tasks with complex, unstructured data.
- Degrees of FreedomFor a set of data points in a given situation (e.g. with mean or other parameter specified, or not), degrees of freedom is the minimal number of values which should be specified to determine all the data…
- DendrogramThe dendrogram is a graphical representation of the results of hierarchical cluster analysis.
- Density (of Probability)A probability density function or curve is a non-negative function ( ) that describes the distribution of a continuous random variable.
- Dependent and Independent VariablesStatistical models normally specify how one set of variables, called dependent variables, functionally depend on another set of variables, called independent variables.
- Dependent EventsSee Independent Events.
- Descriptive StatisticsDescriptive statistics refers to statistical techniques used to summarize and describe a data set, and also to the statistics (measures) used in such summaries.
- Design of ExperimentsDesign of experiments is concerned with optimization of the plan of experimental studies. The goal is to improve the quality of the decision that is made from the outcome of the study on the basis of…
- Detrended Correspondence AnalysisDetrended correspondence analysis is an extension of correspondence analysis (CA) aimed at addressing a deficiency of correspondence analysis.
- DichotomousDichotomous (outcome or variable) means “having only two possible values”, e.g. “yes/no”, “male/female”, “head/tail”, “age > 35 / age <= 35” etc.
- Differencing (of Time Series)Differencing of a time series in discrete time is the transformation of the series to a new time series where the values are the differences between consecutive values of .
- Directed vs. Undirected NetworkIn a directed network, connections between nodes are directional. For example, in a Twitter network, Smith might follow Jones but that does not mean that Jones follows Smith.
- Discrete DistributionA discrete distribution describes the probabilistic properties of a random variable that takes on a set of values that are discrete, i.e.
- Discrete Random VariableA random variable whose range of possible values is finite or countably infinite is said to be a discrete random variable.
- Discriminant AnalysisDiscriminant analysis is a method of distinguishing between classes of objects. The objects are typically represented as rows in a matrix.
- Dispersion (Measures of)Measures of dispersion express quantitatively the degree of variation or dispersion of values in a population or in a sample.
- Disproportionate Stratified Random SamplingSee Stratified Sampling (method ii).
- Dissimilarity MatrixThe dissimilarity matrix (also called distance matrix) describes pairwise distinction between M objects.
- DistanceDendrogram: Statistical distance is a measure calculated between two records that are typically part of a larger dataset, where rows are records and columns are variables.
- Distance MatrixDistance matrix is often used as a synonym for dissimilarity matrix. The “distance” does not necessarily means distance in space.
- Divergent ValidityIn psychometrics, the divergent validity of a survey instrument, like an IQ-test, indicates that the results obtained by this instrument do not correlate too strongly with measurements of a similar but…
- Divisive Methods (of Cluster Analysis)In divisive methods of hierarchical cluster analysis, the clusters obtained at the previous step are subdivided into smaller clusters.
- Dunn TestThe Dunn test is a method for multiple comparisons, which generalizes the Bonferroni adjustment procedure.
Ready to do it rather than read it?
Every course runs on a fixed start date with an instructor who marks your work, and selected ones carry a credit recommendation from the American Council on Education.