Free reference · 596 entries
The words, before the course.
Short, plain definitions of the terms that turn up in statistics and data science — what they mean, and how they relate to each other. Nothing to sign in for.
45 of 596 entries under P
- p-valueThe p-value is the probability that the null model could, by random chance variation, produce a sample as extreme as the observed sample (as measured by some sample statistic of interest.)
- Paired Replicates DataPaired replicates is the simplest form of repeated measures data, when only two measurements are made for each experimental unit.
- Panel DataA panel data set contains observations on a number of units (e.g. subjects, objects) belonging to different clusters (panels) over time.
- Panel studyA panel study is a longitudinal study that selects a group of subjects then records data for each member of the group at various points in time.
- Parallel DesignIn randomized trials, a parallel design is one in which subjects are randomly assigned to treatments, which then proceed in parallel with each group.
- ParameterA Parameter is a numerical value that describes one of the characteristics of a probability distribution or population.
- Parametric TestsIn statistical inference procedures (hypothesis tests and confidence intervals), parametric procedures are those that incorporate assumptions about population parameters.
- Path AnalysisPath analysis is a method for causal modeling. Consider the simple case of two independent variables x1 and x2 and one dependent variable.
- Path coefficientsIn path analysis and structural equation modeling a path coefficient is the partial correlation coefficient between the dependent variable and an independent variable, adjusted for other independent variables.
- Pearson correlation coefficientSee correlation coefficient.
- PercentileIn a population or a sample, the Pth percentile is a value such that at least P percent of the values take on this value or less and at least (100-P) percent of the values take on this value or more.
- Permutation TestsA permutation test involves the shuffling of observed data to determine how unusual an observed outcome is.
- Pie Icon PlotsPie icon plots are a sub-class of icon plots. Each unit or observation is represented by a circle with colored “pies slices” corresponding to variables – the angular size of a slice of pie is proportional…
- Pivotal StatisticA statistic is said to be pivotal if its sampling distribution does not depend on unknown parameters.
- Poisson Distributionk! p(x=k) = lk e–l, k=0,1,2,� (where k! = 1 x 2 x … x k). Both the mean and the variance of Poisson distribution are equal to l.
- Poisson Distribution (Graphical)Poisson distribution is a discrete distribution, completely characterized by one parameter : (where k!
- Poisson ProcessA Poisson process is a random function U(t) which describes the number of random events in an interval [0,t] of time or space.
- Poisson Process (Graphical)A Poisson process is a random function U(t) which describes the number of random events in an interval [0,t] of time or space.
- Polygon Icon PlotsPolygon icon plots are a subclass of circular icon plots in which the rays tend to form a polygon.
- PolynomialA polynomial of order is a function described by the following expression: where are coefficients of the polynomial.
- PopulationA population is a large set of objects of a similar nature – e.g. human beings, households, readings from a measurement device – which is of interest as a whole.
- Post-hoc testsPost-hoc tests (or post-hoc comparison tests) are used at the second stage of the analysis of variance (ANOVA) or multiple analysis of variance (MANOVA) if the null hypothesis is rejected.
- Posterior ProbabilityPosterior probability is a revised probability that takes into account new available information.
- Power MeanA power mean of order of a set of values is defined by the following expression: The family of power mean statistics is often called the generalized mean – because, for different values of the parameter ,…
- Power of a Hypothesis TestThe power of hypothesis test is a measure of how effective the test is at identifying (say) a difference in populations if such a difference exists.
- Power SpectrumThe power spectrum of a stationary random process or a stationary time series is the average of the square of the amplitude of the Fourier spectrum: where is the amplitude spectrum of the realization of the…
- PrecisionPrecision is the degree of accuracy with which a parameter is estimated by an estimator. Precision is usually measured by the standard deviation of the estimator and is known as the standard error.
- Predicting FilterPredicting filters are filters that estimate the next value in a time series from the known previous values.
- Prediction vs. ExplanationWith the advent of Big Data and data mining, statistical methods like regression and CART have been repurposed to use as tools in predictive modeling.
- Predictive ModelingPredictive modeling is the process of using a statistical or machine learning model to predict the value of a target variable (e.g.
- Predictive ValidityThe predictive validity of survey instruments and psychometric tests is a measure of agreement between results obtained by the evaluated instrument and results obtained from more direct and objective…
- predictorsee dependent and independent variables
- Predictor VariablePredictor variable is a synonym for independent variable.
- Principal Component AnalysisThe purpose of principal component analysis is to derive a small number of linear combinations (principal components) of a set of variables that retain as much of the information in the original variables…
- Prior and posteriorBayesian statistics typically incorporates new information (e.g. from a diagnostic test, or a recently drawn sample) to answer a question of the form “What is the probability that…” The answer to this…
- Prior and posterior probability (difference)Consider a population where the proportion of HIV-infected individuals is 0.01. Then, the prior probability that a randomly chosen subject is HIV-infected is Pprior = 0.01 .
- Prior ProbabilitySee A Priori Probability.
- ProbitProbit is a nonlinear function of probability p: probit(p) = F–1(p) where F–1() is the function inverse to the cumulative distribution function F() of the standard normal distribution.
- Proportional Hazard ModelProportional hazard model is a generic term for models (particularly survival models in medicine) that have the form L(t | x1, x2, ¼, xn) = h(t) exp(b1 x1 + ¼+ bn xn), where L is the hazard function…
- Proportional Hazard Model (Graphical)Proportional hazard model is a generic term for models (particularly survival models in medicine) that have the form where L is the hazard function or hazard rate, {xi} are covariates, {bi} are coefficients…
- Prospective Versus RetrospectiveProspective vs. Retrospective A prospective study is one that identifies a scientific (usually medical) problem to be studied, specifies a study design protocol (e.g.
- Pruning the tree<b Pruning the tree: Classification and regression trees, applied to data with known values for an outcome variable, derive models with rules like “If taxable income <$80,000, if no Schedule C income, if…
- Pseudo-Random NumbersPseudo-random numbers are produced by recursive algorithms – i.e. the current number is calculated from one or a greater number of previous numbers.
- Psychological TestingSee psychometrics.
- PsychometricsPsychometrics or psychological testing is concerned with quantification (measurement) of human characteristics, behavior, performance, health, etc., as well as with design and analysis of studies based on…
Ready to do it rather than read it?
Every course runs on a fixed start date with an instructor who marks your work, and selected ones carry a credit recommendation from the American Council on Education.