To request a blog written on a specific topic, please email James@StatisticsSolutions.com with your suggestion. Thank you!

Friday, February 20, 2009

Sample Size Calculation


Sample size calculations are required for a majority of quantitative studies involving surveys and statistics. It is a necessity to consider sample size calculations in order to ensure that analyses have adequate statistical power and that the results obtained are accurate and useful. If samples are too large in size, researchers could waste time, money and resources. On the other hand, if samples are too small, the results obtained may not be accurate or reliable. A sample size calculation is not necessarily complicated or unnerving, though it does tend to strike several statisticians as either a minor technicality or a huge undertaking. There is no mystery involved in estimating the sample size. It is a relatively straightforward task for the equipped and experienced researcher. However, considering it’s tremendous importance in the overall project setting, it may be best left to an expert statistician.

One of the benefits of performing a sample size calculation is that it helps in setting a project on the right foot. A proper sample size ensures that analyses conducted will produce reliable and usable results. Before calculating the sample size, it is necessary to develop a thorough knowledge of the requirements of the project and the nature of the statistical analysis to be conducted. This feeds into the method by which the sample size will be calculated. Nowadays, there is a plethora of online sample size calculators that can simply calculate your sample size. These may seem useful but are more like a band-aid. It is recommended that unless one is an expert, one does not attempt such calculation short cuts without a thorough understanding of the underlying methodologies.

There isn’t a single standard sample size equation. The best equation is one that addresses the needs of the project analysis, types of variables and intended outcomes. For instance, two different sample size equations are available for use in continuous variable and categorical data.

The first and foremost thing that must be kept in mind is that a sample size is essentially a function of effect, significance level and power. In other words, it signifies that effect, significance and power are the three levels on which the sample size is going to depend. If any one of the three measures is changed, then sample size will also change as a result.

Sample size calculation largely relies on the statistical tests that are intended to be conducted. This is because there will be differences in the effect depending on the statistical test(s) being conducted.

In addition to the statistical method in question, there are several other factors on which the sample size calculation depends. These factors are important to consider for any kind of sample size calculation.

· Type of data

· The requisite level of significance

· The desired power

· The standard deviation of continuous outcome variables

· The effect size

· The one and two sided tests of significance

· Other various aspects of design of the study

In case of sample size calculation where continuous data are involved, categorical formulas for sample size calculation must be used. The following formula will be applied to such requirements.

no = ( t)2 * (s)2

_____________

(d)2

Where t = value of selected alpha level

no = required return sample size

s= estimate of standard deviation in the population

d = acceptable margin for error in mean

Where the data or variable(s) are categorical, sample size calculation will differ in terms of approach. The following sample size calculation formula will be applied in that case.

no = ( t)2 * (p) (q)

__________________

(d)2

Where t = value of selected alpha level

(p) (q) = value of selected alpha level

s = estimate of standard deviation in the population

d = acceptable margin for error in mean


Click here for more assistance with Sample Size.

Friday, January 2, 2009

Sample Size Calculation for One-Way ANOVAs in Dissertations and Theses

I'm sure there are some of you out there looking for the minimum sample size necessary to find the analyses of variance (ANOVAs) in your dissertation or thesis significant. Sample size calculation for ANOVAs can be complicated if it's a factorial ANOVA or mixed ANOVA, so we'll start slow and focus on an ANOVA with only one independent variable.

What is the Power used in calculating the sample size of the ANOVA being used in my dissertation or thesis?

For the purposes of this example, we are going to looking for the minimum sample size to give us a power of 0.80. To read more about this, click here. This is going to give us a 20% probability of falsely accepting the null hypothesis, or a 20% probability that we missed something. We're okay with this, since missing something is typically less severe than finding something that isn't really there. Click here for help with determining the appropriate power for your dissertation or thesis.

What is the level of significance used in calculating the sample size of the ANOVA being used in my dissertation or thesis?

This is the probability of falsely rejecting the null hypothesis. The statistical significance for the purposes of calculating the sample size for the ANOVA is going to be 0.05. This means we are looking for less than a 5% probability that our results are due to chance. Get help with determining the ANOVA level of significance for the sample size calculation in your dissertation or thesis.

What is the effect size used in calculating the sample size of the ANOVA being used in my dissertation or thesis?

There are a couple things involved in determining this. Since choosing a small effect size will require that we gather thousands of observations to find our ANOVA significant, and choosing a large effect size will mean fewer people but not a very good chance of finding the test significant if the groups are not hugely different.

What we need here is something in the middle…the medium effect size. For the purposes of the dissertation or thesis, this is definitely acceptable. Get help with determining the ANOVA effect size for the sample size calculation in your dissertation or thesis.

What is the sample size needed for the ANOVAs in my dissertation or thesis?

Using the criteria above, the sample size needed for the one-way ANOVA, testing for differences on one independent variable with two groups, is 128, the same as the independent samples t-test. The sample size will vary with the number of groups in the independent variable, but for the independent variable with 3 groups, you will need 156 or approximately 52/group. Get help with a custom sample size calculation for your dissertation or thesis.

Tuesday, December 30, 2008

Sample Size for Bivariate Correlation, Pearson Correlation, and Pearson Product Moment Correlation

To satisfy some of the requests of my blog readers, I am covering sample size calculation for a bivariate correlation or the Pearson correlation. This test might also be called the Pearson product-moment correlation.

I am going to assume that you know what a Pearson correlation is and its function, if not check out this blog entry on dissertation statistics help featuring bivariate correlation. In a nutshell we are testing for a significant relationship between two variables. Please keep reading, but if you are just looking for someone to help you calculate the sample size for your Master's thesis, Master's dissertation, Ph.D. thesis, or Ph.D. dissertation using bivariate correlation, Pearson correlation, or Pearson product-moment correlation, or to justify the sample you already have, click here.

Sample Size for Bivariate Correlation or Pearson Correlation

There are some things we have to understand prior to calculating the sample size of our bivariate correlation or Pearson correlation. We have to first understand why we are calculating the sample size. If you are looking for some more information on these things, check out this blog entry.

Significance

Sample size is calculated for the bivariate correlation or the Pearson correlation so we know how many people we have to survey, poll, or sample to find the test significant at the level of significance we have set. This is the probability of committing a Type I error. Usually the level of significance is set at 0.05. This means there is a 5% probability that our results are due to chance. Get help with determining the correct level of significance for your bivariate correlation, Pearson correlation, or Pearson product-moment correlation.

Power

Power is the opposite of significance and is probability of falsely accepting the null hypothesis or… in plain English… the probability that we missed something and the test we ran was significant even though the result was not significant. This is the probability of committing a Type II error. Usually this is set at 0.80, making the probability 20% or four times as likely as committing a Type I error (measured by our level of significance). Get help with determining the correct power for your bivariate correlation, Pearson correlation, or Pearson product-moment correlation.

Effect Size

This circumstance is slightly different than other tests, in that there is no causality or direction in a sense. Effect size in this case is measured as r and represents the strength of the relationship. These r effect sizes for the bivariate correlation and the Pearson correlation are 0.10 for a small effect size, 0.30 for a medium effect size, and 0.50 for a large effect size. Just to make sure credit is given where credit is due, these effect sizes are courtesy of Jacob Cohen and his fantastically helpful article A Power Primer. For this example we will use a medium effect size. Get help with determining the correct effect size for your bivariate correlation, Pearson correlation, or Pearson product-moment correlation.

Now that we have determined these factors – and these numbers are the numbers that will be used 98% of the time in a Master's thesis, Master's dissertation, Ph.D. thesis, and Ph.D. dissertation – the rest of the sample size calculation for the bivariate correlation or the Pearson correlation is easy. For this we will refer again to A Power Primer by Jacob Cohen. If you are looking for this journal article you will find it here.

What is the sample size needed for a significant bivariate correlation or a significant Pearson correlation (Pearson product-moment correlation)?

Here it is…. 85. For a significant Pearson product-moment correlation at a 0.05 level of significance, a power of 0.80, and a medium effect size, we need 85 people. This number will fluctuate with changes in any of those measures, including power, which is sometimes set at 0.90. To have me calculate the sample size needed for your bivariate correlation, Pearson correlation, Pearson product-moment correlation, or for that matter any correlation or test, click here.

Tuesday, December 23, 2008

Sample Size Calculation for Dependent Samples t-test

A Priori Sample Size for Dependent Samples t-test

Sample Size Calculation for Dependent Samples t-tests are not as simple as sample size calculation for the independent samples t-test. While the sample size requirement is smaller because the two samples are related or correlated, the calculation is somewhat complicated. In order to calculate the minimum sample size for the dependent samples t-test being used in your Master's thesis, Ph.D. thesis, Master's dissertation, or Ph.D. dissertation, you are going to need some information or have a good idea of values for key pieces of information.

What is my thesis power analysis or dissertation power analysis for a dependent samples t-test or paired samples t-test?

For the purposes of this example, I will refer to a Jacob Cohen book, Statistical Power Analysis for the Behavioral Sciences. We are going to define power analysis for the dependent samples t-test as the sample size necessary to…

  1. Achieve a Power of 0.80
  2. Detect a reasonable difference between the groups with a medium effect size of 0.50
  3. Detect a significant difference between the groups at a 0.05 level of significance

Important to understand is that each of the four items mentioned (Sample Size, Power, Effect Size, and Level of Significance) are a function on one another, meaning that changing any one of these things is going to change the value of the other three. So we have determined values for Power (0.80), Effect Size (0.50), and Level of Significance (0.05), and we are trying to figure out how many people we need to actually achieve all of these values. Get help with your thesis power analysis or dissertation power analysis

How is sample size calculation for a dependent samples t-test or paired samples t-test different than that of an independent samples t-test?

Before, we found the sample size for an independent samples t-test by looking at a simple table. This time, the groups are related and we need to take that into account in our sample size calculation. Before we continue, however, I would like to tell you the good news… The sample size requirements for the independent samples t-test are much larger than that of the dependent samples t-test. If you have calculated the independent samples t-test sample size and power analysis for your dissertation or thesis, can obtain that many pairs of participants, and are not under a lot of pressure from your college or university to justify the sample size, then...

Stop Here.

If you are like the rest of the world, please continue. For your dissertation or thesis using dependent samples t-test, your sample size is not going to be measure in participants, but in pairs of participants. Get help with your thesis power analysis or dissertation power analysis

How do I calculate the sample size for dependent samples t-tests or paired samples t-test?

Since we have established that we are conducting a dependent samples t-test, and are assuming the scores for the pairs are related or correlated, we need to know the degree to which they are correlated. Since this is an a priori power analysis for a dissertation or thesis, we could not possibly know the exact correlation between pairs in the sample we are going to obtain and must estimate the degree to which these pairs are going to be correlated.

If this is a dissertation or thesis following other research designs and there is empirical research information available, then look for a correlation coefficient from some of these other studies. For instance, if four of the studies you have read found a strong correlation (0.90) between childhood obesity prior to dieting and after dieting, then we can assume that we are going to find the same thing in our study. We would assume that the correlation would be 0.90. Get help with your thesis power analysis or dissertation power analysis

What is the equation for calculating sample size for dependent samples t-test?

Here it is in all of its mathematical glory…

n= n_(.10)/(100d^(2 ) )+ 1

where, n.10 = 1571 and is the necessary sample size for the given a or α and Power, d = 0.645 and is the ES index for t tests of means in standard unit calculated by the equation:

d= d_(4^' )/√(1-r)

where
= the medium effect size of 0.50 and
r
= 0.40
is the estimation of the correlation within the pairs.
This yields d = 0.645 = 0.05/√ (1 – 0.40).

OR

Just tell them you need 39 pairs of participants or scores. Better yet, get professional help with your thesis or dissertation power analysis. Let us give you a customized power analysis for your Master's thesis, Ph.D. thesis, Master's dissertation, or Ph.D. dissertation.