To request a blog written on a specific topic, please email James@StatisticsSolutions.com with your suggestion. Thank you!

Monday, June 29, 2009

Sampling

The general idea behind sampling is the extrapolation from the sample to the population. Sampling must be done in such a manner that the sample that is being drawn from the population should represent the population as a whole. The method of choosing the type of sampling is called design.

Statistics Solutions is the country's leader in statistical consulting and can assist with sampling for your dissertation, thesis or research project. Contact Statistics Solutions today for a free 30-minute consultation.

An appropriate type of sampling involves probability. Sampling that is done with the help of probability methods is called probability sampling. Biased results or estimates are serious problems in sampling, and the researcher can get rid of these with the help of the probability involved in sampling.

In order to conduct sampling by means of probability, it is important to identify the population of interest. The next step is then to create the sampling frame.

There is another kind of sampling that is more flexible and easy to understand to the person not familiar with statistics. This sampling is nothing but simple random sampling. For instance, in order to conduct simple random sampling of 100 units of an item, the researcher chooses one unit at random from the sampling frame, and then the second unit, (and so on) until the 100th unit has been chosen by means of simple random sampling. In each step of this type of sampling, every unit has a similar chance of getting selected.

This type of sampling is generally practically feasible in cases where the population consists of business records. The consequence of this type of sampling would not get affected even when the population is of a larger size.

There are two kinds of errors in sampling, namely random error and systematic error.

Sampling error generally occurs in cases where the researcher gets very few units of a desirable sample from the population. The obvious consequence of this type of sampling error is generally quantified by utilizing the standard error or simply ‘SE.’

In the case of sampling involving probability, the SE can be estimated by using the sample design and the sample data. As the size of the sample in sampling increases, then the SE gets decreased. So, if the population on which the sampling is being carried out is relatively homogeneous, then the SE will be small.

In cases of the sampling involving cluster, there is generally a larger SE. However it should be noted that sampling that involves clusters are generally cost effective.

The non sampling error is generally more serious as the non sampling errors are usually harder to quantify and therefore draw less attention in comparison to sampling errors. This problem of the non sampling error cannot be controlled by increasing the size of the sample. The non sampling error can be categorized into three categories: selection bias, non response bias and response bias.

The first category of non sampling error is selection bias and it is a systematic tendency to exclude one kind of unit from the sample. In cases of sampling that involve probability, this type of bias is generally minimal.

The second category of bias for non sampling errors usually occurs in those cases when the respondents do not respond to sensitive questions. In order to minimize this type of bias of the non sampling error, the response rate should be kept high.

The third category of the bias of the non sampling error occurs in cases when the respondent does not answer the question honestly.

Friday, June 26, 2009

Resampling

Resampling is the method that consists of drawing repeated samples from the original data samples. The method of Resampling is a nonparametric method of statistical inference. In other words, the method of Resampling does not involve the utilization of the generic distribution tables (for example, normal distribution tables) in order to compute approximate p probability values. Resampling involves the selection of randomized cases with replacement from the original data sample in such a manner that each number of the sample drawn has a number of cases that are similar to the original data sample. Due to replacement, the drawn number of samples that are used by the method of Resampling consists of repetitive cases.

Statistics Solutions is the country's leader in statistical consulting and can assist in resampling techniques. Contact Statistics Solutions today for a free 30-minute consultation.

Resampling is also known as Bootstrapping or Monte Carlo Estimation. Resampling generates a unique sampling distribution on the basis of the actual data. The method of Resampling uses experimental methods, rather than analytical methods, to generate the unique sampling distribution. The method of Resampling yields unbiased estimates as the method of Resampling is based on the unbiased samples of all the possible results of the data studied by the researcher.
In order to understand the concept of Resampling, the researcher should understand the terms Bootstrapping and Monte Caro estimation.

The method of bootstrapping, which is equivalent to the method of Resampling, utilizes repeated samples from the original data sample in order to calculate the test statistic.

Monte Carlo estimation, which is also equivalent to the bootstrapping method, is used by the researcher to obtain the Resampling results.

There are certain assumptions that are made by the researcher while conducting the method of Resampling.

This method of Resampling is generally based on nonparametric assumptions.

This method of Resampling generally ignores the parametric assumptions that are about ignoring the nature of the underlying data distribution. Therefore, Resampling is based on nonparametric assumptions.

Sample size assumption of the Resampling: In Resampling, there is no specific sample size requirement. Therefore, the larger the sample, the more reliable the confidence intervals generated by the method of Resampling.

In the method of Resampling, there is an increased danger of over fitting noise in the data. This type of problem can be solved easily by combining the method of Resampling with the process of cross-validation.

In SPSS, the researcher can perform the method of Resampling in the following manner:

After selecting “Nonparametric Tests” from the analyze menu, the researcher clicks on “Two Independent Sample tests,” where the researcher finds an "Exact" button. This button in SPSS is used to conduct the process of Resampling, and allows the researcher to make a choice between the types of significance estimates. One such choice the researcher can make includes the method of "Monte Carlo," which is also a Bootstrapping and Resampling method.

Monday, June 15, 2009

Sample Size Calculation

A sample is a subset of the population. It is through samples that researchers are able to draw specific conclusions regarding the population. Sample size is the size of that sample. Sample size is very important in statistics.

Statistics Solutions can assist in choosing the correct sample size for your dissertation, thesis or research. Contact Statistics Solutions today for a free 30-minute consultation.

Sample size calculation ascertains the correct sample size that would represent the population as a whole. A larger sample size is required while making decisions when more information is needed. As the sample size increases, the information obtained has to be obtained with precision. The degree of precision may be measured in terms of the standard deviation of the mean. The standard deviation is inversely proportional to the square root of the sample size. Sample size calculation is very important in statistical inference and findings.






Determining Sample size:

There are many ways to determine the sample size. Sample size calculation for different statistical testing varies depending on the formulae used. Sample size calculation cannot be performed with only one method or technique.

Sample size calculation is legitimate for most relevant tests, like the t test, z test, f test, etc. To show this in an example, let us take an example of hypothesis testing.



Let us assume that Xi (i=1, 2, …n), where ‘n’ is the independent number of observations drawn from N (µ,σ2).



Here, H0: µ= 12 X'> i.e. there is no significant difference in the mean of the sample drawn from the population.



H1: µ= µ*, for some 'smallest significant difference' μ* >0.



While observing some significant differences, the smallest value can be considered.



To estimate our hypothesis, we must do as follows:

Zα = √n ( 12X-µ) / σ'> . Here, Zα is the value of standard normal distribution at α level of significance.



If the tabulated value of Zα > calculated value of Zα , then we accept H0 at α level of significance. Otherwise we reject it.

In order to determine the value of ‘n,’ we have the following formulae:
n= (Zα σ)2 / ( 12X-µ)'> 2

Sample size calculation depends on the different statistical tests that are to be carried out, because with a change in statistical tests, the results are also dissimilar. Depending on the size of the population or the accuracy of the result, the size of the sample in sample size calculation varies.



Sample size calculation depends on many factors that are more commonly known as qualitative factors. These are important to help calculate any kind of sample size calculation and determination. These factors are the importance of decision, the resource Constraints, the number of variables, the sample sizes used in similar studies, the nature of the research, and the nature of the analysis.



In qualitative research, the sample size in sample size calculation is usually small. Larger samples would be required for conclusive research, such as descriptive surveys. Again, if the data collected is on a large number of variables, then the samples should also be large.
In market research, sample size is used for problem solving research, problem identification research, TV, radio, print advertising, test-market audits, focus groups, etc.

Tuesday, June 9, 2009

Estimation

A statistical inference is basically a process that involves the inference of the data in a statistical manner. There are basically two types of statistical inferences, namely estimation and the test of the hypothesis.

Statistics Solutions can assist with estimation and sample size calculation, click here for a free consultation.

Estimation serves the purpose of determining the true value of the population that is based on the observations or the samples that are collected by sampling. To carry out estimation, the researcher needs to utilize certain statistics.

Estimation involves the use of two popular terms that a researcher should understand. The two terms that are used extensively by the researcher in estimation are the estimator and the estimate. These two terms, called the estimator and the estimate, can be explained with the help of an example. It is assumed in estimation that x1 x2 x3 (and so on) are the collection of the sample from the population having ‘s’ as their parameter. If the T=T(x) is a statistic then E(T(x))= s is the estimation. In this manner, estimation of the statistic is done. In this case of estimation, the estimator is the statistic T, and the estimate is the parameter called ‘s.’
It is important to understand the properties of estimators in estimation theory.

In estimation theory, unbiasedness is the first property that is assumed for an ideal estimator.
The Unbiasedness property of the estimators in estimation theory is basically those types of estimators that give their outcome as zero bias for all the values of the parameter. If the researcher considers the example above, then T in the theory of estimation is said to be unbiased only if its estimate is simply ‘s.’

The second property in estimation theory is that of the consistent estimators that involve the estimation that is consistent in nature. In other words, it can also be said that the consistent estimators in the theory of estimation should have a higher degree of concentration as the value of the random variable increases. In the theory of estimation, the sufficient condition of consistency explains that an estimator is supposed to be consistent only if the estimation of its expected value gives an unbiased estimate and the variance of the estimator is zero. In estimation, these two conditions are fulfilled only when the number of random variables tends to infinity.

There is another property for the ideal estimator in the theory of estimation called efficiency. According to the condition of this property in the theory of estimation, the consistent estimators should be distributed by normal distribution. This condition is introduced in the theory of estimation because there is some possibility that the estimators, which satisfy the sufficient conditions of consistency, may not be an efficient estimator.

The last property of the ideal estimator in the theory of estimation is the property of sufficiency. An estimator in the theory of estimation is said to be sufficient only if the joint conditional distribution function of the sample or the observation falls under the condition where T1 T2 T3 T4 (and so on) are the values under the function of the estimator ‘T.’ Thus, this joint conditional distribution in estimation should be independent of the parameter‘s.’

An estimator in estimation is considered to be the best estimator only if it is a minimum variance unbiased estimator (MVUE). By minimum variance in estimation, we mean that the estimator has less variability as compared to the other estimators.