You are on page 1of 58

EDU 5950-1 : EDUCATIONAL STATISTICS

CHAPTER 8 : INTRODUCTION TO
HYPOTHESIS TESTING
-GROUP 5DISEDIAKAN OLEH : NORASALIZA BT NOOR MOHAMED SALEH
(GS45266)
SALEHA BINTI MOHAMAD SALLEH
(GS45439)
MOHD QARIB BIN ZAINAL ABIDIN
(GS45863)
NURUL SYAHIRAH BINTI RUSLAN
(GS45293)
1

CHAPTER 8 : INTRODUCTION TO HYPOTHESIS


TESTING
8.1 The Logic of Hypothesis Testing
8.2 Uncertainty and Errors in Hypothesis
8.3 More About Hypothesis Tests
8.4 Directional (One-Tailed) Hypothesis Tests
8.5 Concerns About Hypothesis Testing : Measuring Effect
Size
8.6 Statistical Power
2

8.1 THE
TESTING

LOGIC

OF

HYPOTHESIS

Introduction
Hypothesis testing is a statistical procedure that allows researcher to
use sample data to draw inferences about the population of interest.
Hypothesis testing is one of the most commonly used inferential
procedures
Although the details of a hypothesis test change from one situation to
another, the general process remains constant
Hypothesis test is a combine of the concepts z-scores, probability, and
the distribution of sample

A hypothesis test is a statistical method that uses sample data to


evaluate a hypothesis about the population
3

HYPOTHESIS
ABOUT
POPULATION
(Concerns the
value of a
population
parameter)

THE LOGIC UNDERLYING


THE HYPOTHESIS-TESTING
PROCEDURE

BUT IF THERE IS A BIG


DISCREPANCY BETWEEN
THE DATA AND THE
PREDICTION, THEN WE
DECIDE THE HYPOTHESIS
IS WRONG

USE THE HYPOTHESIS TO


PREDICT THE
CHARACTEERISTICS THAT
THE SAMPLE SHOULD
HAVE (The sample should
be similar to the
population always
expect a certain amount
of error)

IF THE SAMPLE MEAN IS


CONSISTENT WITH THE
PREDICTION, THEN WE
CONCLUDE THAT THE
HYPOTHESIS IS REASONANBLE

COMPARE THE OBTAINED


SAMPLE DATA WITH THE
PREDICTION THAT WAS MADE
FROM THE HYPOTHESIS

OBTAIN A
RANDOM
SAMPLE FROM
THE
POPULATION

Depending on the type of research and the type of data, the details
of the hypothesis test change from one research situation to another
We examine a hypothesis test as it applies to the simplest possible
situation using a sample mean to test a hypothesis about a
population mean
Hypothesis testing is an inferential procedure that uses the data
from a sample to draw a general conclusion about the population
The procedure begins with a hypothesis about an unknown
population
The a sample is selected, and the sample data provide evidence that
either supports or refutes the hypothesis
5

In this chapter, we introduced hypothesis testing using the


simple situation in which a sample mean is used to test a
hypothesis about an unknown population mean ; usually the
mean for a population that has received a treatment
The unknown population : This is the set of individuals as they exist before.
Note that the unknown population, after treatment, is the focus of the research
question. Specifically, the purpose of the research is determine what would
happen if the treatment were administered to every individual in the population .

The question is to determine whether the treatment has an


effect on the population mean (see Figure 8.1)
The research study involves selecting a sample from the
original population, administering the treatment to the
sample, and then recording scores for the individuals in the
treated sample (see Figure 8.2)
6

The basics experimental situation for hypothesis testing.


It is assumed that the parameter is known for the
population before treatment. The purpose of the
experiment is to determine whether the treatment has an
effect on the population mean.

FIGURE 8.1

From the point of view of the hypothesis test, the entire population receives
the treatment and then a sample is selected from the treated population. In
the actual research study, a sample is selected from the original population
and the treatment is administered to the sample. From the either perspective,
the result is a treated sample that represent the treated population.

FIGURE 8.2

FOUR STEPS OF A HYPOTHESIS TEST :


1) STATE THE HYPOTHESIS
2) SET THE CRITERIA FOR A DECISION
3) COLLECT DATA AND COMPUTE SAMPLE
STATISTICS
4) MAKE A DECISION
9

STEP 1 : STATE THE HYPOTHESIS


The goal of inferential statistics is to make general statement about the
population by using sample data. Therefore, when testing hypothesis, we
make our predictions about the population parameters.
The process of hypothesis testing begins by stating a hypothesis about the
unknown population. Actually there have two opposing hypothesis :
i.

Null hypothesis (H) :

States that in the general population there is no change, no difference, or no relationship


(nothing happened, hence the name null)
The null hypothesis is identified by the symbol H (H stands for hypothesis, and the zero
subscript indicates that this is the zero-effect hypothesis)
In the context of an experiment, H predicts that the independent variable (treatment) has no
effect on the variable (scores) for the population
In symbols, this hypothesis is
H : with stimulation = 80

(Even with the stimulation, the mean test score is still 80)

10

Continued.
Alternative hypothesis (H) :

ii.

The alternative hypothesis is the opposite of the null hypothesis, and it is called the
scientific, or alternative, hypothesis (H).

States that there is a change, a difference, or a relationship for the general population.

In the context of an experiment, H predicts the independent variable (treatment) does


have an effect on the dependent variable.

For this example, the alternative hypothesis states the stimulations does have an effect
on mathematical skill for the population and will cause a change in the mean score

In symbols, the alternative hypothesis is represented as


H : with stimulation 80 (with the stimulation, the mean test score is different from 80)

The null hypothesis and the alternative hypothesis are mutually exclusive and
exhaustive. They cannot both be true, and one of them must be true. The data
determine which one should be rejected
11

STEP 2 : SET THE CRITERIA FOR A


DECISION
Distribution of sample means is divided
Those likely if H is true
Those very unlikely if H is true

Alpha Level, or level of significance is a probability value


used to define very unlikely
Critical region is composed of the extreme sample values
that are very unlikely (as defined by the alpha level)
Boundaries of critical region are determined by alpha level.
If sample data fall in the critical region, the null hypothesis
is rejected
12

The set of potential samples is divided into those that are likely to be
obtained and those that are very unlikely to be obtained if the null
hypothesis is true

FIGURE 8.3

13

The critical region (very unlikely outcomes) for = .05.

FIGURE 8.4

14

STEP 3 : COLLECT DATA AND COMPUTE SAMPLE


STATISTIC
Collect the data, and compute the test statistic. The sample mean is
transformed into a -score by the formula :
=M-
M
In the formula, the value of the sample mean (M) is obtained from the
sample data. The value of is obtained from the null hypothesis
The -score test statistic identifies the location of the sample mean in the
distribution of sample means.
The -score formula can be expressed in words as follows :
= sample mean hypothesis population mean
standard error between M and
15

STEP 4 : MAKE A DECISION


If the obtained -score is in the critical region,
reject H because it is very unlikely that these data
would be obtained if H were true. In this case,
conclude that the treatment has changed the
population mean.
If the -score is not in the critical region, fail to
reject H because the data are not significantly
different from the null hypothesis. In this case, the
data do not provide sufficient evidence to indicate
that the treatment has had an effect.

16

8.2 UNCERTAINTY AND ERRORS IN HYPOTHESIS


TESTING
Hypothesis testing is an inferential process, which means that it uses
limited information as the basis for reaching a general conclusion
Whatever decision is reached in a hypothesis test, there is always a risk of
making the incorrect decision. There are two types of errors that can be
committed :
i.

A Type I error : is defined as rejecting a true H. This is a serious error


because it results in falsely reporting a treatment effect.
The risk of a Type I error
is determined by the alpha level
and therefore, is under the experimenters control.

ii.

A Type II error : is defined as the failure to reject a false H. In this case, the
experiment fails to detect an effect that actually occurred. The probability of a Type II
error cannot be specified as a single value and depends in part on the size of the
treatment effect. It is
identified by the symbol (beta).
17

A TYPE I ERROR
Occurs when a researcher rejects a null hypothesis that is actually
true.
In a typical research situation, a Type I error means that the
researcher concludes that a treatment does have an effect when,
in fact, it has no effect.
In most research situations, the consequences of a Type I error
can be very serious because the researcher has rejected the null
hypothesis and believes that the treatment has a real effect.
A Type I error means that this is a false report thus it lead to false
report in the scientific literature.
Alpha Level is the probability that a test will lead to a Type I error.

18

A TYPE II ERROR
Occurs when a researcher fails to reject a null
hypothesis that is really fast (Is the failure to reject
a false null hypothesis)
Means that a treatment effect really exists but
the hypothesis test fails to detect it (hypothesis
test has failed to detect a real treatment effect)

19

TABLE 8.1 : POSSIBLE OUTCOMES OF A


STATISTICAL DECISION

20

The locations of the critical region boundaries for three


different levels of significance : = .05, = .01, and = .
001

FIGURE 8.5

21

8.3 MORE ABOUT HYPOTHESIS TEST


A summary of the hypothesis test
Step 1 : State the hypothesis and select alpha level
Ex : General population without treatment, has an average test scores
of = 80 with = 20
hypothesis are :
H : with stimulation = 80 (the brain stimulation has no effect)
H : with stimulation 80 ( the brain stimulation has no effect)

we set = .05
22

Step 2 : Locate the critical region


Normal distribution with = .05

z = 1.96

Step 3 : Compute the test statistic (the z-scores)


Sample mean, M = 89
M=4
Z= M
2.25
M

Participants, n = 25 standard error

= 89 80
4

= 9 =
4

Step 4 : Make a decision


Z-scores is in the critical region, which mean that this sample mean is
very unlikely if the null hypothesis is true. Therefore, we reject the null
hypothesis and conclude that the brain stimulation did have an effect
on the students test score.
23

SIGNIFICANT OR STATISTICALLY SIGNIFICANT


If it is very unlikely to occur when the null hypothesis
is true
That is the result is sufficient to reject the null
hypothesis
Thus, a treatment has a significant effect if the
decision from the hypothesis is to reject H
24

The Normal Curve of Means Differences of All


Possible Outcomes if The Null Hypothesis is True

Reject the
Null
Hypothesis

High Probability
Values if the Null
Hypothesis is True

Extremely low Alpha=.025


Alpha=.025
Probability Values
if Null Hypothesis
Two-Tailed Test
is True (Critical
Region)

Reject the
Null
Hypothesis

Extremely low
Probability Values
if Null Hypothesis
is True (Critical
Region)
25

26

Assumptions For
Hypothesis test with z
Scores
Random Sampling :
- Selected randomly
- to generalize our findings from the sample to the
population. The sample must be representative of the
population from which has been drawn.
- Random sampling helps to ensure that it is representative.
Independent Observation :
- Two observation are independent if there is no
consistent, predictable relationship between the first
observation and the second.
- Two events ( or observation ) are independent if the
occurrence of the first event has no effect on the
probability of the second event.
27

The Value of is unchanged by the treatment :


- A critical part of the z scores formula in a hypotesis test is the standard
error M
- Must know the sample size (n)
-Assume the standard deviation for the unknown population.

Normal Sampling Distribution :


- To evaluate hypothesis with z scores, have used the unit
normal table to identify the critical region.
- The table can be used only if the distribution of sample
means is normal

28

Z Scores in a Hypothesis Test


The Z score formula as a ratio
Z = M = sample mean - hypothesized population mean
M
standard error between M and
FACTOR INFLUENCE A HYPOTHESIS TEST
The variability of the scores :
The study used a sample of n =25 students and produced a mean test
score of M= 89. Based on these results, the hypothesis test concluded
that brain stimulation near the parietal lobe has a significant effect on
the ability to learn new mathematical skills.
Standard deviation is = 20. With a sample of n =25, this procedure
standard error of M = 4 points and a significant z scores of z = 2.25.
Now consider what happen if the standard deviation is increased to =
30 . Standard error become M = 30 = 6 points , new z score becomes :
25
29

Continue
z = M = 89 80 = 9 = 1.50
M
6
6
Z scores is no longer beyond the critical boundary of
1.96. So the statistical decision is to fail to reject
the null hypothesis

30

Continue.
THE NUMBER OF SCORES IN THE SAMPLE
Used the sample n= 25, student obtained a standard error of

M = 20

= 4 points, and a significant z

25
Score of z = 2.25.What happens if increase the sample size to n= 100 students,
the standard error becomes

= 20 = 2 points.
100

Z score become :
z = M = 89 80 = 9 = 4.50
M
2
2
31

8.4 DIRECTIONAL (ONE TAILED )


HYPOTHESIS TEST
Definition :
In a directional hypothesis or one tailed test, the
statistical hypothesis ( H0 and H1 ) specify either
an increase or a decrease in the population mean.
That is they make a statement about the direction
of the effect.

32

The Hypothesis For A Directional Test


Two Hypothesis Would State :
H0 : Test score are not increased ( the treatment does
not work)
H1 : Test Score are increased ( the treatment works as
predicted )

33

Continued..
Ex : general population has an average score of
= 80 and H1, state that test scores will be increased
by the brain stimulation.
H1 : >80 ( With the stimulation, the average score is
greater than 80 )
Null Hypothesis stimulation , does not the increase
score

H0 : 80 ( With the stimulation, the average score


is not greater than 80 )
34

CRITICAL REGION FOR DIRECTIONAL TEST

Reject H0
Data indicate
that H0 is
wrong

M
=4

M
80
Z
0

1.6
5
35

Has an expected value of =80 (from H0 ),


sample of n = 25, standard error of

M = 20
25

If Untreated students average = 80.


Critical region one tailed test
Propotions specified by alpha level is not divided
between two tails, but rather is contained entirely
in one tail.
Using = .05 ( whole 5% ) is located in one tail.
The z scores boundary for the critical region is z
= 1.65, ( column C- .05 ) the tail of unit table.
36

EX : mean of M =87 , for the 25


participants who received the brain
stimulation. Sample mean
corresponds to a z-score ;
z = M = 87 -80 = 7 = 1.75

4
4
z, score of z = 1.75 is in
critical region for a one-tailed
test.
The Stimulation produced a
significant increase in scores, z =
1.75, p .05

37

8.5 Hypothesis Testing Concerns:


Measuring Effect Size

38

8.5 Hypothesis Testing Concerns: Measuring


Effect Size
Although commonly used, some researchers are concerned about
hypothesis testing
Focus of test is data, not hypothesis
Significant effects are not always substantial

Effect size measures the absolute magnitude of a treatment effect,


independent of sample size
Cohens d measures effect size simply and directly in a
standardized way

39

8.5 Hypothesis Testing Concerns: Measuring


Effect Size
mean difference
treatment no treatment
Cohen' s d

standard deviation

Magnitude of d

Evaluation of Effect
Size

d = 0.2

Small effect

d = 0.5

Medium effect

d = 0.8

Large effect
40

FIGURE 8.8 The appearance of a 15-point treatment effect in two


different situations. In part (a), the standard deviation is = 100
and the 15-point effect is relatively small. In part (b), the standard
deviation is = 15 and the 15-point effect is relatively large.
Cohens d uses the standard deviation to help measure effect size.
41

8.5 Hypothesis Testing Concerns: Measuring


Effect Size
Example :
A researcher selects a sample from a population with
mean is 45 and standard deviation is 8. A treatment is
administered to the sample and, after treatment, the
sample mean is found to be M = 47. Computer Cohens
d to measure the size of the treatment effect.
42

8.5 Hypothesis Testing Concerns: Measuring


Effect Size
mean difference
treatment no treatment
Cohen' s d

standard deviation

Cohens d = 47 45
8
= 2
8
= 0.25

43

8.6 : STATISTICAL POWER

Definition
The power of statistical test is the probability that
the test will correctly reject false null hypothesis.
The power of a test depends on a variety of
factors including the size of the treatment effect
and the size of the sample.
44

Whenever a treatment has an effect, there are two


possible outcomes :i) Fail to reject Ho
-Type II error with a probability as p=B
ii) Reject Ho

-Probability of 1-B

Rejecting Ho , when there is a real effect, is the power of


effect

The power of hypothesis test is equal to 1-B


45

Example
Find the probability of a Type II error
Power of test is 70%
( 1 B)
1 - 70%
= 30%
Type II error is 30%
46

Researcher usually calculate the power of


hypothesis test before conduct research study.
They can determine the probability that the result
will be significant (Ho)
To calculate power, make assumption about a variety of
factors that influence the outcome the hypothesis test.

47

Example 8.5 (page 233)


-Mean :80
-Standard deviation : 10
-Researcher plan to select n = 25 individuals
from population and administer to each individual.
-Expected the treatment will have an 8-point effect

48

Two possible
outcomes :
1.If
the
null
hypothesis is true,
there
is
no
treatment effect.
2.If the researchers
expectation is
correct, there is 8
point effect.

49

The left-hand side shows the distribution of sample means that


would occur if the null hypothesis is true.
The right-hand side shows the distribution of
sample means that would be obtained if there were an 8-point
treatment effect.
50

Standard error
= 10 = 10
25
5

=2

Calculate exact value for the power of test,


i) determine the portion of distribution.
ii) Locate exact boundary for the critical region
iii) Find the probability in unit normal table.
51

Left hand side:


critical boundary of z = +1.96 correspond to a
location is above mean = 80 distance equal to
1.96
= 1.96 (2) = 3.92 points
Z= 80 + 3.92 = 83.92
*Any sample mean that greater than M= 83.92 is
in critical region and would lead to rejecting null
hypothesis .

52

For the right hand side


Population mean,
= 88
Sample mean, M= 83.92
Z= M
= 83.92 88 = -4.08 = -2.04
2
2
= Look up z=-2.04 in normal table
p = 0.9793 or (97.93%)
-It is in the critical region and we reject the null hypothesis
- The power of test is 97.93%
53

Power and effect size


As the effect size increases, the probability of rejecting H 0 also increases, the power of test
increases.
Example
= 4 point of treatment effect
power test = only about 50%
Z= M
2

= 83.92 84 = -0.08 = -0.04


2

Z = -0.04
Exact power of test is p=0.5160 or 51.60%
-Probability of rejecting Ho increases, the power of test increases.
54

Other factors that affect power


1. Sample size
n = 25 , Power size = 97.93%
n=4.power size will be less than 50 %
*If probability is too small, researcher have option to
increase the sample size to increase power.

55

Other factors that affect power

2. Alpha level
-Reduce the alpha level will also reduce the power of
test.
Example
alpha from .05 to .01
-The boundary is out to 2.58 , moved farther to the
right
-It is smaller portion so lower probability of

56

Other factors that affect power


3. One tailed versus two tailed test
- Changing to a one-tailed test would move the
critical boundary to the left value of z = 1.65

- Moving boundary to the left cause the large


proportion of treatment distribution to be in
critical region so it will increase the power of test.
57

Summary
The power of hypothesis test is defined as the probability that
the test will correctly reject the null hypothesis.
As the size of treatment effect increases, statistical power
increases.
Several factors that influenced power and can be controlled by
the experimenter :
i) A large sample results in more power than a small sample
ii) Increasing the alpha level increases the power
iii) A one tailed test has a greater power than a two-tailed test.
58

You might also like