Skip to main content

Section 25.4 Investigation 5.3: Near-Sightedness and Night Lights (cont.)

Exercises 25.4.1 The Study

Recall from Investigation 3.2 we examined a simplified version of the study published by Quinn, Shin, Maguire and Stone (1999) that examined the relationship between the type of lighting children were exposed to and subsequent eye refraction. Now we are able to look at 3 categories for each variable. Below is the two-way table and segmented bar graph from the study:
Eyes \ Lighting Β Β Β Β DarkΒ Β Β Β  Night light Room light Total
Far-sighted 40 39 12 91
Normal 114 115 22 251
Near-sighted 18 78 41 137
Total 172 232 75 479
Investigation 5.3 introduction image
Rather than considering these data as arising from three independent random samples from binomial processes, we can consider the 479 patients as one random sample, where we have cross-classified the subjects according to the lighting type and the eye condition. But the calculations for the chi-squared statistic can be used exactly the same way.
We will state the null and alternative hypotheses in terms of this association.
\(H_0\text{:}\) no association between type of lighting and eye condition in the population
\(H_a\text{:}\) there is an association
Note that the alternative hypothesis does not specify any particular kind of association (e.g., more near-sightedness with more light), just that the conditional distributions are not the same in the population. This is often referred to as a Chi-squared test of association.

1. Simulation Plan.

Outline a simulation that would be appropriate for this study design.
Solution.
We would need to randomly sample 479 individuals from a population with no association between the two variables. We could compute a chi-square statistic for each table and see how often we find a chi-square value like one in this study or more extreme.
We could choose to carry out a different simulation to model how these data were collected (one random sample, cross-classified, aka β€œmultinomial”), but when the technical conditions are met, we generally use the theoretical null distribution which will be the same if we fix the row and column totals. But be sure to keep the data collection methods in mind when you draw your final conclusions!

2. Expected Count Calculation.

Use the general formula to calculate how many of these 479 children you would expect to find in the Room light, far-sighted category.
Solution.
Room light and far sighted \(= 91 \times 75 / 479 \approx 14.25\text{.}\)

3. Check Conditions.

Do you think the chi-squared distribution is valid for this table? Explain how you are deciding.
Solution.
The smallest expected count is larger than 10, so the validity conditions are met.

4. Technology Output and Contributions.

Use technology to calculate the chi-squared statistic, verify the degrees of freedom, and find the p-value. Also display the chi-squared cell contributions; where do the largest differences lie?

Aside: Technology Detour.

Solution.
Applet output.
The darkness/myopia cell and the room light/myopia cell have the largest contributions. We observed a smaller rate of myopia in the darkness group and a higher rate of myopia in the room light group than we would have expected if there are no differences among the lighting populations.

5. Conclusions Paragraph.

Write a paragraph summarizing your conclusions, being sure to comment on significance, generalizability, and causation.
Solution.
See Summary Conclusion box.

Study Conclusions.

The segmented bar graph reveals that for the children in this sample the incidence of near-sightedness increases as the level of lighting increases. When we have a random sample with two categorical variables, we can perform a chi-squared test of association. Because the expected counts are large (smallest is 14.25 > 5), we can apply the chi-squared test to these data. The p-value of this chi-squared test is essentially zero, which says that if there were no association between eye condition and lighting in the population, then it’s virtually impossible for chance alone to produce a table in which the conditional distributions would differ by as much as they did in the actual study. Thus, the sample data provide overwhelming evidence that there is indeed an association between eye condition and lighting in the population of children like those in this study. A closer analysis of the table and the chi-squared calculation reveals that there are many fewer children with near-sightedness than would be expected in the β€œdarkness” group and many more children with near-sightedness than would be expected in the β€œroom light” group. But remember the main lesson of this study from Chapter 3 β€” we cannot draw a cause-and-effect conclusion between lighting and eye condition because this is an observational study. Several confounding variables could explain the observed association. For example, perhaps near-sighted children tend to have near-sighted parents who prefer to leave a light on because of their own vision difficulties, while also passing this genetic predisposition on to their children. We also have to be careful in generalizing from this sample to a larger population because the children were making voluntary visits to an eye doctor and were not selected at random from a larger population.
You should notice that the mechanics are exactly the same whether you are testing homogeneity of proportions, comparing the distributions across several populations/processes, or examining the association between two categorical variables. The difference lies in how the data were collected and therefore in the scope of conclusions that can be drawn.

Subsection 25.4.2 Practice Problem 5.3

The National Vital Statistics Reports provided data on gestation period for babies born in 2002. The following table classifies the births by the mother’s race and by the duration of the pregnancy:
Gestation \ Race White (non-Hispanic) Black (non-Hispanic) Hispanic
Pre-term (under 37 weeks) 251,132 101,423 99,510
Full term (37-42 weeks) 1,885,189 435,923 692,314
Post-term (over 42 weeks) 149,898 36,896 64,997

Checkpoint 25.4.1. Chi-squared Test of Association.

Consider these observations as a random sample from the birth process in the U.S. and conduct a chi-squared test of whether these data suggest an association between race and length of gestation period. Report the hypotheses, validity of technical conditions, sketch of sampling distribution, test statistic, and p-value. [Provide the details of your calculations and/or relevant computer output.] Summarize your conclusion.

Checkpoint 25.4.2. Largest Cell Contributions.

Which 2-3 of the nine cells in the table contribute the most to the calculation of the \(\Chi^2\) test statistic? Is the observed count lower or higher than the expected count in those cells? Summarize what this reveals about the association between race and length of gestation period.
You have attempted of activities on this page.