Section25.4Investigation 5.3: Near-Sightedness and Night Lights (cont.)
Exercises25.4.1The Study
Recall from Investigation 3.2 we examined a simplified version of the study published by Quinn, Shin, Maguire and Stone (1999) that examined the relationship between the type of lighting children were exposed to and subsequent eye refraction. Now we are able to look at 3 categories for each variable. Below is the two-way table and segmented bar graph from the study:
Rather than considering these data as arising from three independent random samples from binomial processes, we can consider the 479 patients as one random sample, where we have cross-classified the subjects according to the lighting type and the eye condition. But the calculations for the chi-squared statistic can be used exactly the same way.
Note that the alternative hypothesis does not specify any particular kind of association (e.g., more near-sightedness with more light), just that the conditional distributions are not the same in the population. This is often referred to as a Chi-squared test of association.
We would need to randomly sample 479 individuals from a population with no association between the two variables. We could compute a chi-square statistic for each table and see how often we find a chi-square value like one in this study or more extreme.
We could choose to carry out a different simulation to model how these data were collected (one random sample, cross-classified, aka βmultinomialβ), but when the technical conditions are met, we generally use the theoretical null distribution which will be the same if we fix the row and column totals. But be sure to keep the data collection methods in mind when you draw your final conclusions!
Use technology to calculate the chi-squared statistic, verify the degrees of freedom, and find the p-value. Also display the chi-squared cell contributions; where do the largest differences lie?
The darkness/myopia cell and the room light/myopia cell have the largest contributions. We observed a smaller rate of myopia in the darkness group and a higher rate of myopia in the room light group than we would have expected if there are no differences among the lighting populations.
The segmented bar graph reveals that for the children in this sample the incidence of near-sightedness increases as the level of lighting increases. When we have a random sample with two categorical variables, we can perform a chi-squared test of association. Because the expected counts are large (smallest is 14.25 > 5), we can apply the chi-squared test to these data. The p-value of this chi-squared test is essentially zero, which says that if there were no association between eye condition and lighting in the population, then itβs virtually impossible for chance alone to produce a table in which the conditional distributions would differ by as much as they did in the actual study. Thus, the sample data provide overwhelming evidence that there is indeed an association between eye condition and lighting in the population of children like those in this study. A closer analysis of the table and the chi-squared calculation reveals that there are many fewer children with near-sightedness than would be expected in the βdarknessβ group and many more children with near-sightedness than would be expected in the βroom lightβ group. But remember the main lesson of this study from Chapter 3 β we cannot draw a cause-and-effect conclusion between lighting and eye condition because this is an observational study. Several confounding variables could explain the observed association. For example, perhaps near-sighted children tend to have near-sighted parents who prefer to leave a light on because of their own vision difficulties, while also passing this genetic predisposition on to their children. We also have to be careful in generalizing from this sample to a larger population because the children were making voluntary visits to an eye doctor and were not selected at random from a larger population.
You should notice that the mechanics are exactly the same whether you are testing homogeneity of proportions, comparing the distributions across several populations/processes, or examining the association between two categorical variables. The difference lies in how the data were collected and therefore in the scope of conclusions that can be drawn.
The National Vital Statistics Reports provided data on gestation period for babies born in 2002. The following table classifies the births by the motherβs race and by the duration of the pregnancy:
Consider these observations as a random sample from the birth process in the U.S. and conduct a chi-squared test of whether these data suggest an association between race and length of gestation period. Report the hypotheses, validity of technical conditions, sketch of sampling distribution, test statistic, and p-value. [Provide the details of your calculations and/or relevant computer output.] Summarize your conclusion.
Which 2-3 of the nine cells in the table contribute the most to the calculation of the \(\Chi^2\) test statistic? Is the observed count lower or higher than the expected count in those cells? Summarize what this reveals about the association between race and length of gestation period.