{"id":768,"date":"2015-03-30T19:30:00","date_gmt":"2015-03-30T23:30:00","guid":{"rendered":"https:\/\/www.causeweb.org\/sbi\/?p=768"},"modified":"2015-03-31T10:50:44","modified_gmt":"2015-03-31T14:50:44","slug":"why-even-bother-to-teach-the-normal-stuff","status":"publish","type":"post","link":"https:\/\/www.causeweb.org\/sbi\/?p=768","title":{"rendered":"Why Even Bother To Teach The Normal Stuff?"},"content":{"rendered":"<p><a href=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobb.png\"><img loading=\"lazy\" decoding=\"async\" class=\" size-full wp-image-769 alignleft\" src=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobb.png\" alt=\"GCobb\" width=\"150\" height=\"200\" \/><\/a><strong>George Cobb, Mount Holyoke College<\/strong><\/p>\n<p>I\u2019m writing to respond to a pair of questions we often hear from teachers who are considering a simulation-based alternative to the traditional normal-centric course: (1) \u201cIf the simulation-based approach is so great, why even bother to teach the normal-based stuff at all?\u201d (2) If I\u2019m going to include the normal-based stuff, do you have any suggestions about how to make the transition?\u00a0\u201d[pullquote]The \u201ctheory\u201d in what some of us call the \u201ctheory-based approach\u201d is the Central Limit Theorem, actually a cluster of theorems about convergence of sampling distributions to normal (Gaussian). [\/pullquote]<\/p>\n<p><!--more--><\/p>\n<p><strong>Why even bother to teach the normal-based stuff at all?<\/strong><\/p>\n<p>I offer four reasons:<\/p>\n<ol>\n<li>Empirical\/theoretical: When samples are large enough, many summary statistics have distributions that are roughly normal.<\/li>\n<li>Tradition: It\u2019s what almost everybody still does and so you need to know about it.<\/li>\n<li>Informal inference: z-scores and +\/- SEs are often useful as \u201cquick-and-dirty\u201d approximations.<\/li>\n<li>Extensions: Normal based methods generalize easily to comparing several groups at once (ANOVA) and fitting equations to data (regression).<\/li>\n<\/ol>\n<p>Speaking for myself, depending on the local situation, I\u2019d be inclined to keep these four reasons in mind, and deploy them as seems useful in response to student questions. If you have students who have seen the normal and related inference before, these questions may come up sooner rather than later, and may take the form of \u201cWhy bother with all this simulation, when you can just use a formula?\u201d To which I might respond, \u201cSometimes those formulas can give very wrong answers. We can use the simulation approach to understand when the formula can be trusted.\u201d[pullquote]For Stat 101, the bottom line is that the theorems suggest but do not guarantee that normal-based methods may give good approximations.[\/pullquote]<\/p>\n<p>But that\u2019s an aside. I\u2019m trying here to address an opposite point of view: \u201cIf simulation is so great \u2026?\u201d I hope (1) \u2013 (4) offer a useful approach to answering this question. A related challenge is to help students buy into the transition from simulation to theory-based.<\/p>\n<p><strong>Any thoughts about how to present the transition to theory-based methods?<\/strong><\/p>\n<p>This is an opportunity (= challenge) for the teacher to help students over a potentially difficult transition. In my experience, such transitions go most smoothly if they are spread out over time and introduced a little at a time. This gradual approach allows students with differing backgrounds and different learning styles enough flexibility to find their own best way. In that spirit, I\u2019ll set out one possible approach, one which I hope you can easily modify to suit your own situation. I\u2019ll start with an overview, then go into more detail.[pullquote]We can use the simulation approach to understand when the formula can be trusted.[\/pullquote]<\/p>\n<p>The \u201ctheory\u201d in what some of us call the \u201ctheory-based approach\u201d is the Central Limit Theorem, actually a cluster of theorems about convergence of sampling distributions to normal (Gaussian). In essence, and informally, these theorems say that when samples are large <em>enough<\/em>, sampling distributions of <em>some<\/em> statistics are <em>approximately<\/em> normal. Note the three weasel-words: large \u201c<em>enough<\/em>\u201d, \u201c<em>some<\/em>\u201d statistics, and \u201c<em>approximately<\/em>\u201d normal. For Stat 101, the bottom line is that the theorems suggest but do not guarantee that normal-based methods may give good approximations. (Making the three weasel-words precise has been a dominant 300-year research theme in probability theory.)<\/p>\n<p>For students in a Stat 101 course, I prefer to stay informal, to emphasize and rely on the three features of a distribution: shape, center, and variability. If the shape of the null distribution is roughly normal, and you know the center and spread, you can use a z-score to estimate a p-value. When it comes to the transition from simulation-based to normal-based, these three features play important albeit different roles.<\/p>\n<p><em>Shape<\/em>.<\/p>\n<p>Shape is a qualitative feature, which makes it easier for students to recognize by sight, but harder to quantify. In the spirit of making a gradual transition, you can call attention to the <em>shape<\/em> of simulated null distributions as soon as they first appear. For example, the null distribution for testing p = \u00bd will look roughly normal for n as small as 4. Figure 1 shows a simulated distribution with n = 16, p = \u00bd:<\/p>\n<p><a href=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg1.png\"><img loading=\"lazy\" decoding=\"async\" class=\"  wp-image-770 aligncenter\" src=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg1-300x123.png\" alt=\"GCobbg1\" width=\"414\" height=\"170\" srcset=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg1-300x123.png 300w, https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg1-624x257.png 624w, https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg1.png 900w\" sizes=\"auto, (max-width: 414px) 100vw, 414px\" \/><\/a>\u00a0<em><strong>Figure 1:\u00a0<\/strong><\/em>Simulated distribution of the number of heads in 16 tosses of a coin; \u00a0<em>n <\/em>= 16,\u00a0<em>p\u00a0<\/em>=\u00a0\u00bd<\/p>\n<p>Figure 2 shows another null distribution, this time with p = 1\/3 and n = 12:<\/p>\n<p><a href=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg2.png\"><img loading=\"lazy\" decoding=\"async\" class=\"  wp-image-771 aligncenter\" src=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg2-300x114.png\" alt=\"GCobbg2\" width=\"410\" height=\"156\" srcset=\"https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg2-300x114.png 300w, https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg2-624x238.png 624w, https:\/\/www.causeweb.org\/sbi\/wp-content\/uploads\/2015\/03\/GCobbg2.png 900w\" sizes=\"auto, (max-width: 410px) 100vw, 410px\" \/><\/a><\/p>\n<p><em><strong>Figure 2:\u00a0<\/strong><\/em>Simulated distribution of the proportion of successes when\u00a0\u00a0<em>n <\/em>= 12,\u00a0<em>p\u00a0<\/em>= 1\/3<\/p>\n<p>After the first few appearances of plots showing roughly normal shape, my inclination is to use shape as a major way to motivate the coming transition: \u201cThe same shape seems to be coming up over and over. Can we take advantage of that fact?\u201d<\/p>\n<p style=\"padding-left: 30px;\">Notes: (a) I\u2019ve used discrete distributions for illustration because many randomization-based expositions start with proportions, and I think it is is a good idea to call attention to the normal shape as soon as it starts to appear. (b) For discrete distributions (i) the smaller the distance between x-values relative to the SD, the more accurate the approximation, and (ii) there is an effective technical adjustment for discrete distributions if higher accuracy is an issue. (c) It\u2019s good to have this knowledge in reserve in case questions come up, but in the spirit of explaining the birds and bees to kids, my inclination is to be open to questions rather than deliver a gratuitous lecture in advance of curiosity.<\/p>\n<p><em>Center.<\/em><\/p>\n<p>For many students it is intuitive that the center of the null distribution of the sample proportion is the null value, so with students willing to speak up about their conjectures, all you need to do is to confirm that they have articulated a provable fact.<\/p>\n<p><em>Variability.<\/em><\/p>\n<p>Assume a normal shape for the null distribution of the sample proportion, with mean at the null value. All you need for a z-score and approximate p-value is the SD. Unfortunately, the formula for the SD of the sample proportion is almost impossible to make simple and intuitive for students at this level. (It can be done, slowly and empirically, but it takes about a week of simulation activities and discussion. Not a good use of time.) So the formula for the SD must be something of a <em>deus ex machina<\/em>. (Don\u2019t apologize: exploit the opportunity to evangelize for taking a more advanced course.)<\/p>\n<p>Once students have shape, center, and SD, then z-scores and the standard normal are just a step away.<\/p>\n<p><strong>Final thoughts about the transition.<\/strong><\/p>\n<p>My own inclination is to lean heavily on shape, center, and SD, and on the resulting z-scores and p-values, as above. The fact that the results are at best only approximate is critical, but the specific validity conditions are also at best only approximate, and their specificity can be a trap for students who want rules to memorize. What matters much more is the general principal that you can use simulation as a check on the normal approximation.<\/p>\n<p>&nbsp;<\/p>\n","protected":false},"excerpt":{"rendered":"<p>George Cobb, Mount Holyoke College I\u2019m writing to respond to a pair of questions we often hear from teachers who are considering a simulation-based alternative to the traditional normal-centric course: (1) \u201cIf the simulation-based approach is so great, why even bother to teach the normal-based stuff at all?\u201d (2) If I\u2019m going to include the [&hellip;]<\/p>\n","protected":false},"author":14,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[20],"tags":[],"class_list":["post-768","post","type-post","status-publish","format-standard","hentry","category-why-teach-normal-based-methods"],"_links":{"self":[{"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=\/wp\/v2\/posts\/768","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=\/wp\/v2\/users\/14"}],"replies":[{"embeddable":true,"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=768"}],"version-history":[{"count":16,"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=\/wp\/v2\/posts\/768\/revisions"}],"predecessor-version":[{"id":783,"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=\/wp\/v2\/posts\/768\/revisions\/783"}],"wp:attachment":[{"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=768"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=768"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.causeweb.org\/sbi\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=768"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}