Okay, in this video and subsequent videos, we're going to deal with this question. What do we do if our assumptions of equal variance and normality are not met by our data. There potentially three steps to addressing this question. And the first is simply ask, well, is the way in which our data violate the assumptions? Is it a way that we need to worry about a serious violation? And we're going to deal with that first step in this video. In later videos we're going to deal with steps 2 and 3. So step 2 asks, Okay, If our violation is a serious one, can we deal with it by doing something called transforming our data? And sometimes we can, and sometimes we can't. If we can't, if we cannot deal with the violated assumption by transforming the data, then we move on to step three and ask ourselves whether or not there'll be some other way of analyzing our data. There might be more appropriate. On this slide and the following slide, we're going to outline the situations in which violating the assumption of equal variance and normality are actually serious problems. So this slide here deals with the case of equal variance. So when is violating the assumption of equal variance actually a serious problem? When is it something that we need to worry about? Well, in order to answer that question, It's really helpful to understand what violating this assumption tends to do. When we have unequal variances. What that tends to do is it tends to artificially decrease our p-value. And what that means is that when we have unequal variance, this tends to increase our probability of making a type one error. In other words, it tends to increase our probability of concluding that there is some kind of significant effect in our data when there isn't. Or in other words, unequal variance tends to increase the probability of finding a false positive. This is something you want to avoid. So when is unequal variance most a problem? I think covariance is a most acute problem. When we have one exceptionally large, one exceptionally large variance. Or if the variance tends to increase or decrease systematically with our fitted values. What do I mean by that? I want you to think back to the previous video where we were simulating data and we were looking at a plot of our residuals that allowed us to test the assumption of equal variance. And in that figure are in that plot, we had our fitted values along the x-axis and we had a residuals along the y-axis. And what the statement here is saying about variance increasing or decreasing systematically with the fitted value. What that means is, is that if we were to move from left to right along our x-axis. So moving along our fitted values, the variance is a big problem. If the variance and our residuals, or the variation or residuals tends to increase systematically or decrease systematically as you move from left to right along that x axis. And we saw an example of that in our previous video. Third, unequal variance is also big problem in our sample sizes vary between our groups. When is unequal variance not a problem? Well, the first situation is that unequal variance is not a problem if the heterogeneity or if the unequal illness is actually simply random variation. What do we mean by that? Again, I want you to think back to the previous video where we generated data that perfectly met the assumptions of our tests. But even in those cases, even when we generated data that were perfectly applicable to our tests, we could still see that our variance was not identical between our treatments. There was some amount of variation among our treatments in terms of the variance. That's what I mean when I say that the heterogeneity or the unequal illness and the variation is just random variation. In other words, the lack of equal variances not a problem if it's simply due to random variation from sapling error, which is exactly what happened in the previous slide. Or sorry, in the previous video. So if you are looking at your plots in order to test the assumption of equal variance. And you see differences in variance that are pretty similar to the kinds that we saw in the good scenarios in our previous video. Then that falls into this category. That heterogeneity is just random variation and we do not need to worry about it. The other situation where we do not need to worry about an unequal variance is if our significant, if our effect is not significant. Anyway. The logic for that is simply that if unequal variance tends to artificially decrease our p-value, and our p-value is still greater than 0.05. Then in that situation, unequal variances not going to artificially cause us to reject our null hypothesis. And so that's a situation where we do not need to worry about unequal variance affecting our conclusions. I want to note, however, that perspective, it really is a p-value centric perspective. And something I really want to encourage everyone to do is to think about data and you analyze your data in a way that takes effect size into account. And unequal, unequal variance can influence your estimates of your effect sizes. So even if we don't change our conclusions in terms of p-values. Unequal variance still might influence or conclusions in terms of calculating effect sizes. Okay, normality. When is it a serious problem? Well, again, to get a sense of when normalities a serious problem, we should have some sense of how lack of normality introduces issues. Violations of normality tend to have unpredictable effects on conclusions. In other words, they can either increase or decrease your chances of making a type 1 error. It's really important to note, however, that the effective normality is going to be, is going to be very small. However, unless we have a huge violation of normality. That's why I have this here on this slide saying, however, these effects are going to be very small unless the violation is very large. So when is lack of normality most serious? It's most serious when the when our data are clearly not normally distributed. And what I mean by that is that saying that there's a true qualitative difference between our data and a normal distribution and are given as an example here that the residuals are, for example, actually bimodal distributed. Bimodal means you actually have two humps. If you were to plot your data in a histogram, a normal distribution should have a nice, we'll have a particular shape. But it's easily recognized as having one hump that's in the middle. And so we have a one humped symmetrical distribution. If our residuals, I'll generally like that, then we're in good shape. If instead our residuals have two humps. If they are by modally distributed, then that's a case where we would start to really worry about normality. Lack of normality is least serious. If our p-values are either much greater than or much less than 0.05. And the logic there is that if normality has a relatively small influence on your p-value, then if your p-value is already far away from 0.05, then it's unlikely that a lack of normality is going to cause your p-value to change from being just over 2.05 to just under 0.05 or reverse. Okay? It's really when your p-value is just around 0.05 and we have a large deviation from normality that we'd most, that we would most worry. Again though, I want to highlight that this perspective is a, a p-value centric perspective on analyzing our data. And we still want to exercise some caution when we're calculating effect sizes. Again, however, that effect on our effect sizes will likely be pretty small unless you have a very large deviation from normality. I want to close just with a quick comment about how to actually test the assumptions of equal variance and normality. There are formal tests for homogeneity of variance. So things like Bartlett's task Hawkins test flowers and cones test. And there are tests to check for the normality of residuals. Tests like the Anderson-Darling test. But want to point out though, is that these tests only tell you about whether or not there is statistical significance of the violation. These tests do not actually tell you whether or not the violation matters. So for example, let's imagine we had a very large data set and we performed an Anderson-Darling test to check the normality of our residuals. And we found that our data or our residuals were not normally distributed according to the Anderson-Darling test. That might be reason for alarm. If we actually plotted our residuals, however, we might see that yes, they're not normally distributed formerly, but they're actually very, very close to be normally distributed. In that case, from our discussion so far, it should be clear that we do not need to worry about a lack of normality because a small deviation from normality is very unlikely to cause any meaningful changes in our analysis. Likewise, if you have a very small data set, running one of these tests might not allow you to actually detect a problem in your data, just because these tests don't actually have the power to detect it. Whereas if you were to plot your data, it might become immediately clear that there is a real problem with your data. For example, you might have an outlier that might really catch your attention. So the main point here is yes, there are formal tests that can be used to test the assumption of equal variance and normality. But I do not recommend them. The reason I've been showing you in videos how to check your assumptions by visualizing residuals is because I feel, and many people feel that that is the best way to check your assumptions of equal variance and normality. So in general, I would expect that examining your residuals will be the best way to check these assumptions. And that's what we have to say right now about when violating the assumptions actually matters. I hope this video has been useful. Thank you.