Okay, In this video, we're going to discuss the assumptions for one-way ANOVA or one factor general linear model. They have the same assumptions. All statistical tests make assumptions. The assumptions come in various forms and they address different aspects of our being able to make inferences from an analysis. A number of assumptions that we're going to discuss today. For example, equal variance and normality. Those assumptions are tied up in our ability to Estimate standard errors and confidence intervals and p-values accurately. So ANOVA, or one-factor jailing your model has four main assumptions. They are not all equally important. So I've got them listed here in order of decreasing importance. So the most important assumption is random sampling. And then the next most important is independence. And the next most important is homogeneity of variances. And then finally, normality. Normality is the least important assumption. These first two assumptions relate to that the design of the study. So these assumptions have to be checked, not with your data, but by thinking about the experimental design and deciding whether or not your experimental design actually is appropriate for this type of analysis. Just to clarify where I say random sampling here, what I'm referring to is the random allocation of subjects to our various treatments. Okay? So this does not refer to randomly sampling individuals from a population. Although we've talked about in other videos about why that is important. It's strictly speaking, random sampling from a population is not one of the assumptions to conduct a one-way ANOVA, but random allocation of subjects to the various treatments is, okay. So that is our first and most important assumption we talked in other videos, but why it's important? The second most important assumption refers to independence. I'm not going to say much about independence in this video because we discussed independence in depth in our series of videos at discussed pseudo replication. So if independence is unclear to you, I suggest that you go and watch those videos in order to learn exactly what we mean. What I'm going to do right now is just highlight. Some are two slides that I pulled out from that series of videos. Just to remind you of some of the essential features of independence with respect to one factor, general linear models. Okay? So here we have two different experimental designs. Neither of them violate the assumption of independence. These letters here I have a particular meaning. And I've noted there meaning down here, where letters, or sorry, each letter here represents a measurement. And if two measurements share a letter, then that means that those two measurements are not independent of one another. So what I want you to notice here is that in this fairly typical experimental design for one factor general linear model, all of the data points are independent. So this might come from an experiment where we might have randomly selected individuals from a population. And then gone on to randomly assign them to each of the various treatments that they would go into. And so you can see here that importantly, all of the data points within each of our treatments are independent. And so because of that, because the data within the treatments are all independent, these data meet the assumption of independence for this kind of test. Let's look at this experimental design now on the right, you'll, you'll see here that there are data points, are measurements that are not independent. So here we have an a and an a and an a. So these three data points are not independent of one another. But that kind of non-independence does not violate the assumptions of independence for this type of design. Okay? What I mean is if we were to analyze these data with a one-factor general linear model. Because he important assumption of independence for a one factor generally a model is that the data within each of our treatments are independent. And you can see here that that is true. Okay. Because none of the, none of the measurements within each of our treatments share a letter. Here's an example of a case where our data would violate the assumption of independence that you can see here. Now we have multiple measurements that share letters within our various treatments. So we could not analyze data like these directly with a one-factor gel linear model or one-way nova, we'd have to do something additional to help the data meet the assumptions for. So for example, what we could do is we could take the average of the data points that share letters and then use those averages in our analysis as one example. So just be clear, we take the average of a, B, average of C, and so on. Okay? So those are the most important assumptions that are assessed by thinking about the experimental design. Our last two assumptions relates to the data. And so we check these assumptions by looking at the qualities of the data. And specifically, we test these assumptions by examining what's called the residuals. So before we go any further, we need to define what a residual is. It turns out you've already seen a residual in these videos. We've just not use that term yet. Okay, so to illustrate what a residual is, let's imagine we have an experiment where we're looking at the factor three different fertilizers on crop yield. And for each R3 fertilizers, we have ten fields. So we have a data set with 30 independent data points. Okay? So here is what our experiment looks like. We have 10 data points for fertilizer, a tent for B, ten for C. And when we run one way ANOVA or one-factor linear model, you will recall that one of the things that happens as our software is performing the analysis is it will fit a mean value in each of our three different levels. So mean value for a, for B, and for C. And then it will calculate the distance between that mean value and each of the data points. Okay? Um, so this is just to say, say this is a mean value, can also call it a fitted value. And I said that we calculate the distance between that fitted value or the mean for our given treatment and each of the data points. And we did that to calculate our sum of squares. That distance also has a particular name which is referred to as a residual. So it's these distances between our fitted values or our means and our data points. It's those distances that we assess when we're trying, when we check our assumptions of equal variance and normality. Okay? So for our assumption of homogeneity of variance, what we're referring to is the assumption that the variance of the residuals is similar between our various groups. So what that means is if we were to calculate all of these residuals. So there's a residual, there's a residual, there's a residual, and so on. And if we looked at the distribution of those residuals for treatment a, the variance in those residuals would be similar for treatment a and B and C. Okay, and we assess that assumption visually by looking at a plot of the residuals. And the plot looks something like this. Where we have along the x-axis here is the fitted value. So the fitted value here refers to, oops, wrong direction. Fitted values refer to these fitted values here. So the mean values for each of our treatments. Okay? So don't want you to notice here that the fitted value for fertilizer B is equal to 4. The fitted value for fertilizer C is equal to about 4.7 or so. And the fitted value for fertilizer a is just over 5.5. Okay? So what you can see here is that we have these points line at the values for which corresponds to the mean value for fertilizer B. Fitted value of about 4.5, and that corresponds to the mean value for fertilizer. See, add another fitted value for fertilizer, right around 5.5, which corresponds to the mean value of fertilizer a. Okay? And what these points, so, sorry. So these points here represent the residuals associated with fertilizer B. These represent the residuals for fertilizer C, and likewise for fertilizer a. Okay? And what we're looking for. Is to see whether or not the overall variation in these residuals is similar among each of our three treatments. Okay, that just to help you connect these plots with, sorry, with the previous plot. Points that lie above this 0 line. These represent the data points that lay above the mean. Whereas negative residuals, residuals that fall beneath the 0 on the next slide, those refer to the data points that fall below the mean. Okay, so these residuals here corresponds to data points that fall below the mean value, in this case, or fertilizer B. Okay? So what we're looking for here is to see whether or not the variation within each of these groups is fairly similar. Okay? So what should the residuals look like? In general, we're looking for a lack of pattern in the residuals are random cloud of points. Or another way of saying that is to say that we want the variation within this group, this group, and this group to generally be similar. Okay? Now our last assumption refers to normality, where this refers to the assumption that the distribution of the residuals is going to be normally distributed. So if you focus on our data in treatment a, then we're saying that we want this, the distribution of these residual values. So a calculating each of these residuals. And if we were to plot those residual values, we want those residual values to be normally distributed, which is what this normal distribution refers to here. And we want that to be true for all of our data. Okay? Now, when we test our assumption, we don't actually look at the data within each of these treatment separately. Instead, we can pool all of these date and look at them simultaneously and what's called a QQ plot. Okay? And this is what a Q-Q plot looks like. And we test this assumption by seeing how well our data points fall along this diagonal line. And so the data will be perfectly normally distributed. If the data, all the points fall exactly on this line. Okay? You can see in this case, that's not true. And I'll say this example would be kind of at the limit of what we consider to be normally distributed data. Okay? So that's how we check all of our assumptions. There's one last essential point to make. And that is that it's really important to check your assumptions before you check your statistical results. And by that I mean, before you look at p-values, before you look at effect sizes, why is that? It's important to check your assumptions before you check your results. Because we wants to make decisions as to whether or not our data meet the assumptions. In a way where we're not going to be biased when we make those decisions. Imagine that you are performing an experiment that was really exciting to you. And you looked at your results before checking your assumptions. And you found that based on your p-values and your effect sizes, you got really excited because the results you thought were really interesting. And then you check your results, sorry, and then you check your assumptions. I mean, the fact that you got so excited about your results could influence your decision as to whether or not your data, Mickey assumptions. If your data were borderline, then having seen your results before, checking your assumptions, that could bias your decisions about, about whether or not your data actually meet the assumptions, might lead you towards saying You idiot data are okay. Whereas if you'd done the reverse process, if you check your assumptions before looking at your results, then you should be able to look at those assumptions more objectively. And you might be less likely to call results that are a little bit controversial. Consciousness is probably the wrong word. You might be less likely to say that results that are borderline in terms of whether or not to there they accepted, acceptably meet your assumptions, you're less likely to say that your results are OK, Hey, when they aren't actually necessarily okay. In other words, we want to make sure that we can check our assumptions without any preformed bias about how we want our assumptions too. About whether or not we want our assumptions for a particular analysis to be met. Okay, so that's why it's important to check your assumptions before you look at your results. There are no statistics police that are come knocking on your door if you do it the other way around. But it's a really good practice just to keep ourselves honest. And so I always try to check my assumptions before I look at my results. And I try to make sure that my data meet the assumptions before I check my results. And so then I say, okay, if my assumptions are met, then I can proceeds to interpreting my p-values and my effect sizes. We'll stop this video there. Say, I hope it's been helpful. And I'll say, thank you very much.