Okay, in our previous video, we started our discussion on how Analysis of Variance works. And we talked about how analysis of variance works by taking a total amount of variation and dividing it into two sorts. Or the variation within groups and variation among groups and then comparing them. And we talked about how to calculate that variation using something called sum of squares. We ended our last video by saying that in order to actually perform an analysis of variance, instead of analysis of variation, we needed to take a next step or we needed to convert our estimates of variation, which were made through our calculation of sum of squares into measures of variance swing to go from variation to variance. And that's what we're gonna do in this video. We're going to talk about how we can make that step from variation measured a sum of squares to variance. And to do that, we're going to introduce a new concept called degrees of freedom. Where degrees of freedom refers to the number of independent pieces of information that are included in an analysis. In other words, the number of independent data points that are free to vary. It might be a good first instinct to assume what the number of independent pieces of information and analysis will simply be the number of data points you had. Especially if you collected your data points to make sure that they were statistically independent, That would be a really good first instinct. What I want to talk about in this slide is why that instinct is not quite correct. And the reason why it's not quite correct is that when we go through the calculations that are involved in an analysis, some of those calculations leads to a loss of independence or a loss of the number of data points that are free to vary. In particular, when we calculate sum of squares, we calculate a mean. And on this slide we're gonna talk about how calculating this mean has certain consequences. Before we do that, I just wants to take a slight tangent and explain this notation here. And just because I don't want there to be anything on the slide that's a mystery to anyone. Just want to point out that this notation or this equation refers to exactly what we did previously to calculate sums of squares. If you remember to calculate sums of squares, what we did is we took each of our data points, subtracted that from a mean or found the difference between each data point and the mean, found that difference squared the difference, and then added them up. That's exactly what's meant by this notation here. What we're saying here is for our AI data points, we take each of our individual data points, subtract from that a mean value, which is given by this x with a bar over it. So that's our difference. Then squaring that value. And we repeat that for all of our data points. And then this symbol here just tells us that we're going to take all of those results and add them up. Okay? So this just refers to exactly what we did previously when calculating sums of squares. Okay? No, that's enough of that tangent. Back to our task at hand. We were talking about why calculating a mean and the course of our analysis. So in the course of calculating sums of squares removes some freedom before our data points to basically be free to vary. So why does Cathy mean have that effect? Well, let's go back to basics. We calculate a mean value by calculating the total of all the numbers that we're interested in are it's our total for all of our various observations. And then we just divide that by the total number of observations. Okay? Let's say that this total is a fixed value. Ok, so for sake of example, let's imagine that we have three observations which will call x1, x2, and x3. And let's imagine this total is equal to 100. What I'd like you to realize is that when we go through the course of this calculation, not all of our data are free to vary given that we're going to have this particular total. So for example, if we're totals a 100 and we say the first two data points, they can be whatever they quote unquote wants their free to vary. Let's say the first one is ten, second one is 80. They were free to be anything. This could have been 10 million, not just ten, but whatever they are. This last value is constrained to make this total equal to a 100. So in this example, this last value of X3, it must be equal to ten. So this is the last value is no longer free to vary. Only the first two are free to vary. Likewise, if the first two values had been 1515, then this last value must be 70. So if we have three data points, only two of them are free to vary. Okay? So that's what's referred to as our number of degrees of freedom. The degrees of freedom when we're calculating the sums of squares or a number of pieces of information for calculating our sums of squares is equal to r number of data points for that type of sum of square minus one. As we saw here, we had three different data points. But our degrees of freedom here is actually equal to two, because only two of our data points were free to vary. So how does this carry forward for Our analysis of the sums of squares within and between treatments. Well, with this information at hand, with our understanding of degrees of freedom, we can now calculate the degrees of freedom for both are within treatment sums of squares and are between treatment sums of squares. Let's start with within treatment sums of squares. Our degrees of freedom for this portion of our data. So for treatment a is going to equal to the total number of data points that we have, which we'll just call an a. So it's the number of data points we have treatment a minus one. So that's the degrees of freedom for this first part. To calculate the sums of squares for this first treatment. Same general idea is true for the next treatments. So for achievement be our degrees of freedom for this treatment is equal to r number of data points that are in treatment b minus one. And similarly for treatment C, our number of degrees of freedom he here is equal to the number of data points and treatment C minus one. So if we want to determine the overall degrees of freedom for the within treatment sums of squares, we could just add these up, which is a number of data points in a minus one data points and b minus one plus NMR data points in c minus one. And of course, the number of data points in a, b, and c, Those just add up to the total number of data points in your experiment. So if you want to calculate the, or the total, if want to calculate the degrees of freedom. For our within treatment sums of squares, what we do is we just take our total, our total sample size and subtract from that the number of treatments that we have. And that will give us our degrees of freedom for the within treatment sums of squares. And that's these arrows just say we get this minus one, minus one, minus one, minus one for each treatment. And so that's why we get minus three. Or more generally can say just subtracting the number of treatments you have in your experiment. Okay? Our degrees of freedom for the among treat, for the among treatment sums of squares is very similar. In this case, we have 123 values that we use to calculate hour among groups, sums of squares. And so our degrees of freedom for this term for the among chickens, sums of squares is just equal to our number treatments, which in this case is three minus one. So our degrees of freedom here is going to be equal to two. So this is where we can make our transition from variation to variance. And we're going to do that by calculating something called the mean square. And this mean square is made up of all the things we've seen so far. On the numerator, we have this term which we just described about two slides ago, where this represents. Our sum of squares, okay? And we just divide our sum of squares for whatever type of sum of squares it is that we're interested in by the appropriate degrees of freedom. So we, if the sum of squares is for the within group sum of squares, then we would take that sum of squares and divide it by the within group sum of squares. I did say that this was the within group, right? We want within group divided by within group. So sums of squares, degrees of freedom. That's what we're doing. Okay? So we take the within groups sums of squares, divide that by the corresponding within group degrees of freedom. And that gives us something called the mean square. That's this. Ms stands for mean square. And the mean square is just a measure of the amount of variation that we have for part of our data, which is corrected for the number of data points. And we're going to calculate two types of mean squared depended on whether or not we're looking at the within group or between-group variation. So we'll get one that mean squared for the between treatments variation and one for the within treatment variation. I'd just like you to make this realizes comparison. Ok. Let's compare what we've just calculated as mean square. We just talked about what that is. And I'll just remind you of our equation for calculating the variance when we're estimating variance of the population. That's who we have here. How are they related? To stop and have a think about what they're actually identical. So this Mean Squared is actually a measure of variance. So by calculating the degrees of freedom, we can, we divide our sums of squares by that degrees of freedom in order to get a measure of variance. And so by doing this, we can obtain two measures of variance. Variance between the treatments and variance with in the treatments. And what we're going to do next is we're going to take those two different measures of variance. So this mean square, which we have on top of this fraction. So we have MS one divided by MS to MS one refers to the between treatment variance. And remember I said before that this term will also include within treatment variance, okay? But the main type of variance are most interested in here is the variance that's due to variation. That's due to the treatment effect, okay, which is going to be included in this term. And we're gonna divide that by this mean squared for the within treatment variance. So broadly speaking, we're taking the between treatment variance and dividing that by that within treatment variance. And this ratio has a special name. It's called EF, where f is the test statistic that we calculate for analysis of variance. And what we can do is we can take this value of F and compare it to a null distribution in order to calculate a p value. And that's what we're going to talk about in the next video. So I hope this video has been helpful. Just to sum up, we've talked about how you can use degrees of freedom in order to convert sums of squares into an estimate of variance, which we call Mean Squared. We either have a mean square for r between treatment variance or mean squared for within treatment variance can use those values or ratio of those values to calculate F, which is our test statistic for analysis of variance. We'll pick up on that in the next video. Hope this has been helpful and say, thank you very much.