Okay, In this video, we're going to continue our discussion of experimental design. And we're going to focus on a topic called covariance. I'd like to first remind you of this really nice book, Experimental Design for the life sciences by Ruxton and coal grave. They talk about covariates in a really nice fashion. So if you'd like to read more about this topic, then this is one of the places where you could go to get additional insight. Before we get into covariates, I wants to quickly revisit an example that we've discussed in some of our previous videos. So for example, when we talked about blocking, where we imagined an experiment where we wanted to test whether or not petting puppies would change the average rate of tail wagging. And we imagined this experiment would have two treatments and no padding treatment that I'm heading treatment. And we said that these would be some of the subjects involved in our experiment. And we importantly realized that our subjects would likely differ in a number of characteristics. So in their age or their size, their breed. And in the context of the blocking video, we discussed how variation in these traits could influence our ability to understand the differences between these groups if these traits are associated with our Y variable. So our tail wagging. I'm just going to quickly, at this point in the video, I'm just going to quickly revisit those ideas. Because those ideas are central to our discussion of covariance that's going to be coming up. Okay. So let's use an example. Let's imagine, let's focus on size, body size. And let's imagine what we're gonna do is we're going to think about how variation in body size could influence our ability to understand differences between these groups. So let's imagine that body size is related to the rate of tail wagging in some way. So for example, let's imagine that larger puppies tended to wag their tails more than smaller puppies. Okay, if that's true, and let's imagine that our experiment involves both small and large puppies in our two groups. And we include a variety of size of dogs because we want our experiments to have fairly general conclusions that would apply to a wide group of dogs. Then that variation in the size of dogs within groups, we'll increase the amount of variation that we have in tail wagging compared to a situation where if we're only studying, say, a single size of dogs. And that's because if you are only considering a single size of dogs, then there'll be very little variation in size. As in size influences tail wagging. There'd be very little scope for size to influence tail wagging very much in these data. If in contrast we had large puppies and small puppies. And if that difference in size can leads to substantial differences and tail wagging. Then the variation in size within these groups can increase the amount of variation in tail wagging that we have within this group and also within this group. Now, as we talked about in our blocking video, having more variation within groups can affect how we draw conclusions for our study. Specifically, if we have more variation within our groups, then that can make it more difficult to detect evidence for differences between our groups. Alternatively, if you're analyzing our data in such a way where we wanted to estimate the size, the difference between these groups. Having more variation within these groups would mean that our estimates, the differences between these groups, would likely be less precise and less precision is undesirable. Okay? So in order to remedy this situation, in other words, in order to control for the effects of things like size on our attempts to understand the differences between these groups. In this previous video, we suggested that we could use blocking to those ends. So if we blocked by, by body size, so let's say we had small dogs here, medium dogs here, and large dogs here. Then by including body size as part of our expense experimental design that allows us to more effectively detect differences between these groups and to be able to estimate the differences between these groups more precisely. As we talked about in our previous video on blocking. In this video, I'd like to discuss an alternative approach for controlling for variation among our subjects. So I want to provide an alternative approach that still meets many of the same goals as blocking, but that does not actually require blocking. This alternative approach involves using something called covariates, which I'll explain in a moment. First, I'd like to point out though, that this alternative approach can only work for certain types of variation among our subjects. Specifically, we can use this alternative approach of covariates. If the variation among our subjects is continuous variation. If there's a natural order to that variation, and if the relationship between that variable and the variable that we're measuring. So like the rate of tail wagging, if there's a linear relationship between those two things. Okay? So if those three things hold, then we can use this alternate approach of covariates to allow us to better understand the differences between our two different treatments. I know all the sounds really vague right now because I haven't even told you what using covariance means yet. But we're getting there, okay. My description here of these three different. Criteria that need to be met in order to use this co-varied approach might seem a little bit vague. So let's make this more concrete. Okay? So we mentioned earlier that these are some of the ways in which our subjects could differ. And I'd just like to briefly consider whether or not all of these different ways in which our subjects can differ can all be analyzed using this covariate approach. Okay? First of all, let's consider breed as an example. Can we use breed or can we analyze the effect of breed using this co-varied approach that I still haven't explained to you yet? And the answer is no. And that's because breeds are described in a way that is not a continuous trait. Breeds are usually described in terms of categories. You are a Labrador retriever or you are poodle over your a labradoodle, which is a combination between a Labrador and a poodle. Okay? So breed is not continuously distributed. So we could not use covariates in order to control for variation among breeds. What else? Something like size though, or age. Both size and age are continuously distributed. They both have a natural order to them. In other words, if we were to measure size in terms of grams, we know that a large number of grams, or if we know that measurements that have a larger number represent a heavier organism than those that are measured with, that yield a smaller number. Okay, so there's a natural order to age measurements and to size measurements, okay? So both of these variables meet these first two criteria. Whether or not age or size has a linear relationship with our Y variable. So in this case, the rate of tail wagging, that's something that we'd have to determined by just plotting our data and seeing whether it looks like there's a linear relationship between size and the rate of tail wagging or age and the rate of tail wagging. And if it does, then we could include size or we could use size or age. Or I could say, I want to say we could use size or age in the approach and we're about to describe. Okay. So just to summarize this, the approach that I'm going to describe where we use covariates is not applicable to all types of variation among our groups, are among our subjects. But for certain types of variation that can be described according to these ways, we can use this approach which involves covariance. Now let's get a little bit more specific about what we mean about using covariates to help us understand differences between levels of a factor. We use covariates. What we're specifically doing is we are measuring each of our subject for some particular trait, like age or size. And then we would include that information as a variable in our analysis. In other words, we might include size within a model that we're using to understand the effect of our padding or not patting treatment. Okay? And when we do this, well, we include these variables within a model or within our analysis. We now refer to these variables as being covariates. Okay, So while we include these covariates and a modeler or inner analysis that can help us to understand the biology that we're specifically interested in. So in this case, our padding or not petting treatment. Okay? Now I'd like to point out that covariates are these, these values themselves might or might not be direct biological interest. So for example, if we already know what they were interested in whether or not petting influences tail wagging. We might or might not be interested in whether or not size influences the rate of tail wagging. Whether or not we're specifically interested in whether or not either of these covariates influences rate of tail wagging is immaterial. Okay? If we are interested in how they influence tail wagging and that's great. Including them in our analysis might give us some additional biological insight. If we're not interested in them specifically, then they can still be helpful by helping us to get a better understanding of these other factors that are of focal interest. Okay? What we're going to do now is we're going to delve into how including these covariates and analysis can help us to understand differences between levels of a factor. That's we're going to spend the rest of this video focusing on. And to do that, we're going to use this cartoon illustration of some data. And what we have here is a y variable, which had just called response. Something along this x axis, which has to just labeled x. And I've color coded two sets of data to be either red or blue, which corresponds to two different levels of a factor. We're going to describe this scenario using our puppy experiment. But it just wants to give you some other examples as well for where this might apply. Just to get your mind, just to get your mind going and to help you. Extrapolate from this specific cartoon to maybe some data that you might have already that might be applicable to the situation. Okay, So I'm just going to describe a num, a number of different types of experiments that might fit to this cartoon. Just as a way to help you see how your own data might apply to this particular scenario. So, what are some examples? And biomedical science? Let's imagine that our response variable was the degree of gene expression. Okay? And let's imagine we wanted to compare gene expression between two different genotypes. Let's imagine maybe blue is wild type and read it was some mutant genotype. Okay? So ultimately we are interested in understanding whether or not there's a difference in gene expression between our mutant and sorry, mutant and our wild-type genotypes. However, we might recognize that some other physiological factor might also influence gene expression. So let's imagine something like that degree of stress, which we might quantify as the level of a stress hormone that's floating around in our subject's blood. Okay, So if it turned out that the level of a stress hormone influenced the gene expression, then we might want to include, That's the level of our stress hormone in our experiment in order to help us understand the difference between a gene expression for our mutant and our wild-type genotypes. Alternatively, let's imagine we're working with plants and let's match you want to understand things that influence the number of flowers and plants produce. So you might have flower number on the y-axis. Our experiment might involve a fertilizer treatments where we might have a fertilizer given to the blue subjects and no fertilizer given to the red subjects. And we might recognize that the size of plants might influence a number of flowers they produce. So we might have size along our x-axis. Okay, so those are just two examples to get your mind going on, on this topic. In the context of our puppy experiment, we can say that our y variable would be the rate of tail wagging. Our x variable might be size of the puppies. And the red points might corresponds to the Nope heading treatment. And the blue points might represent data from the padding treatment. So just looking at this, what we might initially guess is that larger dogs tend to wag their tails more. Add, on average, it's possible that the padding treatment, it leads to a higher rate of tail wagging than the non-paying treatment does. And that's, I just say that because these blue points tends to be above the red points. Whether or not we have strong evidence for a difference between these groups. That's something we would conclude based on a formal analysis here. Okay. So now that we're kind of oriented to the kind of biological scenario that we're faced with. What I'd like to do now is explain how including this covariate in our analysis will help us to understand the difference between the blue group and the red group. And in order to do that, I want you to first see the amount of variation that we would have within the blue group and the red group. We were not considering the covariate at all. And we can get that simply by taking each of these data points and just slide them over to the side. Whoops, like we're going to see on the next slide. Actually, I forgot to mention this, this point here. So we're just going to mention this, this point here. Before we go on to visualizing the data, I just want to point out that we should really, we can only really use this approach when we have substantial overlap along the x-axis for our two different groups. Can see here that along the x axis, all of the data from the blue group largely overlap with the x axis for our data from the red group. And that's excellent. That's the kind of thing that we want in order to perform this kind of analysis that does not have to be complete overlap between the groups. So for example, these blue data could be shifted to the right somewhat so that some of the range of the blue data did not overlap with the range of the red data. That's okay so long as there's sufficient overlap along the x axis between our two groups. I apologize for not mentioning that earlier, that that's an important additional criteria for using this approach. Now that I've said that, let's go back to what we're what I was about to do, which is explain why including the covariate in our analysis helps us to understand the difference between these groups. And I said that we can start that understanding by just recognizing what the data would look like if we did not account for the covariate. Okay? This is what our data would look like if we just produced a scatter plot of these data without considering the covariate. Okay? And all I've done here is I've taken each of these data points and just slid them over here. So the, this data point has been slid from here to there. This data point has been slid from here to there and so on. Okay? So if we just produce a scatterplot of these data, are data would look like this. And you should see that there's a fair amount of variation within each of our groups. There's a lot of variation within the blue group and is a lot of variation within the red group as well. So the highest value for red is, their lowest value for red is there. And when we have a lot of variation within each of our groups, we said earlier that and that's going to make it more difficult for us to understand the differences between our groups. So how does including a covariate help us? That's a number different ways of illustrating this. I want to illustrate it in this way by having you imagine that we can slide our data points along each of the lines that correspond to their group. Okay? So essentially when we perform an analysis that includes both a factor and a covariate. We can understand how the model, I'm sorry. Let's be careful how I say this. If our lines are parallel as they are in this case, then we can understand how our analysis compares the mean of the blue group to the mean of the red group. By imagining that we're just simply sliding the data points along their respective lines. So here we have a red point, and so we're just going to slide it along the red line up to this point along the x-axis. Okay? And we slid it along so that it's this arrow is parallel, this line here. And by doing that, we're making sure the distance between this point and this line is the same at this point along the x axis as it is at this point is along the x-axis. And we've done the same thing here for this blue point. We've just split it along this blue line however. Okay? So we've moved that point there up to there where we now have this open circle. Okay? Now if we do that for all of our data points, we move them all to this position along the x-axis. Then what we can essentially do is we can see. This allows us to appreciate how our model allows us to compare the blue data to the read data in a way that accounts for the variation in the, that accounts for the variation in our data due to the covariant. Okay? So what I want you to notice here is that when we include the covariate in our analysis, we can essentially imagine that we're comparing all, we're comparing the data from our blue group to the data for our red group. All for one particular level of the of our covariate. Okay, I don't want to be careful here and say that we can really only adopt this perspective if our lines are parallel, okay, and we're going to talk about why we can't have this perspective. Lines aren't parallel shortly. Okay? So for the moment, I just want you to imagine this case for the lines are parallel. And in this case, we can understand how including the covariate and our analysis allows us to essentially. Slide all the points along to one position along the x-axis. So that now we are comparing the mean value for the blue against the mean value there of our red. For a single value of size or for a single value of r covariant. Okay, that's, that's one way in which we can understand this approach. And what I want you to notice here is we have a lot less variation in these data within our groups than we do over here. Now I recognize that it's a little bit hard to see that that here because everything so scrunched together. So what I'm gonna do is I'm just going to match these open circles and slide them along to the right. And that's what we have here. I want to point out that all I've done is taken these exact points here and slid them over. I'm not meaning to say that we're now examine these points for different level along the x axis. Okay? So when we're looking at this, at these points here, I want you to forget, where are these data lie along the x-axis. Really what I want you to recognize or what I want you to focus on is the amount of variation that we have within these groups now compared to what we started with. And what you should see is we have a lot less variation within these groups come when we account for the covariate. Then we saw when we did not account for the covariate, and that's especially true for the red group. The red group. Now we only have this amount of variation within the red group. Whereas previously we had a lot more variation within the red group. What this will do is this will help us to better compare the mean values for our red group compared to our, our red group compared to our blue group. Okay? So that is how including the covariate in our analysis can help us compare between our groups. Essentially by including the covariate in our model. We can essentially statistically control for the variation in the response that's due to the variation in our covariate. And as a result, we can compare the mean values from our blue with the mean value of our red essentially for a single value of our covariate. Okay, and again, that perspective holds when we, when we can say that our lines are parallel. Okay? In a moment, we'll talk about the case where lines are not parallel. What I'd like to point out next is that when we analyze our data in this way, we are asking a slightly by a slightly different, by a logical question. Then when we do not include the covariate. So when we do not include the co-vary, what we're simply asking when we compare the mean of the blue to the mean of the red is, we're simply asking, is the mean of the blue group different from the mean of the red group? Okay? Well may include a covariate in our analysis, we're asking a slightly different biological question, which is. Is there a difference between the blue and red group having accounted for the variation due to our covariate. Okay. So it's a slightly different biological question. But that won't necessarily get in our way of doing interesting biology. In fact, the point of this video is by including covariates that can actually help us to understand interesting biology by controlling for variation in our data that makes it easier to detect differences between our groups. Now it might turn out that our covariate might also be a biological interest. And if that were the case, then including the covariate in our analysis also allows us to answer this biologically interesting question, which is, does the covariate affect our response? And if so, how? So does the response increase the cope with the covariate or decrease with a covariate, for example. Okay. Now, I've, I've made a big deal about the fact that we can use this interpretation that I've been highlighting it when our lines are parallel. And that's because when our lines are parallel, this basically means because the definition of parallel, this means that we essentially have 11 effect of our factor. What do I mean by that? Well, what I mean is if we were to measure the difference or the distance between the blue line and the red line for this level of the covariates. So at this point along the x-axis, we would get an effects that's this big. Okay? So that's the difference between the blue line, the red line that gives us the size of our effect. When these lines are parallel, the size of the effect here will be the same over here. And that's just really follows from our definition of parallel, okay? And so because this, the size, the difference between the blue and the red is the same up here as it is down here. And that really highlights the fact that we really have one single effect of, of our factor of the difference between the blue and the red regardless of the value of the covariate. Okay? So in this case, our interpretation of our results is really straightforward. We can say that really there is one average effect of the, of the blue versus the red, which we can understand more easily by taking into account the effect of error covariance. The interpretation is not so straightforward. However, if we have an, if we have lines that are not parallel. So if we have lines that are not parallel, specifically, if we have statistical evidence that the lines are not parallel, then we call this situation an interaction where we say that there's an interaction between our factor. So between the blue and the red and our covariate. And in this case, our interpretation needs to be more nuanced. It's a little bit more complicated. But complicated does not mean necessarily difficult. Okay, we might, instead of saying complicated, we might just say interesting. So I'm going to use this figure now to, to illustrate why our interpretation has to change. Okay? And I'm gonna do that by ill, by having you look at the size, the difference between the red and the blue lines for different values along the x-axis. Okay? So in this scenario where the lines are not parallel, we could move our points along their respective lines, say to this point along the x-axis. Okay? And if you were to do that, then what we'd essentially we'd be doing is we'll be asking, what's the average difference? Sorry, what's the average difference between the blue group and the red group? For this level, the covariate. And we would see that we might have evidence that the blue group is higher than the red group for this level, the covariant. What happens if we consider other levels of the covariate? So for example, what would happen if we consider the covariate at this level? Well, when x has this value, you can see that the red and the blue lines actually cross one another. Which means that we expect there to be no difference between the red and the blue for this value of the covariate. Alternatively, if he considered the difference between the red and the blue at this value for the co-varied, then you could see that we would expect there to be an even larger difference between the red and the blue here. Then we found in our first case. Finally, let's go way over to the left. And I want you to notice that at this point along the x-axis, we can see that might be a difference between the red and the blue. But in this case, the red has a higher average than the blue, which is the opposite of what we find at this end of the x-axis. Okay? So when we have non parallel slopes are when we have statistical evidence for there being a difference and between the slopes for our two different levels of our factor. So between the blue and the red, then in this case, it makes no sense to try to compare the red versus the blue. Because not one single difference between the red and the blue. When the slopes are not parallel, we can have many different differences between the red and the blue depending on which level of the covariate we're considering. Okay. So that's just what I've highlighted here. I've said it makes no sense to compare the levels within a factor. Sorry, I should say I should say between factors if the slopes differ. Okay, So it makes no sense to compare between the red and the blue. If the slopes, the red and blue are different. Because in this case, any difference between the red and the blue will depend on a particular value, the covariate that we're considering. Okay. That's really important. Does that mean all hope is lost? I would say absolutely not. If this kind of situation arose in an experiment where you include a covariate and a factor. What this tells you is it might just actually reveal some additional interesting biology. Because it would tell you that the effect or the size, the difference between our different levels depends on this covariate that can be biologically interesting. Conversely, it could tell us that the effect covariate, in other words, the slope for the relationship between the covariate and the response depends on which level the fact you're looking at. Okay? And that again, could be biologically interesting. Okay. I'm in the context of our puppy experiment. If the red corresponded to the no padding treatment and the blue corresponded to the petting treatment. What these results would tell us is that size or an increase in size will leads to an increase in the rate of tail wagging faster when we are putting the dogs compared to the case when we don't pet the dogs. That's how we would interpret these results. If x was dog size and the red and blue represent their chief two different padding treatments. And that could be biologically interesting. I've summarize these different types of questions on this slide here, where I've just noted that if we include, if we allow our analysis to include the potential for different slopes or to have an interaction between a covariant HER factor. Then we can ask these biological questions. We can ask. Does the effect of a factor depend on the value of the covariate? In other words, is there an interaction? Or equivalently, we can ask, does the effect of the, does the effect of the covariates? In other words, the slope of the covariate depend on which level the factor we consider. These two questions are just two sides of the same coin. They're both essentially asking, is there an interaction between our co-vary factor? If we find evidence for an interaction, then when we're trying to understand the biology of our system, we should focus on understanding how these effects have arisen. Okay, so describing the kinds of results that we see here, describing how this slope is different from that slope and what that would mean biologically. If it turns out that we don't have any evidence for an interaction, then we could say that we have evidence to say that our lines are parallel. And in that case, we can interpret the results of our experiment in light of our original case, we are considering where we had parallel lines. And in that case we said that we could simply, I'm interpret our results in light of this question, which is, do the levels of a factor differ from one another after accounting for variation due to the covariate. Okay? So what I'd like to point out here is that covariates, as well as allowing us to detect differences between our factors more easily. They can also open opportunities to reveal new interesting aspects of biology. Okay? And so I'm just going to close with this quick summary when we include covariates in our study and in our analyses. Then this can help us do two things. First of all, can help us to detect and estimate the effects of a factor that specifically interests us. But it also might open opportunities to ask additional biological questions. There'll be more difficult to address. If our experiments only included factors and did not include covariates. I'm going to end the video there and say, I hope it's been helpful. And I'll say, thank you very much.