Okay, in this video, we're going to introduce a new type of model called mixed effects models. Mixed effects models is, it's a vast topic. There are entire books written on mixed effects models. And in large part that's because mixed effects models represent a very flexible tool that can be applied in a wide variety of problems. Due to that flexibility and the wide number of applications for mixed effects models. I'm going to be producing a, a large series of videos to explore mixed effects models in some depth, in some depth and some detail. However, I wants to point out then that the first three video is that you're going to see about mixed effects models. This one, the next one, and the third one, which will provide an example analysis. Those three, these three videos, I really meant as a brief introduction to the world of mixed effects models. And my goal is to make you familiar with the components of a fixed, of a mixed effects model. And to get a sense of the ways in which they can be useful to see how you can code and mixed effects model and how to interpret the output. We'll be building on that greatly in further videos. Our goal for this video is to answer these questions. What are fixed effects and what are random effects? As the name implies, a mixed effect model involves a mixture of different types of effects. And those different effects that are being mixed together are fixed effects and random effects. So defining it mixed fixed effects and random effects is really a, a logical starting point for introducing mixed effects models. And that's, that's what we're going to do. Let's start with fixed factors or fixed effects. Whether or not you realize it. You're actually already very familiar with fixed effects. And that's because all of the independent variables we have considered in our models up to this point. So where we've looked at one factor general linear models, multiple factor general linear models, general linear models that have covariates. In all those circumstances. We've only dealt with fixed effects. You just haven't realized that because I haven't giving you that term yet. So what I'd like to do at this point is just quickly revisit a number of those independent variables that we considered. And then ask what the general qualities are that they have in common. Okay. I'm gonna start just with calling something generally treatment. So if we had a factor in a model that we just called treatment and I've put quotation marks just to indicate them try to be very general here. Then usually we're talking about a situation where we might have a few specific levels that we want to compare. So for example, we might have study where we administer a drug, some subjects, that's, that's one level. A second level might involve subjects receiving nothing. Whereas a third-level might involve subjects receiving a placebo as a type of control treatment. So that would be an example of a treatment factor that we might model as a fixed effect. We've modeled data where we had either a gene present or gene absent. We've modeled data where we're interested in some aspect of biology as a function of different types of diet. We've compared aspects of biology between females and males. We've compared a number of species specifically to one another. We've examined aspects of physiology to either high altitude or low altitude. We've considered experiments where female or male just saw flow were manipulated. So they would either have high or low fecundity. So these are all examples of fixed effects that we've considered in our models. What do they have in common? Fixed effects are often described as having relatively few levels, like in this drug example for treatments, there were three levels with sex, there were two with altitude. They were too high and low altitude. Importantly, these levels are predefined. In other words, the researcher has decided ahead of time exactly which, which types of levels she or he will include in their experiment. And it follows from that that each of those levels is of direct interest. So back to our drug example where you had drug nothing being given or a placebo. An experimental design like this is very deliberate. And each of those treatments allows us to infer something specific about the biology and comparison among those various levels. Also allows us to learn specific lessons about the biology. So all of our levels are of direct interest for understanding of biology. So I've given you this general list and these general descriptions of fixed effects largely as a guide to help you decide in the future whether or not you might want to model a particular factor as a fixed effect or a random effect. Ok? But this view of fixed effects really doesn't tell us very much about what it means to be a fixed effect in the context of the actual analysis? In other words, how is a fixed effect different from a random effect? In terms of how it's treated by the analysis, in terms of the calculations. That's what I want to get into next. Jared Hadfield is at the University of Edinburgh and he's written, done some beautiful work. One of the things that he's done is produced this package called Mick, Mick Glynn for R, which is a very nice package that uses Bayesian approaches to analyze mixed effects models. And he provides a document which you can download with Mick, Mick limb called course notes. And these course notes provide guidance on how to use Mick, Mick limb. And at the very introduction, he describes fixed effects and random effects in a manner that I particularly like. And so I'm just pulling out some partial quotes from what he said to present to you here. Ok, so our goal here is to answer the question, how do we define a fixed effect with respect to how it is modeled? And what Jared points out is that with fixed effects, when we're trying to infer the value for some level of a factor, then we believe that the only information regarding the value of that level comes from the data associated with that particular level. What does that mean? Let's, let's walk through that with an example. Let's imagine that we wanted to study mallard ducks. And we want to compare some aspect of biology between females and males. And we wanted to analyze sex than as a fixed effect. If we were doing that. If we were to analyze sex as a fixed effect, then what that would mean is that the value for R trait that we are estimating for females will only use the information from the females. In other words, when we model sex as a fixed effect, we are implying that all of the information that's required to infer the value of our biological trait for females comes from females. And similarly. We're saying that all of the information that's required to infer the value of R trait for males will come from males. That might seem obvious. And if that seems obvious, the thing that I'm going to say next will seem even more obvious. But we'll make in a moment and we start looking at random effects. You'll start to see why these obvious things are so important. Okay? So the other way that I was going to express this is that when we model sex as a fixed effect, what we're saying is that when we determining the value of R trait for females, we only use information from females to determine that value. We are not using any information from males to determine the value of the trait for females. Similarly, if we're trying to infer the value of R trait for males, we only use information from males to infer that value. We are not using information from females. As i said, that might seem incredibly obvious, but will make its random effects are going to see why that's not quite as obvious as, as you might assume. Okay, so let's turn to random factors now. This is really the new concept for this video because as I said earlier, you're already familiar with random effects, whether or not, sorry, you're already familiar with fixed effects, whether you knew it or not. So to break the ice with random factors, I'd like to just start by considering some common examples of terms that are considered random effects in mixed effects models. So family, if you had a study that had that use multiple individuals within families and you had a number of families in your study, then you might include family as a random effect. Subject or individual. If you measured individuals more than once, then you would likely model individual as a random effect in your model. If you were working with mice or rats and you're using multiple individuals with that come from the same litter. And you had multiple litters in your experiment, then you might model litter as a random effect. Similarly, if you had a series of cages in your experiment, and within each cage you had multiple individuals. In other words, we have more than one data point for each of our multiple cages than we might model cage as random effect. Genotype is something we might model as a random effect. And I've noted here that whether or not we model genotype has a random effect would really depend on the studies goals. What I'm imagining here when I say we might model genotype as a random effect, is I'm imagining a study that might be conducted in the field of quantitative genetics. Where in that field, where very often interested in measuring some aspect of biology for multiple individuals that come from different genotypes. And then what we want to do is we want to ask how much variation there is among our genotypes. Because that can give us a sense of how much the traits that we're measuring is influenced by genetics. Ok, so in that context, we might consider genotype to be a random effect. And Biomedical Sciences we might consider genotype as a fixed effect. And one of imagining in that context is an experiment where we might have, say, a particular gene in mouse or a rat. And the normal form of that gene we'd called wild type. But then often the researchers are, are interested in understanding how a specific change to that gene influences biology. And so they will manipulate that gene and put that mutated gene into a number of other individuals. And then they specifically want to compare the biology of the individuals that have the wild-type form versus the biology the individuals that have the mutant form. And from that perspective, we would model genotype as a fixed effect. Ok. So again, how we model genotype, aren't we model as a random effect or fixed effect really depends on our studies goals. The same we true widths, region or country. If we wanted to specifically compare some aspect of society among say, two or three countries. And we're specifically interested in whether or not Country a is different from country B, then we'd probably model those data as a fixed effect. If instead, we had an example, or sorry, we had a situation where we had performed say, the same experiment multiple times within a country and repeated that over many countries. Then in that context, we might not be interested in comparing specifically among our countries who might just want to know whether or not country influences the outcome of the experiment. And so in that case, we might consider country as a random effect. So what are these various examples have in common? Well, we can often describe random effects as first of all, being effects that have many levels. So we would expect a study that has families to have many more families, then we would have levels within a treatment group of a fixed effect. Similarly, we might have many individuals within a study where each individual is measured multiple times. So random effects often have many levels. Really importantly, when you model something as a random effect, we are essentially saying we're modelling the data with the assumption that our different levels of our random effect will have been sampled from a larger population of possible groups. So again, back with our family example, if we had an experiment, say with 20 families, we would want those 20 families to have been sampled randomly from a larger population of families. And when we've done that, we can model family as a random effect. The last main point that I'm going to list here is that usually where when we model something as a random effect, we're not interested in the mean values of particular groups. So bags for our family example, if we had family, a family be through to family zed. We're usually not interested. When we model family as a random factor or as a random effect we're using are interested in comparing the mean of family a versus the mean of family b. So we're not so much interested in comparing the means of our different levels of our factor. Instead, what we're usually interested in is calculating the amount of variation in our data that can be explained by our random effect. So we might ask how much variation in, say, height can be attributed to variation among families as an example. Okay? So that's the kind of question that we might be particularly interested in and answering when we include something as, as a random effect, you know, we might be wanting to estimate variance for some particular reason that startup all the time in say, crop research. So when we're trying to breed new, new crops, I want to touch on another point here because I've started to open a can of worms about why we would actually want to measure something as a random effect. So one reason for why we want, might want to measure something as a random effect would be to measure how much variance they can account for in our data. We've already said that. Another good reason to model something as a random effect is that it can allow us, is that is because doing so can allow us to model independent data from an experiment in a way that gives us reliable results. So let's consider an example where we measured multiple individuals, each of them multiple times. And let's imagine those individuals were subjected to different types of treatments. If we were to analyze those kinds of data with say, a one factor general linear model. Those data on their own would violate the assumption of independence or they could violate the assumption of independence. And as a result, we wouldn't be able to trust the results from a one factor general linear model. When we have non-independent data like that, like multiple measurements from individuals, then we can use random effects in order to model the variation among those individuals. And doing so account, we account for the non-independence in our data. And that's a huge advantage, are huge. But that's one place where using mixed effects models can be incredibly useful. Because it can allow us to analyze data that are non-independent in a way where we can avoid problems with pseudo replication. That's a little bit of a tangent. So the main point that I was trying to get to here, just to bring myself back in rain myself in. Is that what we're trying to do here is we're trying to describe some very common aspects of random effects. So random effects usually have many levels. Those specific levels, like the various families we might include in a study, will have been randomly selected from a larger population. And usually we're not interested in the mean values of particular levels. Instead, we'll be modeling something as a random effect, either to estimate some aspect of variance that can be explained by that random effect or for some other practical reason like wanting to account for non-independence. Okay? So again, I've given this, this list of descriptions basically as a guide to, to help you to decide whether or not a variable might be modeled as a fixed effect versus a random effect. But again, this description doesn't really give us any insight into what it means to have, what it means to call something as a random effect with respect to the analysis. In other words, how do we define a random effect with respect to how it is modelled? We're going to return to Jared again for some advice on this. And I, he expresses the answer to this in this way. So he says, like a fixed effect when we're trying to measure the value for a particular level of a random effect. Or when we're trying to measure the value of a particular random effect. Then the information from that particular level will inform the value of that level. Okay, so that's just like a random effect, sorry, that's just like a fixed effect. With the fixed effects, we said that if you want to understand the biology of females, then they can use the information from females to understand that value. And we're saying the same thing is true here. So if we're trying to infer the value of some level of a random effect, then the information from that level is useful for inferring the value that level. Bot. We can take that information from that particular level and we can weight it by what other data tell us about what the likely values of our focal level would like, could possibly take. Okay. What does that mean? We're going to try to explain what that means with an example. Let's imagine we wanted to study kittens and we wanted to study body size of kittens. And to do so, we wanted to collect families of kittens for our study of body size. In our analysis, we might want to include family as a random effect. And if we're doing that, we want to make sure that we had randomly selected our families from a larger population of families. Okay? So here is our distribution that represents the effects of family on body size for this larger population. Ok? When we are randomly selecting our families for our study, what we are doing is we are randomly selecting. But I've just tried to illustrate that here. So we have 11 dots here, 11 points where each of these points represents the effect of a particular family. And so we have one effect for each of our various families. And these 11 families, or the information from these various families came from this larger population of families where we have this normally distributed. Sorry, I should have said that before. The original population, we're assuming that the original population from which our families was drawn will have, will be normally distributed. In other words, the distribution of the effects of families will be normally distributed within our source population. And so here are our various families are the values of our various families that we have selected from this larger population. I just want to bring in some terminology here. I want to point out, we would see that each family has its own effect. Okay, and as a result, it's better to say that the family effects are random. And that's because when we use that terminology, it emphasizes that it is the effects that are random. We have randomly selected effects associated with families that come from a larger, normally distributed population of effects due to family. It's less correct to say that family is a random effect. And it's less correct simply because this doesn't reflect the fact that it is the effects themselves that are random. Having said that, you will catch me using this term all the time. I've probably used it in this video already. And this is a very common way of phrasing it, even though it's not the most correct way of describing these things. Okay. So here's our where are we? We're generally to situation where we're saying we want to study body size for kittens were using families. We've said that there'll be a larger population of family effects on body size and those effects will be normally distributed. And we have randomly sampled families. Each family will have its own effect since we've randomly sampled these effects of family on body size from this larger population. There's some really important consequences of modelling random effects as being drawn from a population. Okay, there are two really important consequences of this perspective. This perspective of where we have, where we're saying that we're drawing random effects of family from a larger population of a family effects. Okay? And the first really important consequence is that now when we measure or when we infer the value of a particular family effect, we're no longer going to use only the information from that level, which is what we did with fixed effects. Instead, when we infer the value of, of an effect for a particular family, we're going to be use the information from all of the levels to determine the values for all of these effects. Okay, so now those remodelling the values of all of these effects jointly with one another. And I'll just point out to situations that I think will help you to understand why this is relevant. Okay? So let's imagine first that we have high variation among our family effects. Okay? If there's lots of variation among these family effects than most of the information that will be used to infer the level of a particular family will come from that particular family. Okay? And that's going to be especially true. We have lots of replication within each of these families. This is something that probably should have said earlier. When we're sampling families, what am I imagining is we're sampling multiple individuals within each family. So if we have 11 families and we pulled out ten individuals for each family, then that would mean we have a 110 individuals here. Okay. Sorry for backing up there for just a moment to the main point that I want to get at here is pertains to this. Which is that when we model our data in this way, we're using information from all of our various levels to infer the value of the levels at each of our, uh, to infer the values at each of our levels. And we've just walked through a situation where the variation is high among, are among our levels. If the variation is high among the levels. And especially if we have lots of replication within our levels than most of the information used to calculate the value of that level will come from that level itself. And that'll be true for all levels. Let's imagine instead that we have relatively low variation among our family effects. Okay? And let's also imagine that we have relatively low replication within each of these families. So we only have, say, two or three individuals per family as opposed to 20. Okay? In this context, then it's going to be more likely that extreme values and our distribution of these effects are going to be due to sampling error. And in that case, what the analysis will do is it will use the information from all of these levels together to shrink the more extreme values in towards the middle. It's not just going to shrink the extreme values. This shrinkage will happen to some degree for all of the different levels. Okay, so I don't mean to paint a picture where the analysis will pick, will say one particular value like we have here and only shrink it. This shrinkage will occur for all of the levels, okay? And that is that shrinkage is going to happen more when there's low variation or a little variation among the levels of our random effect or a random factor. And when we have little replication within these. So these two situations of high variation among the effects versus low variation among the effects present a kind of continuum to understand how these models can use information from all the levels that are used to determine the values at these levels. So we have lots of variation and lots of replication. Then the values for individual levels, we'll mostly use information from each of their respective levels. If we have low variation and low replication, then when the analysis is determining the values of each of our levels, it will be shrinking somebody's values, using or shrinking the values and bringing them more towards the center. Using information from all of these points, from all these families collectively. So with that discussion, I hope that this point is little bit more clear. So I'll just repeat this point that you made earlier. What's special about random factors is that when they're being modeled, when the values of our levels are being modeled. Then we say that as we have in a fixed effect, we can use the information from a particular level to inform the value of that level. But we can also make use of the information from the other levels for random effect in order to inform the value of not just our focal level, but for all of the levels. So that's, that's the first consequence of modelling data in this way, modelling data as a random effect being drawn from, or say it's comes from a larger normally distributed population. The second consequence is that when we model the data in this way, what we're actually doing is we're modelling the original population from which our families came, or whatever our random effect is. Okay? And as a result, when we model our data in this way, we are able to make general statements about the original population, which is usually what we're trying to do when we conduct a study. Usually when we conduct a study, we want to be able to make general statements about the world, or at least about a larger population that interests us. Modeling the data in this way through random factors allows us to do that. And I'm going to go into that a little bit more in the next videos. That's going to be the subject of the next video you're going to see. Let's wrap up our discussion of fixed and random factors. We'll just quickly recap what we've talked about. We've said that fixed effects for fixed effects. When we model something as a fixed effect, we're essentially saying that we believe that the only information that's needed to determine the value of a particular level of a fixed effect comes from that particular level. And with fixed effects, they typically have few levels that are pre-defined in there of direct interest. And we're already familiar with fixed effects, whether we know it or whether we realize it or not. With random effects, we say that when we're trying to estimate their values, then what we're saying is that we believe that the information for, for determining the value of a particular level of a random factor. That information from that level is relevant to determining the value of that level. But we can use additional information to infer the vout. To infer the value of that level we can use. We can weight the information from a focal level by other data for other levels to tell us about what the likely values of our focal level could be. And again, I'm not trying I shouldn't I'm not trying to say that that that this that this analysis focuses on one factor at a time because that's not how the analysis works. The analysis will be going through this process simultaneously for all the different levels of our random factors. Random factors typically have more levels than fixed effects will. With random factors were saying that we will have drawn the various levels of our random factors from a population. And we're typically not interested in the mean values of random factors. Typically we model random factors for other reasons. Will stop our discussion of random if fixed effects there. I hope this video has been useful and I'll say, thank you very much.