Okay, In our previous video, we introduced the concept of independence. In this video, we're going to build on the concept of independence by considering independence specifically in the context of experiments. And we're going, in this context, we're going to understand how non-independence can lead to something called pseudo replication if we do not analyze our data correctly. And then the last thing we're gonna do in this video is we're going to understand why incorrect analyses when you have non-independent data can lead you to false conclusions. Okay? Those are general goals for this video. This video has a particular structure where what I'm gonna do first is I'm going to introduce a series of core ideas. That's why have core ideas mentioned here on this slide. So we're going to introduce a number of core ideas that we're then going to put into use when we consider three separate scenarios are three different experimental designs later on in the video. So to recap, what we're going to do is I'm going to introduce a series of core ideas. I'm going to summarize them and then we're going to put those ideas in practice when we consider three different experimental designs. So let's get started. Throughout this video, we're going to consider a series of experimental designs that look like this. So this is just kind of a cartoon sketch of a, of a particular experimental design where we're imagining we have an experiments that considers two treatments. We have treatment 1 here and here we have seven data points, a, B, C, D, E, F, G, Cs are seven data points that were collected for treatment 1. And we have another seven data points collected for treatment to as a point of vocabulary. We might call this an experiment with two treatments. Or we might call this a one-factor experiment with two levels. Where one level is called Schmitt one and another level is called treatment to. What I really like you to focus on in this slide is just kinda the nomenclature for these data points. Because I've noted down here that measurements are not independent if they share a letter. Okay? You can see here that all of the data points here have unique letters. And this tells us that we have no reason to worry about non-independence anywhere in this experiments. In other words, this, as far as we know, all of the data within treatment one should be independent. All the data in treatment 2 should be independent. And we have independence between our treatments. That's not going to be true. And some of the scenarios we consider later on, where later on we're going to find some cases where we might have a listed three times in a row. And that would indicate that those three A's are three data points that all have something in common that would lead us to worry about non-independence. Okay? So that's, that's basically how we interpret these cartoon sketches of our experimental designs. What I'd like you to appreciate next is that in this first scenario that we're considering, and in all the snares that are going to come up. We're always going to imagine a situation where the null hypothesis is true. In other words, we can imagine that a scientist might have set up an experiment like this because they were interested in testing whether or not the mean value for treatment one is different from the mean value treatment 2. Okay, That was their motivation for conducting this experiment. But the truth is, for this experiment, the end for all the ones that follow. The truth is that there actually is no difference between the true mean of treatment 1 and the true mean of treatment too. Okay? Having said that, even though the true mean for treatment 1 and the true mean for treatment two are the same. We still don't expect that the mean values that we obtained for samples for treatment one to be the same as a mean value for a sample of treatment two. And that's because of something called sampling error. We've mentioned sampling error in some previous videos. So for example, if you've seen the videos on, for the forearm experiment, where we measure the forum lengths of females and for males. And we wanted to determine whether or not there was evidence that The average forum lengths of females different from that for males. In that context, we introduced this concept of sampling error, where we said that if we had one population, which we represent by a normal distribution. So my hands here trying to take the shape of a normal distribution. We could take a sample from that single population fro. So from that single normal distribution, say we pull out seven data points and then we take, we determine the mean value for the seven data points that some holding that mean value here, okay? We could repeat that process by pulling out another seven data points from the same population or the same normal distribution to obtain another mean value for those other seven data points. And we would not expect those mean values to be identical even though they came from the same distribution. And that's because of the random processes that are associated with obtaining a random sample from a larger population. And the random processes that lead to these different samples enter these different means. We refer to that as sampling error. Okay? What I'm going to do now is I'm going to expand on our understanding of sampling error just a little bit by running some simulations. There are other videos available where I discuss something called standard error, where I perform exactly the same kinds of simulations as these, but we discuss them in greater depth. So I'm just mentioning that now because if you've already seen those other videos than what I'm about to show you will be familiar. If you've not seen those other videos, then I'll just say that what I'm about to explain in this video is going to be a relatively quick treatment of these ideas. In particular, when I explain the R code here, I'm not going to explain it in the same depth as I do in the videos that deal with standard error. Just because I don't want to repeat myself too much. And if you're interested in thinking about these, these code more than I suggest you go and watch the videos on standard error. Okay? So our goal here with the, with this code that I already have set up is just to illustrate what I just described for the forearm practical way, you can have one normal distribution. You can pull out a sample. Determinants mean pull it another sample determined it's me and you'll see that they're not the same. Okay, I just want to go through that practice explicitly and we'll do that a couple of times and then we'll do something a little bit more fancy. So first of all, I want to show you the distribution that we're going to be drawing our data from. We're going to be using a function called our norm, which allows us to pull random numbers. That's the R stands for from a normal distribution, okay? And our normal distribution, they're pulling the random numbers from has a specific shape. Care it has specific properties. We're going to say that the normal distribution has a mean value of 0. So the center of our distribution is going to lie right on 0. And it has a standard deviation of one. Where the standard deviation refers to. Generally, you can think of it as being the breadth of the distribution. If we want to just think of standard deviation in very loose terms. Okay? So I'll just illustrate this here. What I'm going to do here is I'm going to pull out 5000 random numbers from the normal distribution, which has a mean value of 0 and a standard deviation of one. Okay, and I'm just going to plot those random numbers in a histogram. So let's just run this. So this provides us with a glimpse of what the distribution is going to look like from where we are going to be drawing a random numbers. So you can see that the mean value of this district, of this distribution is right on 0 and it has a particular width. Let's just play around for a second just to show you how we can create different normal distributions. First of all, let us change the mean of this distribution, okay, so instead of having the mean value be 0, let's make the mean value 10. And we'll run this, okay? Now you can see they've just shifted the distribution over. So the mean is now. Then, let's go back and return it to a mean of 0. There's our original finding. Notice that the distribution tends to range for the random numbers that we obtained, tend to range from about minus four to about positive four. Okay? Roughly speaking, we can make the distribution wider by increasing this number here. So let's change it from say, one to five. And now if you run that, you can see now we're getting a much wider distribution. Okay? So that's just to illustrate. How we can draw random numbers from a normal distribution with a particular shape. And what we're doing is we're going to imagine that this is, this distribution represents a population that we're interested in studying. It could be a population of trees, it could be a population of hamsters, it could be a population of slime molds, okay? And what we're gonna be doing is we are going to be randomly sampling slime molds or hamsters are trees from the population that we're interested in. In order to obtain some mean value for something like maybe for studying hamsters, we're measuring the length of their tails, Okay? And this distribution would represent whatever it is we're interested in. Okay? Obviously for doing hamster tails then we couldn't use exactly this distribution because hamsters can't have negative numbers. But that's, that's not really the point. The point is that this distribution represents the population of, represents the trait values for some population they were interested in studying. And what we're gonna do is we're now going to pull out 10 random numbers from this distribution. Okay? So instead of point out 5 thousand, we're just going to pull out 10 and we're going to save those ten numbers. This object called sample. Now, if I just highlight sampling run that you can see here are our 10 random numbers. And now we can take the mean of those ten random numbers and we get a mean value of negative 0.17. Let's repeat that process. Let's reach into our population again to obtain another ten measurements and will take their mean. Okay, there we go. You can see that we've got a different mean this time. So instead of negative 0.17, we have negative 0.19 and do this again. And this time we have a positive mean value. Okay? So here are three, what we might call experiments, where each experiments just involves taking a sample of 10 individuals from our population, looking at their measurements and calculating the mean value. So there's one experiment which gives us one mean, another experiment which gives us another mean. And a third experiment that gives us a third mean. And you can see that each time you run these experiments we get a different mean value. Let's a little bit more fancy. Instead of just running this one at a time. Let's just for a few times, Let's run this experiment. Let's run the series of experiments many times. What we're gonna do is we're going to perform the same thing that we just did. But we're going to perform it 5000 times. Okay? And we're going to say you all of the, all of those mean values. And then we're going to plot the mean values from these 5 thousand experiments. Okay? So I'll just show you how we're doing that. This pit, this bit of code here does exactly as what we were doing above, but we just collapsing these two lines of code into a single line. We're just pulling out ten samples from our population. And we just calculated the mean of those tamped 10 samples. And then we're going to store that mean in an object called sample means. Okay, I'll just explain a little bit more about how we're doing that. So sample means is just an object, which is, it's an array which I'm creating. So it will have 5000 places to store data. That's what I'm doing here. Okay, So if I run this and if we look at sample means, you can see that it has 5 thousand spaces as listed the first 1000, it says it's left out the next four thousands that sums at the 5000. But what we have here is 5000 N A values and NA just means we have a missing value. In other words, we've created an object that has 5000 places that can hold data. But all those places are currently empty. And we're going to populate this object with the mean values that we get from this from this part of our code. Okay? We're going to repeat this process many times by putting our data or by putting this line of code within what's called a for loop. And what that's going to do is it's going to just run this code for 5000 times using a counter which we call i. Okay? And so the first time you run this code r is going to be, sorry, I is going to be equal to one. And so the first mean by that we obtain will be placed in the first position within within this array. The second time we do this, we'll store the mean of the second position. The third time we run this, we'll store the mean, the third position, and so on. Okay, so what we're gonna do here is we're going to create an object that holds the results from 5000 of these experiments that we ran earlier, early we just ran three experiments. Okay? So we're going to run, do those experiments now 5000 times and we're going to create a histogram to give us all of those results. Let's run this code. Here is our histogram. Okay? So these are the results of our performing this experiment 5000 times, where each time each experiment involved pulling out a sample of 10 individuals from a single population, calculated a mean for that sample. And what you should notice here is that most of the time our experiments gave mean values that are pretty close to the true value. Remember the true value of our population is 0. And most of the time, our experiments give us a mean value that is pretty close to that true value. Okay, that's, I'm saying that because the, the part of this distribution that's highest is centered right around the true value of 0. But we can certainly see that that our mean values are not always the same. In fact, in some cases, we end up getting mean values that are around one or as far away as negative one. It's still very probable to get a mean value of around minus 0.3 or a positive 0.3. I'm, in other words, we can see variation in the mean values. Okay? This, I'm showing you this for two reasons. The first is just to really emphasize the point that when we draw a sample from a population, we expect variation in the mean value that we obtain. Okay, so if we draw more than one sample, we do not expect those mean values to be identical. That's the meat, that's one of the main points that it wants to emphasize here. The second thing is I want to point out that this distribution has a particular name. It's called the sampling distribution, and we're going to refer to sampling distributions again later on in this video. Okay, so that's what I wanted you to get from our simulations. So let's return to our slides. Okay, we've now use some simulations to really hammer home this point that when we obtain samples for treatment 1 and treatment 2, we do not expect the mean values for those samples to be identical. Even though the truth is that there's no difference between cheaply than one achievement to, and we expect there to be differences because of sampling error. Okay? So I'm going to show you figures like this throughout this video, where these distributions represent the sampling distributions that we just created in our simulations. In other words, this represents the distribution of possible mean values that we could get from samples of our larger population for this treatment. In other words, treatment one will corresponds to a population of individuals. We can obtain a sample of seven individuals from that population and calculate a mean for it. Most often that mean will be somewhere around the middle. But if we repeated that experiment, we could very easily get a mean there or there. Or it's still quite reasonable that going to mean well, even as far out as here, okay, it's less likely to get a mean, um, that's light in this table, but it's still something that can happen. Okay? So that's what these distributions represent. They represent the sampling distribution that we created from our simulations. Now, we've been so far thinking about our thinking so far has involved imagining a series of many experiments. Okay, What I'd like you to do now is focus specifically on what we could get if we just performed one experiment, okay, so we obtain one sample from treatment 1 and one sample from treatment to you. And want you to imagine that our one sample for treatment 1 gives us a mean value that lies about there, which still lies pretty close to the heart of this distributions. This is a reasonable value to expect to arise just by random chance. And our sample for treatment 2 gives us a mean value that lies right here. Okay, so this is really close to the mean of the distribution. This mean value is very likely to rise just due to random chance. Ok? So you can see here that the mean values for treatment 1 and treatment 2 are not identical. That's a point that we. Beaten to death so far, what I'd like to emphasize on this slide, there was something slightly different, which is that the difference between these means, which I've depicted here by this orange dashed line. This is a difference between these means and a difference of this size is perfectly reasonable to expect even when the null hypothesis is true. Okay? In other words, what I'm trying to, What I'm trying to illustrate is that differences as big as this would be perfectly reasonable to expect when the null hypothesis is true, when our sampling distributions look like this. Okay? I don't feel like I've said that in a very articulate manner. So I'll try I'll try one more time. We found one mean here, another mean here. This gives us a difference of this big. What I want to point out is that a difference this large is perfectly reasonable to expect with sampling distributions like this when the null hypothesis is true. Okay? So we've now really emphasize these first two points. This next point builds on everything we've talked about so far, which is simply to say that a correct analysis of an experiment like this will properly account for the sources are properly account for sampling error that will be involved in your experiment. And specifically in the context of this video, I'm thinking about a correct analysis with respect to accounting for independence or non-independence in our data. The last main point that I really want to emphasize here is that the sheep of our sampling distribution is going to depend on our sample size. So if you have a small sample size, your sampling distribution is going to, tends to be wider. Whereas if you have a smaller, sorry, if you have a larger sample size, then your sampling distribution is going to tends to be more narrow. So small sample size, you have a wide distribution, wide sampling distribution. Small, large sample size, you have a more narrow sampling distribution. And what that means is that when you have a smaller sample size, it's going to become more likely for us to get relatively large differences between our treatments simply due to sampling error or simply due to random chance. Even when the null hypothesis or one that I hope null hypothesis is true. We're going to give you an example of that exactly. Very shortly. Can I just reiterated this points that I've just made here. With small sample sizes, sampling error is more, sampling error will more easily lead to larger differences between the mean values for our different groups, even when the null hypothesis is true. And this provides us with a point of foreshadowing, which is to say that if we analyze a data incorrectly with respects to independence, then what we can end up doing is we can end up confusing the effects of sampling error as actually being due to treatment effects. That might not make a whole lot of sense right now, but it will make more sense very shortly. Okay, so let's just summarize the core ideas that we've been building up before, actually put them into practice. So we've been hammering home that all the scenarios that we're going to consider in this video. Imagine scenarios where the null hypothesis is true. But we said that even when the null hypothesis is true, the mean values for the samples that we obtain will almost certainly be different between our treatment groups simply due to sampling error. Okay? We've also said that we can expect there to be larger differences between the means of our different groups. Simply due to random chance when we have a smaller sample size. Finally, we've said that a correct analysis. So what analysis that correctly accounts for non-independence? We'll properly account for sampling error. Okay? So to say that again, a correct analysis, SWOT analysis that correctly accounts for any sources of non-independence in our data, will properly account for the effects of sampling error and then be able to give us. Reliable results with respect to trying to determine whether or not there's likely any differences between our treatments. So those are our core ideas. We're now going to consider three separate scenarios. So here's our first scenario where we have an experimental design. It looks a lot like the first design we're looking at. But now you can see that we have non-independence in our experiment. So these three data points are, have something in common that would lead us to worry about non-independence. So do these, sort of, these, sort of these. So perhaps these three data points all involve measurements from siblings within a family. Okay? So these represent one family, these are another, these are another. These are another. Or maybe these represent three data points at all experienced a common environment. So maybe these all represent colonies that came from one Petri dish. And these were three other colonies that came from another Petri dish and so on. Maybe these all represent animals that came from one cage. So we have three data points from each of four different cages. Maybe these represent multiple measurements from the same individual. So we've measured the same individual at three different times to give us three points, three points and so on. Okay? The point is that we have some cause of non-independence in these data. Now I'm going to tell you ahead of time that if we did not analyze these data correctly, then we will introduce pseudo replication and that will lead us to some incorrect conclusions. So I'm just summarize the main points I want to start with for this, which is already made this point quite firmly. So we've established that we have non-independence in this experiment. And I said that the next point I want to build from that is that if we were to treat all of these data as being independent, then what we'd be doing is we would poorly characterized the effects of sampling error when we're trying to understand the, what, when we're trying to determine evidence for differences between our two treatments. Okay, and we're going to illustrate that specifically here. Okay? So these plots represent the sampling distributions for two different scenarios. The first here represents the true sampling distribution for one way in which we could legitimately analyze these data. So one way that we can handle these data is to simply take the mean value of the data from a, the mean value for b, the mean value for C, and the mean value for d. And by doing that, we be taking our 12 data points and collapse them down into four data points to data points for treatment 1, 1 average for a what I wish for B, and two data points or treatment 2. So one average for C and one average for D. Okay? And that's why have here n equals two. So in this case, these represent the correct sampling distributions for a correct analysis of these data. If instead, we were to imagine that all of these data points were independent and ignore the fact that we have non-independence here, then what we could, we could do, or what we could do incorrectly, is we could say that we have a sample size, not have to, but a sample size of six. And you remember earlier that I said that when we have a larger sample size, that will tend to lead us to have a sampling distribution. It's more narrow. And that's why in this case, given you sampling distributions that are more narrow than in the correct case. So again, these are what the sampling distribution should look like. But this is what our sampling distributions would look like with an incorrect analysis of our data where we're assuming that we have six independent data points where the reality is we really only have to, okay? Or we can analyze the data in a way where we treat them as having two independent data points. Let's imagine for this experiment, we got mean, a mean value for treatment, one that landed here. And I mean value for treatment to the land here. Okay? You can see then that this effect size or this difference between treatment 1 and treatment 2, which is given by this orange dotted line. This difference between these two treatments is perfectly well expected with this particular set of sampling distributions. Because you can see that it's perfectly reasonable to expect there are saying that poorly. There's a high probability. Getting a mean value of this within this distribution, because this mean value still lies relatively in the center of this distribution. And the same thing happens here. We have a high probability of getting this mean value for this sampling distribution because this mean value still relatively lies relatively within the bulk of the center of this distribution. As a result, it's perfectly reasonable to expect to get a difference between these two means that's this large when the null hypothesis is true. Okay? Do we reach the same conclusion when we consider this scenario? Okay, well, we can answer that question just by taking these arrows and this dotted line, just essentially sliding them over to this scenario and we'll just do that here. Okay, So what I've done as judge, a slit this era over there, this arrow over there and moved the dotted line over. Okay? What you can see now is that in this scenario, it's now much less probable for us to get a mean value of this for achievement one and a mean value of this for treatment to you. Because now these mean values are aligned more in the tails of these distributions for treatment 1 and treatment 2. And as a result, it becomes much less likely that we could get a difference between these mean values just due to random chance compared to this scenario. Okay, so I'll just say it again. In this scenario, you can see that it's less likely for us to get a difference this large between our treatments. Just due to random chance when the null hypothesis is true compared to what we saw in this scenario. As a result, if we were to analyze our data in this way. So using a way where we're using an approach that assume that all the data were independent when they weren't. We'd be much more likely to conclude that a difference like this, a difference this size, would be unlikely to arise due to random chance when the null hypothesis was true. And as a result, we would conclude that we would have more evidence to reject the null hypothesis. In other words, if it's relatively unlikely for us to get a difference this large between these treat, two treatments. Just due to random chance when the null hypothesis is true. If it's unlikely for this result to arise when the null hypothesis is true, we can take that logic and turn it on its head and say, Okay, if it's unlikely for us to get a difference this big with the null hypothesis is true, then that counts as evidence against the null hypothesis. And we now say that we have some reasonable evidence to reject the null hypothesis and to say that the null hypothesis is false. And we're more likely to conclude that there is actually a difference between treatment 1 and treatment 2. Okay? And that would be a false conclusion because we've already said that the truth is that there is no difference between treatment 1 and treatment 2. So let's just wrap up that, that, that, that discussion. If we were to treat data like this is all being independent when they're not, then what our analysis is likely to do is it's likely to essentially mistake and the effects of sampling error for being effects of the treatments. And what that will lead to is artificially small standard errors and artificially small p-values and others will get p-values that are smaller than they should be. If we'd actually calculate, if we'd actually don't, we're just getting p-values that are smaller than they should be. In other words, if we were to analyze these data correctly, we, we tend to get larger p-values than we would get if we analyze these data incorrectly. And as a result of all of this, treating it non-independent data as independent will lead is a smaller p-values and will lead us to increase the chances of rejecting our null hypothesis when we shouldn't. And that's going to increase our type one error rate or increase our rate of false positives. Okay? That in essence is the big danger that comes from not accounting for non-independence correctly. That's our first three scenarios. Let's move on to our second scenario. So in this case, so here's our second snare. This, the very first snare that we, that we started with. This is, was the snare that we use to establish our core concepts. But we're focusing now on the scenario on the right. In this case, you can see that all of the data points here, we have six data points. They, none of them share a letter within treatment 1 or within treatment 2. That means that within each of our treatments we have independent data. Okay, that was not true in our previous scenario. And this scenario we had non-independent data within our treatments. In the second scenario, we have independent treat, we have independent data within our treatments, but instead we have non-independent data between our treatments. I'm going to tell you that this kind of experimental design does not worry us with respect to C two replication. In other words, if you were to analyze, if we were to run an experiment like this, given what I'm presenting to you here, there, we wouldn't have any real reason to worry about pseudo replication. And that's because that's because we have independent data within our treatments. And as a result, these data will effectively characterize the sampling error that we expect to get when calculating a mean for this treatment and for that treatment, okay, in other words, I've said this. What I've said here is that the within treatment variation will effect, will effectively represent the variation in the population. Once we write, it will effectively represent the population. And as a result, we're able to effectively determine the effect of sampling error without having to do anything fancy like taking mean values like we did in this previous scenario. Okay? So in this scenario we do not need to worry about pseudo replication. Having said that, I want to point out that even though the first experiment that we considered and this experiments, neither of them raise any concerns with respect to pseudo replication. But that does not mean that these experiments are identical, okay, there are important differences between these experiments, even though neither of them lead us to worry about C to replication. The main difference lies in the fact that we have non-independence between our treatments for this experiment here. And that actually can give us some specific advantages compared to this experiment on the left. This kind of experimental design on the right can be analyzed in a way where we can leverage the, the non-independence between our treatments in such a way that we can have a more powerful test of the differences between the treatments here than he could hear. So for example, we could analyze these data with what's called a paired t-test, which would be a more powerful test of differences between these treatments. Then we could analyze with, with this experiment, which we'd probably analyses something like a two-sample t-test or Welch's t-test. Okay? Exactly what test you would use is not really the point. The point is that even though these two experiments, neither of them give us any reason to worry about C to replication, but they're not. Equivalent experiments in this experiment has certain features that we can leverage with particular analysis tools in order to give us a more powerful test of differences between our treatments. That's bit of a tangent. But really the main point is I'm beating a dead horse here, okay? The main point here is that this experimental design does not worry us with respect to C two replication. So now on to our third treatment and we're going to our third scenario. And I'm going to do something slightly different with the snare. I'm going to make a different kind of point. Okay? This experiment represents an extreme force, sorry, an extreme case of non-independence. We can see in this case that all of the data within treatment 1 are non-independent. At all the data within treatment 2, I know one independent. So for example, maybe these data are not independent because these all represent measurements from animals that were living in one cage, for example. And these all represent animals that come from another cage. Okay? Or maybe these all came from the same mother and these all come from another mother, for example. Okay. What I'd like to point out is that there's two main things I want to point out in this experiment in this slide. The first is that whatever the cause of this non-independence, that cause of non-independence is completely confounded with the treatment effect. Okay, So in other words, let's imagine that the cause not a dependence was the cage at the animals were living in. So these animals all lived in one cage. These animals all lived in another cage. In that case, cage is entirely confounded with our treatment effects. In other words, let's imagine that we analyze these data. And we got some evidence that there was a difference between the mean of this and the mean of that. Because of how this experiment is set up, we would not be able to know whether or not these mean values were different because of these treatment effects. And that's because these mean values could be different for two different reasons. One is they can be difference because these animals experience different treatments. Or they could be difference because these animals experience different cages. That's what I mean by saying that treat, that are Cage is completely confounded with treatment. In other words, there's two differences here between these groups. One is the treatment and the other is the cage, okay? And so as a result, if you have an experimental design like this, you could not actually determine whether or not there is truly difference between the treatments. Simply because you couldn't disentangle the treatment effects from the effects of your sources of non-independence. That's the first thing that, that's, that's a first less than I want you to get from this slide. That might seem a bit strange because you might say, who would actually do an experiment like this? Believe it or not, there are published studies that do exactly this. Okay? So if you are reading through the literature, it's perfectly reasonable to come across in areas like this. In which case the people that I've published those results are publishing the results that are entirely meaningless. Ok. So this kind of mistake is made in the literature. So lookout for it. That's the first thing. The second that I want to make from this slide is a bit more general. Where it's much more general and it's a, it's a point that I'm only going to touch on. Okay. And that's this point here that I have highlighted, which is that we can think of pseudo replication as being related to these confounded effects. Okay? And that's me sound like a fairly profound thing to say. But I'm not going to say anything more about it. I'm just going to plant that seed in your mind to say that all of our discussions about pseudo replication can be thought of in terms of confounded effects. I'm going to sing more about this here. But I'm going to point you to a resource in a moment where you can learn more about that perspective. Okay? Right now what we're gonna do is we're just going to summarize all the points in this video, okay? So pseudo replication occurs when we treat non-independent data as if they were independent. When we do that, we increase the risk of concluding that we have differences between treatment effects. When in fact, the differences we see are actually just due to random chance or that just actually due to sampling error. And so overall, what we end up doing is we end up increasing the chances of finding a false positive or making a type one error. That's what happens or that's what can happen when we treat non-independent data is independent. The third here is just a point I just very briefly introduced on the previous slide, which is that we can think of CDO replication as being related to confounded effects. Okay, and I'll say a little bit more about this on the next slide. The last few points that I want to make or just really kind of points that I want you to carry forward when you're designing your own experiments and analyzing them. And that is that when you're designing experiments, it's really important to be conscious of sources of non-independence. And either design your experiments in a way to avoid them or to make note of them so that you can analyze your data appropriately to account for those sources of non-independence. And this last point here really just piggybacks on the previous point that I made. Which is that when you are at the stage of your project where you're analyzing your data, you want to be very conscious of sources of non-independence and account for them when you analyze your data. So I want to end by pointing to this book. It's an excellent book by Ruxton and Cole gave. This book is highly accessible. It's written for an undergraduate audience. It's written beautifully. It's very accessible for undergraduates. Having said that, I know of PhD students who've benefited from me in this book. I know about, I know of, of site. I know about professors who have benefited from reading this book. I have benefited from this book. It's an excellent book and highly accessible. And this book has an entire chapter devoted to see to replication. And in that chapter, they adopt this perspective of thinking about SR replication in terms of confounded effects. So if you're interested in learning more about pseudo replication and in particular, getting this other perspective on to replication. A highly recommend getting your hands on this book and reading the chapter on pseudo replication. On that note, I'm going to stop this video here and say hope it's been helpful. And I'll say, thank you very much.