Okay, and our previous videos, we talked about what independence is and what non-independence is with respect to data that are used for an experiment or data for any type of analysis that we might be wanting to conduct. And we also talked about how non-independence can leads to pseudo replication. We talked about how pseudo replication can influence our ability to make correct conclusions when we're analyzing a data set. In this video, we're going to build on those themes in two ways. First, we're going to talk about sources of non-independence by considering some specific examples. So I want to start introducing you to various ways in which non-independence could creep into your experiment and make your life more difficult than you'd like it to be. Second, we're going to shift gears a bit. And we're going to talk about how pseudo replication can change an interpretation of an experiment. So let's start by considering some examples of non-independence and experiments. And we'll focus to start with on how shared environments can leads to non-independence. So let's start by considering this experiment where let's imagine we wanted to test whether or not a drug influences weight gain in mice. And our experiment looks like this, where we have two cages, cage one and cage. To add each cage, we have a number of mice. We can imagine that the mice themselves are independent. And within cage one, all the mice are drug free, so they did not experience the drug. Whereas in cage to our series of mice did experience the drugs. They were treated with the drug. And let's imagine we ran this experiment for awhile and we found that over the course of the experiment, on average, the mice and cage one gained one gram of mass. Whereas in cage to on average, the mice gained five grams in mass. And we analyze these data. And our analysis gave us a p-value of that was less than 0.01, which would then lead us to conclude that there was evidence that the drug might have influenced or leads us to conclude that the weight gain was different between our two different situations. Okay? If you stop and think about this experiment for a moment, hopefully, you'll notice that this experiment is terrible. Our p-value might call, give us reason to celebrate yea or drug does what we want it to do. But with closest scrutiny, you'd find this experiment cannot allow us to come to a conclusion regarding the effect of the drug itself. And the reason for that is because the drug treatment is confounded with the cage differences. So ideally, when we conduct an experiment, we want all the conditions that individuals experience to be the same on average between individuals that experience one treatment versus individuals that experienced another treatment. And you can see that that's not true here. You can see that the individuals that did not experienced the drug were different from those that experienced, that were different from those that did experience the drug in two ways. First of all, they differ in whether or not they received the drug. But they also differed with respect to the environments that they, that they experienced. Specifically, they had a different cage. And so what we can say, what we can notice here is the drug treatment is confounded with the cage differences. And as a result, even though our p-value told us that there is evidence that the weight gain maybe different between this set of mice and this set of mice. We cannot conclude that, that difference, or that this evidence for the difference in weight. We cannot conclude that this evidence for differences in weight gain are due to the treatment. And that's because it's also possible that the difference in weight gain could be because the mice experience different environments. Maybe this cage was particularly lush. Meat was much warmer environment than this first cage, so that the individuals in the first cage needed to spend a lot of their energy on keeping themselves warm and they couldn't spend that energy instead on growing. For example. The point here is that this experimental design confounds the cage effects with the drug effects. So how can we fix this type of experiment? Well, there are three, at least three different ways to fix this. The first is just to put all the mice into a single cage. Ok? And by doing that, we ensure that all the mice have the same environment. This might not be very ethical for the mice because crowding the mice in this way might just not be good for their well-being. And so a second solution would be to have cages with mixed mice. In both cages, we have mice that do and do not experience the drug. But we've been able to space the mice out into multiple cages. This is also a reasonable experimental design in order to test the effect of drug on the weight gain. Alternatively, we could keep mice. That experience. We could assign all the mice within a particular cage to a particular drug treatment, but we replicate the cages. So in this case we have one cage where all the mice were drug free. Another cage where all the mice experienced the drug, et cetera. In this case, the, the mice, that in this case what we can do is we can recognize that the three mice within this cage are not independent. And so what we might do is we might just take the average weight gain for the three mice in this treatment, right? In this cage, to the same thing for the three mice in this cage and this cage and so on. Which will give us then three independent measurements of weight gain for each of our different treatments. One measurement here for drug free, another there, another there. Likewise for a three cages where the mice did not experienced the drug. And so this gives us then three independent measures of weight gain for each of our different treatments. And what we expect when we replicate our data in this way over the various cages is what we expect is that on, on average, the average condition that the mice experience among these various cages for the drug treated individuals will be similar on average to the environment that the drug free mice experienced in their cage environments. Let's consider a different experiment now. Let's imagine that we wanted to determine whether or not temperature influenced stem cell development. And we know that incubators can be very expensive pieces of equipment. So let's imagine we just had to incubators. And we set the first incubator at 37 degrees. And we had 20 tissue culture is and incubator one. And then we set incubator to, to a higher temperature of 40 degrees. And again, we have 20 tissue cultures within each experiment. If we were just to compare the cell development here versus here, could week. And let's imagine we found a difference between here and there. Could we conclude that the differences in cell developments that we perceive here are due to differences in the temperature. Stop and think about that. Well, the answer is no, you can't. And the reason is the same reason that we saw on the previous cage experiment. And that is that we actually have two differences between these sets of tissue cultures. One is the difference, we're interested in, the difference in temperature, but the other is just the incubators. And as much as we might hope that Incubators are identical, the fact of the matter is they will not be identical. And so it's possible that any differences that we see between these treatments could be due to differences between the incubators in knots because the temperatures. So how can we fix this? Well, what we might do. You might repeat the experiment by switching the temperatures the incubators. In other words, we can look within incubator one that can do one trial where you have the temperature set at 37, and then you repeat it, which with the temperature at 40 degrees Celsius, ideally, you'd repeat this multiple times. And so you don't just have one measurement per incubator at each temperature. But the main point here is that by swapping the conditions within each incubator, they can then tease apart the effective incubator versus the effect of temperature. Okay? So those are a number of examples that we've walked through in detail. To continue this, I'm not going to go through a whole series of experiments. Instead, I'm just gonna list a number of different ways in which non-independence can creep into your experiments because of shared environments. If you're working with, let's say drosophila, where you are producing growth media. Then you have to recognize that we can have non-independence through different batches of growth media. So if we use different batches of growth media for different treatments than we would be confounding the batch the batch of our growth media with our different treatments. And Matt could severely compromised our experiment. If we were working. We've already talked about incubators. Let's imagine that we were doing an experiment with just one incubator. And we had different shelves within that incubator. I've heard from a good source, a very reliable source, that if you measure the conditions among shelves within an incubator, those conditions among the shelves are not identical. So if we did something like we set up an experiment where we had different treatments set on different shells within an incubator, then the different treatments would not be independent of the conditions on the shelves. And that non-independence between the treatment and the conditions in the shelves would could lead us to incorrect conclusions. Similarly. So here we've talked about different conditions within an incubator. We might have different conditions within different areas of an animal house. So that if you set up different treatments within different parts of the animal house, then those treatment effects would not be independent of the conditions within the animal house. Or in other words, the treatment conditions would be confounded with the environmental conditions that are set within the animal house and that could lead to incorrect conclusions. A few more, a few more examples. Let's imagine that we're making different cultures from an animal. If we have different cultures than each of those cultures can represent different environments. And we want to be careful about how we use those cultures and our experiment. Finally, I'll just point out that animals themselves, or if you're working with plants, plants themselves, can represent particular environments. So the environment within my body is going to be a different environment from the environment within someone else's body. That might be because of how I've treated my body or where my body is situated. So when a hot location versus another location, or just because the genetics of my body, the genetics themselves create a particular environment. Okay? So if you are making multiple measurements from the same animal or the same organism for any species, not just animals, then those multiple measurements from the same organism will not ness, will, are very likely to be non-independent. Okay? An example I've given here is just if you're measuring, say, neurons that are recorded from the same animal, it's very possible that those neuron measurements will not be independent because of the shared conditions within an animal. Okay? So that's a whole bunch of different ways in which shared environments can lead to non-independence. I'm just going to devote one slide to the next source of non-independence. And that's just relatedness. We, we know from our basic understanding of the world that individuals that share genes are more likely to be similar to one another. Because genes influence how organisms develop and influence the traits that they express. So here's my one slide. Here's a fam, here's a photo of a family. They will share genes and most likely. And as a result, they tend to have more similar traits to one another than to other randomly chosen individuals with Anna population because of the non-independence that's brought in from similarities due to genes. So whenever you have relatedness in an experiment or a study, you need to be conscious of how that relatedness can, can impact the, your, how you, how you analyze your experiment, and how you interpret your experiment, and indeed how you design your experiment. So that's our discussion of how non-independence can be introduced into our studies. We're now going to shifts topic a fair amount and we're going to talk about something that's both very related and also both very different. And we're going to ask how pseudo replication can change our interpretation of an experiment. In many cases, the experiment might just become uninterpretable because of pseudo replication. In other words, if you analyzed your data incorrectly. And by that I mean you analyze your data in a way that did the way you treated non-independent data as being independent, then the results of that analysis will be unreliable and we will not know whether or not to interpret the results. We won't know how to we won't know how to interpret the results of an analysis like that. In the most basic sense, we won't know whether or not to trust the results from an experiment that's been analyzed and appropriately. Ok. I want to point out that in some cases, pseudo replication, well, not necessarily completely destroy our interpretation of experiment. In some cases, we can still salvage some conclusions from an experiment, although we have to be careful on how we interpret the results. And so I'll give you an example here. Ok, we're going to explore this topic with this one experimental example. Let's imagine that we wanted to examine the effects of logging on the diversity of small mammals in a, in an environment. And so what we can point out here is that logging can be pretty devastating to an environment. And we would expect that when we change the environment in such a drastic way at this may influence the types of species that will now live in this environment. And so we can expect that the number of species of small mammals that live here might be very different from the number of species of small mammals that might have lived here when there was a full standing forest. Ok. So that's the general biology that I want you to imagine that we're going to investigate. In this case, our null hypothesis would be that there is no effective logging on the diversity of small mammals or there's no effective logging on the number of species of small mammals. And I want you to note that when we're answering this question, we're really thinking about how logging affects species diversity in general. Okay? In other words, we want to know about how logging influences the diversity of small mammals for the average forest. So we're thinking about this very generally. I want you to consider this experiment and ask yourself whether or not this experimental design can answer our question. Specifically, let's imagine we went into this region. I'll tell you this. I took this picture. This is a picture that I took for us an experiment that I was doing had nothing to do with logging and species diversity, but it's still useful. Let's imagine that we went into this region, which happens to be in northern England. And you can see that we have an area that was logged on the left and an area that has not been logged or at least not recently. On the right. And we want to know whether or not logging influenced the number of species of small mammals. And so what we did is we went into this logged area and we set up four different areas, 1234, where we were able to count the number of species of small mammals in those four areas. So that gives us four measurements. And then we did the same thing in this area that has not been logged or at least not been logged recently. So given this experimental design, can this experimental design help us understand the effect of logging on the diversity of small mammals. Stop and think about that for a moment. Okay. Now that you've given that, I think you'd say the answer is no, or at least the answer is no. We cannot use this experiment in order to understand the effect of logging on, on on the diversity of small mammals in general. And the main reason for that, I've kitten given two bits of text here that that speak to this issue. I'll start with the bottom one. The main issue here or one of the main issues is that logging, in this case is completely confounded with any other differences between these two different regions. So for example, we can recognize, yes, this area was logged recently. But, and so we might imagine that if we found that on average the number of species of small mammals was lower in this region than in this region. We might be tempted to say that logging decreased. That logging itself was responsible for decreasing the diversity of small mammals. But it's also possible that there can be many other differences between this region and this region. So imagine, for example, that this region on the left was used as pasture for large mammals. So like for cows. In that case, it could be at the presence of cows themselves, influence the diversity of small mammals. And if that's the case, then we couldn't actually tell whether not a decreased abundance or a decreased number of species of small mammals on the left is different from that on the right. Because of the effect of logging, or instead because of the effect of the presence of counts as an example. Okay? So one of the major points here is that with this experimental design, where we just have one representation of an unlocked forest and logged forest. In this case, logging or not, is going to be completely confounded with any other differences that there might be between these two different regions. And as a result, we cannot conclude that any differences that we see in the diversity of small mammals between these two regions is because of logging. It could've been due to some other effect that is systematically different between these two different reagents. So how do we get around that? Well, we would get around that problem by having multiple regions that are law, that are logged versus unloved. And by looking at multiple unlocked forests and multiple logged forests, then we would be able to say, then we will be able to argue that on average, the main differences between those multiple unlocked regions and those multiple large regions would likely be because of the effective logging itself. And so in that case, if we replicate our logging an unlocking achievements multiple times, in that case, we could actually infer that the differences in species diversity between those two treatments could be attributed to the effect of logging. Ok. So this, this other piece of text that I've written here just notes that in this experimental design that I've depicted on this slide, we have no true replication of the logging treatment. And in order to be able to make any conclusions about logging itself, then we would need to have multiple logged and unloved forests. We need to have more than one on log in logged forest and more than one logged forest. So those are the problems with this experimental design. If we want to understand the effective logging on species diversity. Okay. Does that mean that this experiment that all of our work here is completely lost? Well, no, not if we interpret our results appropriately. So what we've done here is we can recognize that we've conducted our experiment into one single forest, okay? Because of that, we do not have the appropriate information or the appropriate data in order to say anything about the effects of logging in general. So in an experiment like this, we have to give up any efforts to try and say something about logging in general, we lose that generality. However, we can say something about this particular forest. So this experiment would actually be fine if our interest was not about logging in general, but was about understanding the effects with but what was about understanding differences within this particular region. In other words, if our true interest for this experiment was to test whether or not the small, to test whether or not the diversity of small mammals and this region was different from this region, then that is absolutely fine. If that was our question. If we want to know whether or not the diversity of small mammals on this side of the fence was different from the diversity of small mammals on side of the fence. This experimental design would be totally appropriate for that question. Okay? It's just not appropriate for making more general conclusions about logging for the average forest. So the main point here is to say that the merits of a particular design will depend on the studies, particular goals. If you want to understand something in one particular region, then it's fine to study that one particular region. If you want to be able to make conclusions much more broadly than you need to be very careful about how you design your experiment. You will probably want to be able to, you'll probably wants to study your phenomenon and a much more broadly and consider a number of different environments. We'll stop there. We've covered a lot of ground. We've talked about how sources of non-independence can creep into your experiments. And we've also talked about how even when we do have confounded variables within an experiment, that does not necessarily mean that our experiment is entirely worthless. What it can mean in certain circumstances, like the example that I've given you, that we still can learn something from our experiment. But we're just going to have to, we're probably going to have to change the nature of the question that we're asking. And I will stop there and say, I hope this video has been helpful and thank you very much.