Okay, in this video, we're going to analyze some data in order to demonstrate how we should approach and analysis when we do not have a significant interaction in our results. We spent so much time in our series of videos focusing on interpreting interactions that we've really neglected the very common situation aware an interaction doesn't actually occur in our data, and that's what we aim to address in this video. The biology of this example involves sexual conflict. And the data come from this paper here. And that is where this paper, I'm just going to describe the experiment to just second. Dah dah dah dah da, where are we? So the actual data themselves, I obtained the actual data from Whitlock and shooters text, The analysis of biological data. They provide data for many examples, but those data were pulled out of that original paper by Whitlock includer. So what this paper by Barnes at all looks at is they're trying to understand how a particular ejaculate called SP influences the lifespan of mated females. It's very common in nature for ejaculates from males to influence the females sick, they can change their metabolism. They can change all different aspects of how the female body works. And in for software melanogaster, it's founds that made it females often live shorter lives in unmade ID females. And there's evidence that this protein, that the males inject protein, which was just called SP, has a role in reducing female lifespan. The question from this paper is how, how does this protein actually lead to reduced female lifespan when they mate? And one of the possibilities, one of the possible ways in which SP could reduce female lifespan is by causing females to produce more eggs. And then this extra cost of producing these additional eggs. That additional cost is what causes the females to live shorter lives. Basically, there's a large cost of reproduction. So to assess this hypothesis, the, the researchers in this paper designed a 2-factor experiment where they, where they manipulated two factors. First of all, they use mutant forms of females. To where they were. The mutant forms, our feet are sterile so the females do not produce any eggs. So in this case, we do not expect the females to be susceptible to this SP protein if this hypothesis for the effective SP is correct. So we have females that are either fertile, Which do you produce eggs or stairwells, They did not produce eggs. They also manipulated the, the frequency which with, which, with which females mate with males. They did this by manipulating the males. I won't go into details if you're interested in that, you can go look at the paper. But they basically manipulated the males in such a way that females would experience a high cost of mating with males by creating situations where females mate frequently with males. Or females would be around males, but where medium was less frequent. And so we expect in this situation there to be a lower cost of mating for those females. Okay, so where this experiment manipulates both the female sterility and the frequency with which females mate, which would then expects to also mean the amount of S P that the females receive from males. We'd expect that females that mate more frequently with males should receive more of this SP, ejaculate. So that's the general experimental design. These are the numbers for the sample size in each of our treatment combinations. And you'll notice that our data are unbalanced slightly, but there are unbalanced. So before we dive into the analysis, let's just think a little bit more about the biology. And this is something that I would like to do more in our examples. Because the biology and the analysis itself should really be going hand in hand. And this particular data set lends itself really well to doing that. So before we analyze the data, let's just set out some expectations. Will ask three different questions. First of all, if pretty seem more eggs is what reduces longevity, then do we expect the fertile females and the sterile females to live the same amount of time? Thinking about that for a moment. And the answer is no. So if producing more eggs is what reduces longevity than we would expect that sterile females should live longer than fertile females. Our second question is, if mating, if it's mating per se, that's costly to females, then and, and this we expect. So I'll say it again. If mating per se is what's costly to females and that cost of mating is what decreases female lifespan. Then do we expect females to have similar lifespans when they're exposed to lots of mating versus less Meeting. We expect them to have different lifespans. We expect that the females, if this hypothesis is true, we expect that females that mate less frequently should be able to live longer. If it's mating per se, that decreases female lifespan. And finally, if the SP increases egg production and as a results that egg production lowers longevity, then we can ask, do we expect the low and high cost males? In other words, do we expect the effect of exposing females to more or less Meeting two affects the lifespan of fertile or sterile females. Similarly, think about that. That's just, no, we don't. And that's because if S P changes female lifespan because of a cost induced through egg production, then we do not expect SP to affect lifespan in the stairwell females because a sterile females cannot produce anymore eggs. And so any cost of SP, through egg production will not be realized in the female, in the sterile females, but it can be realized in the fertile females. And this slide here just summarize what we've said. So base and our first question. We expect that sterile females to live longer than fertile females. We expect low cost, low meeting. Females that experienced low cost meetings to live longer than females that experienced high cost meetings. And we expect the high versus low cost male treatments to effect lifespans of fertile females but not the sterile females. Okay, so those are our predictions that's driving our biology. Now with all that in mind, let's go look at these data. So the data set are in this DataFrame or this CSV file called female lifespan. Let's just start by exploring these data. Let's use our STR function. So you'll notice here that we've saved the, the data in this DataFrame called Live. And you can see that we have three different columns. One called lifespan, one called treatment, and one called fertility. Where treatment has two levels, one called high cost and one called low cost. And then for our fertility treatment has either fertile females or sterile females. Let's look at these data another way. As a summary, live. And here's our summary of our data. So again, just highlighting that we have three different columns of data. Fertility has these two different levels of fertile and sterile. Treatment has two different levels of high-cost trust is low cost. Some of my students had brought to my attention that when they use the summary function, they get output that looks different from this. I'll just say, I'm not exactly sure why that would be, except it's probably just because I'll bet that those students are using a more up-to-date version of R than I am. And there might be small differences with the output for our, among versions. That's my guess. Okay? And finally, let's just actually take a look at the top of the DataFrame. So we can say head. And so there's a top of our dataframe. And just for overkill, you can see we have lots of data. R is only giving us the first 333 lines. But as we scroll up here, you can see both fertile and sterile females and you can see high cost. Oh, we only see high cost data. Okay, so these data are organized such that we can only see the high cost data at the top. If you want to see the bottom of this spreadsheet, the bottommost DataFrame. We can say tail, whatnot, pale tail. And there we go. So you can see the low cost treatment is organized in the bottom section of this DataFrame. Okay? As always, let's I was about to say, let's start by plotting the data, but there's one more thing. It's very handy to be able to test whether or not your data are balanced or not, because that influences how we analyze the data. I've already given you a description of the dataset which tells us that the data are unbalanced. But I'll just walk through how we can do that very easily in R. Sometimes if your dataset smaller, you could just inspect the dataset yourself. And even just counting things. Although that's a little bit error-prone and perhaps take too much time. Instead we can, I'm just going to use the library, library called du by. It's not essential that you do this, but it's not a bad idea. And so we'll say Summary by where the By is capitalised, you'll notice. And then what we want to do is we want to give our Y variable, which is going to be lifespan days. That's because our hypothesis is that the treatment or that the The mating frequency, treatment and the fertility females, those'll be factors that we're hypothesizing will determine or, or affect lifespan. And so lifespan then means it's, This means that lifespan is our dependent variable, which we plot on the y-axis. So we'll say lifespan and tilda, say treatment plus fertility. Now and I list these two factors on the right side of the tilde a in the function summary by what that means is that whatever calculations we ask summary by to do, it will perform those calculations for every combination of these two treatments. So we should get results for all four of our treatment combinations. Say data equals, live, and we'll say function or fun equals. And then c, let's just ask for the mean, just for, just for the fun of it. And then we'll ask for length. Well, what length does is just counts the number of data points in this, whoops. The number of data points in this variable that we've named on as our dependent variable. Okay, so it's this variable length. That's what's most informative to us if we want to be able to determine whether or not we have balanced or unbalanced data. So if you run this, we can see that we get output for all for treatment combinations. High-cost fertile females, high cost sterile females, low-cost fertile females, and low-cost sterile females. And again, you can see this is really what we're looking for. You can see that our data are unbalanced. Okay? So now what we want to do is we want to plot our data. So we're going to say boxplot. And I'm just going to save myself some typing here. Copy that, put this down here. And you may remember from a previous video that if we put an asterix between our two factors that we're interested in. Then the boxplot function will plot our dependent variable as a function of all the combinations of these different factors. That's where he put this asterix here. Now, we're going to say data equals lie, not love. That's a nice sentiment. Will say live, I'm not going to bother specifying the, the, the range for the y-axis here because I'm not trying to produce a publication quality figure. We're just producing this figure in order to generate some expectations for what we see next in terms of our assumptions, in terms of the results, et cetera. So let's start by running this. We don't get everything listed here, so it's just expand this a bit, okay. So let's start just by thinking about our assumptions. First of all, these box plots are very nicely, normally to the very nicely symmetrical, which suggests that our data are likely normally distributed. We do have a couple. There's some slight asymmetry in a couple of these treatment combinations because we have a number of potential outliers at the bottom here. But by enlarge, normality seems okay. Variance is generally pretty similar. It's not exactly similar. What we'll get a better sense of this when we look at our residuals and when we explore our data. But you can see the box plots are not tremendously different from one another with their widths. So that's okay. What about our predictions? We predicted earlier that females that experience more matings will probably live shorter lives. So to test that idea, we can look within a particular fertile treatment for females. Here we have data for fertile females and here idea for fertile females, but we have high versus low cost mating environments. And you can see in this case that the low cost females up here to live longer than the high cost females. So that's interesting. That's consistent with original expectation. We see the same thing when we look over here, when we look within the treatments where the females are both sterile. So when, when, when we look at female stirred and we look at sterile females and the high cost mating environment. You can see that they live less. That their lives are shorter than females. It sterile females that experienced the low cost meeting environment. So this result here seems to be consistent with this result. Also, we can see that it looks like the difference between the high costs and low cost environments seems to be pretty similar between the fertile and stereo individuals. That would suggest, if that's true, that would suggest that we do not have a significant interaction. Which would suggest that our hypothesis that S p decreases lifespan because it causes females produce more eggs. The lack of any signaling interaction here would suggest that that hypothesis is, is wrong. So, so far, these data suggest that our, that we do not have any support for that hypothesis. Our other question was whether or not there was a difference in life span between stereo and fertile females. To figure that out, we'd have to take the average of our fertile females over our two mating environments and compare that. To the average for the sterile females averaged over the two medium environments. And that's a little bit hard to do by eye, but given how similar the results are, it doesn't it looks like if there's any difference in life span between the sterile and fertile females, it's not going to be a very big one. Okay, so those are our predictions. We could use Chip charts to add points to this. But just for the sake of time, since my videos that and that analyze data often end up being very long. Just for the sake of time, we're not gonna do that. Okay? So at this point we wants to start modeling our data. So let's now create our modal. So we'll say live dot lm. That's a book called The object that saves the output from our linear model, say LM. And then we wants to have the same womb, the same information as we had in our box plot. A point out, this is a shortcut that I have not discussed yet. If we use an asterix between our three different treatments in the lm command. That's actually a nice little shortcut because this way of writing the model is a, means exactly the same thing as saying treatment plus fertility plus their interaction. So I could write this out as female plus sterility plus treatment, colon fertility, as I've done previously. But just to change things up a bit, I'm just going to demonstrate that if you replace all of that with just an asterix, you'll see we end up getting that the model runs and exactly the same way. Now remember we had data that were unbalanced and so we want to account for that in our analysis. So we're going to list are contrasts. So a contrasts I equals, and then we have a list. And we have to name each of our factors. So treatment, we're saying is contrast some. And fertility equals the same. Contrast some. And that's pretty much what we need. Just don't forget. Well remember that when we check our P values, we're going to use the ANOVA function with a capital a, which comes from the car library. So we'll just submit that right now as well. Okay? Now, let's give ourselves a bit more space. We don't need to look at that. Figure so much right now. Now before the p-values, you want to look at our, at our residuals. You want to see whether or not our data and meet the assumptions. So we'll say plot live lm. And let's see, we get, okay, so this first plot allows us to test the assumption of equal variance. And you can see that the data points are slightly more tight at this end than at this end. But really that's not enough of a difference to really make me worry. We can try log transforming the data. Just to explore the data a bit more to see whether or not that improves things. But this is really not enough for me to be concerned about. The data are really nicely normally distributed in the grand scheme of things. Also with such a large sample size and with so much of our data being expressed nice, he knows box plots. I'm not worried about normality in this case at all, especially since we have such large sample size. Mos skipped these last two. Alright, so generally so far we're happy with our assumptions, but let's just see what happens if we log transform the data. And now we'll just look at this again. That's not nice. So when we log transformed our data, this is, this makes things much less happy. So we will not log transform our data. Now I have to remember to run all this again. Or we don't have to run the, we don't have to plot the data again, but we have to make sure that when we're looking at this object stored as the output from LM function. They want to make sure that we use, we have the output from the non transformed model that we were looking at. Just want to think of it. I want to point out that we could have actually explored the effect of log transform indirectly in boxplot, but could have said log and then just taking the log of our dependent, dependent variable right there. And you can see that when we log transform the data, the data are much less pretty. Whoops. I should say this is still a new computer and I'm still getting used to the keyboard. So that's a, that's a reason for some of the surprises that I experience when I'm making these, these live videos. So what we're learning here is that log transforming the data is not really something you want to do. All right. Did I say earlier 2t, I can't remember face submitted library car. I'll just do it again. Ok, that's fine. We have, alright. So now that we're satisfied that our data meet the assumptions, we can look at our p-values. I'll just point out that often when we do these analyses, I will say, want to look at the summary of this output. We're not gonna do that here. Enlarge part because when we specify contrasts in this way, as I've mentioned in a previous video, doing that changes the interpretation of the coefficients in, in the output for this model. And so we're not going to look at the coefficients in this, in this analysis for that reason. All right, back to looking at the p-values. To look for the p-values, we're going to use our ANOVA function. And our data or our output is in the object live dot lm. And we want to calculate p-values based on type three sums of squares. And here, so we get, okay, so what did we get here? Let's start at the bottom. First of all, we can see from this p value that we have very, we have no evidence whatsoever for an interaction between treatment and fertility. And that really matches what we saw when we looked closely at our box plot earlier. So this output really confirms our intuition from when we were looking at the data earlier. Because there's no interaction. That means, what that means is that the effective treatment does not. We have no evidence that the effect of treatment will depend on the level of fertility and we're looking at, and vice versa, effective fertility. We have no reason to think that the effect if the fertility treat that the effective of fertility will depend on which level of the median environment treatment we're looking at. So since effective treatment is consistent over levels of fertility and vice versa, what that means is we can interpret these p-values where, what, what's happening here? These results here allow us to test the effects of the treatment and fertility when we're kind of averaged over the effects of the other treatment. Sorry, treatment is a terrible word to use here. I'll try to say that again more, more carefully. So. This line of output allows us to make inferences about this factor treatment, which refers to the differences in mating environment when we've averaged over the effects if fertility. Similarly, this P-value here allows us to assess the, the effective, the fertility factor when we've accounted for any effective treatment, but averaging over those two difference scenarios, okay? Because you don't have an interaction, we can interpret these two lines of code very freely. Case, what did we learn from this? What we learn from this is that this first factor, treatments, which refers to the meeting environment, we have a very small p-value, which is what we inferred or which is what we might expect from looking at the size, the difference between the high cost made environments and low cost media environment when we're looking at box plots earlier. So this small p-value suggests that we have a lot of evidence are we we have a strong reason to believe that there is a difference between our media environments in terms their effect on life span. We also see that the p-value is less than 0.05 for our fertility treatment. So this suggests that female, that we have evidence that suggests that we have evidence to conclude that females from the US, that sterile females and fertile females do not live the same amount of time. This p-value suggests that sterile and fertile females live different amounts of time. So those are our big conclusions. It looks like in terms of our original hypothesis, our interaction here, or a lack of evidence for an interaction, suggests that our original hypothesis that SP decreases lifespan by increasing egg production or by causing females produce more eggs and that additional cost would cause females to live less long. We have no evidence to support that hypothesis, which is fun. But we do have evidence that egg production per se will influence lifespan. And that's because of this low p-value for fertility treatment. And we also have evidence. That mating per se affects female lifespan given the results from this treatment. So what we'd like to do now is we'd like to understand these outcomes a little bit more. So we're going to do some post hoc tests. Ok? But before we do that, just want to point out that we do want to report these results. These are the main results from the experiment. And we would want to report the p-values, the f values that are here, and the degrees of freedom. So the degrees of freedom is one for each of our, each of our various effects that you main effects and the interaction. And then we also want to report the, the degrees of freedom for our residuals. Ok. I'll present a slide at the end of this that illustrates a way to report these data. Right now we're gonna focus on doing some post-hoc tests to understand the differences between the various treatments and the size of those differences. So as always, we're going to use a library. Library. Em means this is annoying me. I put this there by accident earlier. Okay, so let's start by trying to understand the difference between the two mated environments. So let's try to understand the effect of fertility. I'm sorry that the effective treatment. So we're going to say, we'll save our output and something we'll call it live EM means. And we want to look at the effect of the media environment which is in treatment. So I'll just summarize that with TR T0 for treat. Let's use our IEP means function. And to do that, we need to first of all give the object where we store the output from the lm function, which we called Live dot lm. And now we want to specify which treatment we want the results for and we want the results for. So I should say we factor, I keep saying Treatment When I mean factor, it's really unfortunate that this column has been named, the column any output has been named treatment. It's because that keeps confusing me. And there we go. So we're getting this warning that the results may be misleading due to involvement in interactions. It's giving us this, it's giving us this warning because we have an interaction term in our model. But you can see both from our plots and from our results from my p values, that we have really no evidence for an interaction whatsoever. And so, given that we Both are plots. And our p-value suggests that there's really no reason to think that there's an interaction present. We can ignore the interaction and we don't have to worry about this warning that that EM means is giving us. Okay. So let's start now just by getting the values of our EM means. So you can see here that, so this column here gives us the marginal mean or the estimated marginal mean for females that experienced either the high-cost treatment or the low-cost treatments. So what this means is these are mean values that account for the level of our other factor fertility. And so these results are averaged appropriately over the effective fertility, which is what's given here. And so we have mean values and standard errors and confidence intervals. And we could report these values in our output, or we could report these values when we, when we read up our results if we wish. Okay? And what we can see here is pretty much what we expected from the plot we saw before. From our box plot. I mean, you can see that females that experienced the low cost environment lived about seven days longer than females. That made it a lot. Okay. Let's think about that in terms of an effect size. So females in the high cost treatments of the high cost, females that made it a lot. Lived 20, roughly 25 days. Where as females that experienced the low cost made Environment lived about 30 days. I'm shifting these numbers ever so slightly because it's easier to, in my mind to divide by five. And so that means that this experience, that experiencing a high cost median environment relative to low cost made environment decreased lifespan by a little bit more than a sixth. Because this value here of 24 is about 5-6, that of this value of 30. So a change in your life span of a sixth is pretty big if you thought of that in human terms. If the average human lifespan was 60 years, than a sixth of your life would be living only 50 years. Okay, so we've gotten those mean values. Let's now look at our effect size. We can, so I was just doing those calculations in my head there. Let's now formalize that by saying by using the Paris function. And so this will give us our contrast. So what we're doing here is we're estimating the difference between the average lifespan of females in the high-cost meeting environment compared to the low cost media environment. And you can see that the difference was equal to 6.98 days. So almost exactly seven days. And so we've already expressed that. We've already kind of tried to put that into real terms when we're discussing, just means a moment ago. And so I would say that this effect is really a large effect. So not only does mating environment have a significant effect in women thinking about our results and these traditional terms of Just something means significant or not. You can also see that the magnitude of the effect is really quite great. We can get a standard error for this difference. And both of these things are things we'd wants to present. If we were to write up these results. Okay. We'd want to do want to present the standard error and the standard error for the size, the difference, and the size of difference itself. There's really no point in reporting this P-value because this result here is kind of redundant. Because that result there, we already learned that result from this result up here. Okay, so those two results are redundant. And that's because we only have two levels within this factor called treatment. Okay, and now finally, let's get some confidence intervals for this difference. We'll just say that just get the confidence intervals directly. And here we go. So here's our estimate again. But now we have generated confidence intervals for that. Okay? Now we can do the same thing for looking at the effect of female sterility. So we'll say live. Em means we'll call it stereo. En means live dot. Lm. Whoops, forgot the dot lived on LM. And we want to look at the effect of fertility. Again, we get that warning. Again, we don't need to worry about it. Here are our, here are our estimated marginal means. And you can see here that the difference in life span between fertile and sterile females is relatively small. So in this case, we only have a difference of about two days. Whereas earlier we had a difference of about seven days. So this is an example of a case where we have a significant effect of the treatment of female sterility. So we have evidence that sterility can increase female lifespan. So there is something about being fertile that makes females live longer. Or at least there is something about sterilizing females that allow them to live longer. We do have decent evidence for that. But the size of that effect for increasing lifespan is much smaller than we saw relative to the effect of the mating environment. This is, I think it's a really good example of two different cases where we have formerly statistically significant results. But kind of the biological impact of those two different results is not the same. The biological impact of the media environment is much greater than of sterility. And that really emphasizes why it's so important to look at effect sizes and not just look at p values. Which is another reason why I really like this example here. We're just going to run the Paris function. So let's just run that. So here is our effect size for female sterility. So the difference in life span between sterile and fertile females is equal to 2.18 days. And here's our standard error. So we can again report these in our right up. And finally, let's get some confidence intervals for this. Okay? So here's our confidence intervals. This I think is a good place to end our discussion of these data. Because let's just look at, I've been emphasizing the importance of effect sizes. Let's look at the effect size is based on these confidence intervals. So the effect size of sterility and its impact on on lifespan, we can reasonably say lies within a range of about one to about three days. Okay. So we do have good evidence for there being a impact of sterility on female lifespan. But that effect is most likely within this general range. Given that we're looking at 95% confidence intervals. Let's compare that to what we got up above. I'm not going to search for it. We'll just enter this again. We can compare that result of sterility with what we get for. The median environments, you can see that the effect size for the media environment is likely to be at least twice as big as the effect of sterility. Because our 95% confidence interval for this effect size is somewhere between six days and eight days. So if we took the lower bound, so sorry. If we said that the effect size for fertility was 3.3 days, which is kind of at our, at our upper limit. That's listed here as a lower confidence limit. That's just because these confidence intervals are negative. And that's just because we've taken a big value and subtracting that from a little value. So if we say that the effect of sterility was 3.33.3 days, then that's still only about half of the lower level of our estimate for the conference. Or 3.3 is about half of the or a lower estimate of the effect size of the media environment. The point here is that both of our different factors, we have evidence that both Ready Front factors influence lifespan, but one of them seems have a much higher impact than the other. Ok, so we've interpreted our results to death at this point, I think. How would we write these up? We'll just end with that. Okay? We could report the results with something like this. We could create a figure like this one. And I won't read all this out. You can pause the video and do that yourself. But you can see in this blurb, I've tried to explain how to interpret box plots, etc. And here I've just tried to outline what we would generally like to report in a results section. First of all, he wants to tell the reader what kind of analysis we use. So we analyzed our data using a two factor general linear model, accounting for unbalanced data. In other words, saying we analyze the data assuming by by using type three sums of squares when calculating our p-values. Let me say the effects of fertility appears consistent across levels of treatment and vice versa. So that's our first main conclusion. And this is a results section, which is why I've not gone on to say how this influences our interpretation of our hypothesis. I've not gone on to say that this result here really means that we have no evidence that SP influences lifespan by influencing egg production. Okay, that's something you put in a discussion. But I've backed up the statement by name in the test. Stating we're talking about an interaction given our F value and giving our degrees of freedom for the interaction. So 1838. So giving the, both these degrees of freedom, the degrees of freedom for a particular model term and the residual degrees of freedom. And then giving the p-value. I follow this by saying on average females and the low cost treatment lived 6.98 days. I should say, lived on average. So I've I've not indicated this is mean that something should do. But we're going to leave it there. And then we provide a standard error for that difference. Okay. So I've just, I've just tried to reread what I wrote here. So the main point here, so same females on average. So on average females than a low cost treatment lived 6.98 days longer than females and the high cost treatment. Okay. I just realized I made a mistake, that I made a mistake earlier. This can happen when you're making a video. I do have my own average here. So this first statement here refers to the effect size of the mating environment. So we've said this is the difference in longevity due to the effect of media environment. Here's a standard error for that difference, and here is our confidence interval. And then afterwards, I've given the marginal mean and the standard errors for each of the media environments. So either in the low cost or high-cost media environments that I've referred to Figure one to depict the data. And then we've given the output from our general, our overall model. So we have a two-factor general linear model. This is referring to our treatment term. We report our F value with our degrees of freedom and we give our p-value. And then finally, it's a stairwell. Females lived 2.18 days longer than fertile females. Here I should say on average. Okay. So this again is referring to the effect size of the sterility treatment. And here is a standard error for that effect size and confidence intervals for that effect size. And then we have means and are marginal means and standard errors for those marginal means, we're referring the reader to the figure to see the data. And we report our F-value, our degrees of freedom, and our p-value. We have covered a lot in this video. I hope that this video has been useful and hopefully interesting, if not for the data analysis. I hope it's been useful for the data analysis and interesting, but also been interesting for the biology. And we'll stop it there. And I'll say, thank you very much.