Okay, In this video, we're going to discuss the first of three different categories of questionable research practices. Cherry picking. Want to say cherry picking? I'm not talking about what we see in this nice picture where you outside on a beautiful summer's day, collecting beautiful delicious fruit. What I'm talking about is a process of doing science where we're selective about which results we choose to report. In particular, cherry picking refers to the practice of failing two reports, the results for variables or conditions or treatments that relat, that had relatively high p-values. And instead, I'm tending to only report results for variables, conditions, or treatments that are associated with low p-values. In other words, results that are traditionally called statistically significant. Before we go on, I want you to just stop and ask yourself for a moment what this practice of biasing towards reporting results that have low p-values or a statistically significant effects. What consequence do you think that practice would have for interpreting the results of a particular paper and for the scientific community in general. Just stop and think about that for a second. Okay, I don't want to depress you too much. So I'll just say we're going to return to that topic at the end of the video, okay? But I think it's important to just let that sink in before we go on and talk about this any further. It should be self-evident that if as scientists were trying to be objective about our studies, it should be self-evident that any biases in terms of what we report to other scientists would be problematic. So we'd like to know how often is cherry picking occur. Frazier at all performed a study to quantify exactly that question. And they focused on ecologists and evolutionary biologists. They contacted over 800 ecologist and evolutionary biologists and ask them to complete survey. Where they asked him ten questions that all pertain to questionable research practices. And in this video and the two videos that follow, we're going to talk about the results from that survey. This type of survey has been performed in other areas of science as well. So for example, in psychology, I'm focusing on the results from this paper. Because this paper consolidates their findings with comparable results in psychology. And so I thought that it's nice to be able to make that direct comparison. So I wanted to keep things simple and focus on on a few papers as possible. And so I've chosen to focus on this one just because it makes that connection to other research areas. I should say this paper came out in 2018 and is published in PLOS One. It's a nice paper. Go read it. So they asked three questions to these ecologist and evolutionary biologists that pertain to cherry picking. The first of all, so these questions, these three questions pertain to questions 124 in their survey. And I'll just walk through what these questions are now. The first question asks the respondents about not reporting studies are variables that failed to reach statistical significance. So there's haggling P being less than 0.05. Apologies that my less than sign seems to have disappeared. Or some other disease statistical thresholds. So P being less than, say 0.01 for example. Okay? See what they're talking about here in this first question is cases where let's imagine a case to explain what they mean here. Let's imagine that some scientists were studying humans and they were measuring, I don't know, Hajj, diameter, hand grip strength, add heart rate. Okay. And let's imagine they wanted to perform a separate analysis for head size, hand strength, and heart rate. And let's imagine that hand, your head size and hand strength gave small p-values, whereas heart rate did not. What they're talking about here for this first case is the practice of excluding the results for heart rate. Just because it tended to be associated with a higher P value. Whereas instead they would just report the results for hand strength and head size. The second case, as I understand it, is talking about something a little different, where they're talking about not reporting covariates in an analysis that fail to reach statistical significance. What they're talking about here is, or at least as I understand it, talking about situations where an analysis might include the effects of several different independent variables on say, one dependent variable. So let's imagine that we wanted to look at the effect of sex and body size on, let's say heart rate. Okay? And it's possible to perform analyses that would, or that can look at the effect of sex and bottom body size simultaneously. So in one analysis on heart rate, okay? And what they're talking about here or at least my understanding is. That what they're talking about here is cases where a model might include multiple independent variables like sex and body size, but then only reporting the results for the model terms that had significant p-values. So imagine that only sex had a small p-value associated with it, and body size did not. Then it believe what they're talking about here is cases where the scientists would fail to report the results of body size because it had a large p-value. To understand why that might be relevant, we should put ourselves in the minds of the scientists, where typically someone who's analyzing data appropriately would only be choosing variables to include in their model that they might actually think are relevant for the study that they're doing. So in other words, if someone included body size in an analysis of heart rate, than they would presumably only do that in the first place because they would think that body size could be a relevant independent variable. And so if they fail to reports the effects of body size than what they're doing is they're leaving out. Results that they themselves, as biologists thought, could have been important in the first place. Okay, That's the kind of situation we're getting at here. The third situation that we're considering is a case where scientists might report a set of statistical models as the complete tested set when other candidate models were also tested. So what do they mean here? What my understanding is that we're talking about is where a scientist mighty performed a series of analyses that are all relevant to understanding a particular phenomenon. But then they would only report a subset of all of those analyses. Even though they might have thought a priori that a number of those analyses could also give insight. Okay. And so we're talking about cases here, or at least my understanding is they're talking about cases where the researchers chose only to report a subset of all of the analyses that they performed. Okay? So that's what we're talking about here. Okay? So what we'd like to know now is how frequently scientists might partake in these various questionable research practices. And I'm going to show you a series of figures that all come from this Frazier at all paper. This first figure here pertains to cases where the scientists report having done this questionnaire research practice at least once. Okay? So what you can see here is that we have four different practices. So this first set represents the first question. So I've not Reporting analyses a particular variables. So like leaving out analysis of heart rate in our example earlier. The second prefers to leaving out covariates that had high p-values associated with them. And this last set here, this refers to not reporting all the analyses that were actually performed. Okay? This third aspect here refers to something called harking, which we're actually going to refer to in our third video. I just chose not to block it out in this figure because I just thought that would look too ugly. And I thought that it would be better if I just told you that when we're not going to talk about these harking results now, we're going to talk about them in the third video. What I want you to see is that among the people who responded to these questions, for all three of these cherry picking exercises, about 50 percent of the respondents have or report having performed these questionable research practices at least once. Which is worrying. Okay. Any more worrying if that people who perform these procure, if the people who have performed these practices at least once, do them all the time. So what we'd like to now to know now is that among the people who have done this at least once, how often do they do it? And so that's what's reported next. So I have results for our three different forms of cherry picking. First of all, only are we leaving out variables that gave high p-values, leaving out covariates that had high p-values and also not reporting all of the models that were run. Okay? And what do I want you to see here? So what I want you to see what it should do first is orient you to these figures. What freezed all I've done is they've reported the fraction of respondents that describe their behavior is occurring either never, once, occasionally, frequently, or almost always. Okay. And what you can see here is that most of the respondents that said that they have performs this. So I should say that the people who said never, they represents this empty space in this figure on the left, okay, say they represents the people. Notch perform these practices at least once, okay, that they're represented here among the people at how perform these practices at least once. They're represented in these various bar charts. Once, occasionally, frequently, an almost always. And so among the individuals that have performed these practices at least once, we see that they would describe themselves as performing. These are the most individuals perform these practices at least once. Say they do so occasionally. A smaller fraction. So about 10 percent of the respondents say that they do partake in these practices frequently. So maybe 10 percent of the respondents around that number say that they partake and these question Rhys research practices frequently. That's that's worrying. Okay. That about 10 percent of the respondents to the survey would perform these practices relatively frequently. It's also worrying that about 40 to 50 percent of the respondents say that they perform these practices occasionally. Which would mean that once you have mean that occasionally we should worry about these practices. Skewing our ability to interpret the results from, from studies that were reading. Okay, There's one last aspect of these results that I want to point to, and that is the meaning of this light blue that I'm pointing to here. Where you see this color of the light blue. This refers to the fraction of the respondents that actually think that these practices are just fine. Okay? So there's a couple of things to see here. First of all, you can see that individuals that only partake in these practices occasionally. You can see they probably only do this occasionally because most of the time they think that it's wrong. Okay. So for the individuals that leave out analyses of certain variables, most of those individuals who do that also think that it's a bad thing to do. It's also interesting to note that down here for individuals that do not report all of the models they perform, the individuals that do that frequently also tends to believe it that practices fine. Which probably explains why they partake in this practice frequently. It may not be good, but it's logical. Okay. So that's our picture of how frequently these practices occur. At least according to these respondents, among ecologists, evolutionary biologists. How does this compare to psychology? I should go back. We can compare some of these results to psychology where specifically we can compare this result here, okay, Where into other studies. Psychologists were asked how often. Sorry, not how often. Psychologists were asked whether they had ever failed to report the results of a variable because a variable had a large p-value associated with it. So where we're talking about this kind of result here, the Have you ever kind of question, specifically for this first topic of not reporting all the variables you might have investigated. And what we can see from these two other studies in psychology, one performed in Italy and when performed the United States. You can see that the fraction of respondents that it performs this practice at least once, is pretty similar in psychology versus ecology and evolution. And that's, that's not a strong Paris. And as we would like, between psychology and ecology and evolution, because we're really just comparing one type of data, this type of data on the left, for one type of questionable research practice. So we were far from a complete comparison between these two, between these two disciplines. But these results we see here should be worrying. And that's because as I mentioned in a previous video, there have been some really exemplary psychologists who have gone to great lengths to try to understand the extent to which psychological research is not reproducible. And they found that about 40 percent are only about 40 percent of the psychology experiments that were examined were reproducible. So a large fraction, there's worry that a large fraction of the psychology literature might have IR, IR reproducible results. So if fast we find for psychology and we have at least a match between psychology and evolutionary ecology for at least one measure where we can make this comparison. That's worrying for ecology and evolution. I should say that, that, that red flag that I'm waving around here right now. That's not my insight. That's one of the insights at Frazier at all make in their paper. So again, I encourage you to go and read that paper to see the kinds of insights that they're able to make. So what are the consequences of cherry picking? This is the thing that I asked you to dwell upon the very start of this video. Well, overall, what cherry picking does is it means that readers are not getting the complete picture for the research that was performed. And not having a complete picture is going to have a number of important consequences. The first is by not knowing the complete picture, this impedes interpretation. Okay, basically readers lack all of the information that they should have access to. And without having access to all the information they should be able to see. They're not able to interpret the results fully. Presumably, the researchers that performs the studies are performed the analyses oring that included covariates in analysis that they failed to report results for. Presumably they did those experiments or they did, they included those covariates, or they performs those models in the first place. Presumably they did all those things the first place because they thought that they would be relevant. Okay. They thought that they would be important to understanding the results. And so given that, that probably is true, it's probably important that the readers actually see all of those results that the original, or see all the results that the researchers produced in the first place. Just so that they can get a full sense of what the full picture of the data is. Okay? One obvious component here is just a comment that nonsignificant results are important to. This should be particularly clear when we think about effect sizes. In previous videos, I've pointed out arguments that other people have made that just because a result is not statistically significant, does not mean that it's not important. And we can, we can come to that kind of conclusion by focusing on effect sizes. Instead of focusing on p-values. Cherry picking takes us one step in a more terrible direction by not reporting the p-value or the effect sizes in the first place. And so by leaving out all that information, we can't even tell whether an effect size might have been important in the first place. Okay. I feel like I'm droning on a little bit too much here. The main point here is that when we don't report all of our results, readers cannot properly interpret the results. They cannot understand the results in, to the extent that they should be able to. So as a result, cherry pick and should not occur. This creates a situation that's not only relevant for understanding the focal paper that we might be reading, but also for understanding the literature as a whole. And that's most, most obvious when we're considering meta-analyses. Meta-analyses perform a crucial function in science by synthesized and the results of multiple studies to be able to give us some overall sense of the overall picture for an area of study. The meta-analyses are only really useful. Or meaningful, I should say, when the data that they're drawing from our unbiased cherry picking obviously biases the information that exists in the literature. So when we cherry pick, we diminish our ability to understand a whole research area, which is obviously bad. I'm going to lump these last two comments together about this practice being unethical and leading to redundant investigation. I'm just going to say that biasing which results we choose to present, it should be self-evidence with that. So unethical. But there are also unethical downstream consequences which had just want to elaborate on. When someone's reading a paper. If they're reading with a keen eye, they might notice. That's the researchers do not comment on some particular aspect of the study. And that might lead the reader to say, the authors didn't consider this. I'm now going to go and look at this thing that they have not commented on. And as a result, I'm going to take research money. So money that came from the public him was given to me most often through government funding. I'm going to take that public money and I'm going to use it to investigate this thing that the authors didn't comment on. If I'm performing an animal research, I'm going to use animals lives to investigate this phenomenon. That's the authors have not commented on it. Okay. That's what can happen when someone's reading the paper and that, and that's a good thing. Okay, we want authors or we want readers to be getting new ideas for things they can look at, or things that can study when they're reading someone else's work. It's a big problem. However. If people are reading the papers, decide to start a new study without knowing that the authors have already looked at that aspect. Because if they had known what the original authors had already found, then that might have dissuaded them from actually taking this next step. Okay? And as a result, when authors leave out some of their results, this can lead to redundant investigation, which has ethical consequences in terms of bomb particular just using resources including animals, lives that should not have been used. And on that sad note, I'm going to wrap up this video and say that's all I'm going to say about cherry picking. In the next video, we're going to talk about our next type of questionable research practice. I hope this video has been helpful. I'll say, thank you very much.