Okay, In this video, we're going to talk about the second of three types of questionable research practices. We're going to talk about something called p-hacking. P-hacking involves a series of practices that all serve to change p values. Again, we're going to focus our discussion in video on the results of a survey that was published. And plus 12018 by Frazier at all. Where in this study, Frazier at all surveyed about 800 ecologist and evolutionary biologists, or 800 biologists total. And asked them about their experience at very fought with various forms of questionable research practices. So they asked them for questions about p-hacking, which I've listed here. These were questions five through eight in their survey. So let's just jump into them. So first, they asked whether the respondents had an experience with rounded off a p-value or some other quantity in order to meet some pre-specified threshold. So for example, round and a p-value of 0.054 down to 0.052. They asked the respondents if they had ever edited their dataset. So they asked if they had experience with deciding to exclude data points after first checking the impact on statistical significance or some other desired statistical threshold. Okay, so we have a situation where a biologist would analyze their data, check the p-value, and then afterwards decide to remove some data points. In the third case, we have the opposite scenario where instead of removing data, we're talking about adding data. So they asked whether or not they had experience with collecting more data for a study after first inspecting whether the results were statistically significant. Okay? So you have a biologist who's running a particular experiment. They collect the data as they go. They analyze the data base. At that point of analyzing it, they know whether or not they have a statistically significant effect. And after knowing that they continue to collect more data. I'm trying to say this in a way where I'm trying to not guess exactly what the researchers thinking was. And I'll just say that, okay, I'm just trying to take these questions absolutely. At face value. And their last question was to whether they want to know whether or not the respondents had experience with changing to another type of statistical analysis. After the initially chose an analysis failed to reach statistical significance or some other desired statistical threshold. So in this case they're talking about a biologist that would be, that would have collected their data, analyze their data, found that, that the first analysis that they had selected did not yield statistical significance. And after that discovery, they decide to try another form of analysis. So one of the respondents say, these are the results for all the questions that were asked. I've just plotted them or given the entire figure here, because it's the tidiest way of presenting the data for these four questions. And we're just right now going to focus on the results under these red arrows. So for these four questions here, and this plot indicates the proportion of the respondents for ecology and the dark blue and for evolution in a lighter blue, where the respondents had participated in these practices at least once. So this does not tell us how often the respondents might use these practices, but it tells us how many respondents had use these question research practices at least once. So for rounding p-values, we can see that ecologists release the, the ecologists that were surveyed. About a quarter of them are slightly more than a quarter. I've rounded down p-values, whereas the number is closer to around 20 percent for evolutionary biologists. It looks to me like about a quarter of the surveyed ecologist and evolutionary biologists have participated in this practice of excluding data after having first analyzed it. And so they've done, about a quarter of respondents have done that at least once. It seems like more ecologist and evolutionary biologists have gone to the practices of analyzing the data and then adding more data to their dataset. So it looks like around 35 percent. That's just my eyeballing this figure earthbound 35 percent of ecologists have done this at least once and around 50 percent of evolutionary biologists have done this at least once. And it looks like about 50% of both ecologist and evolutionary biologists have switched analyses after the first analysis that they chose, indicated that there is no statistical sorry, that there was no statistical significance for the effects they're investigating. Okay. So that's, those are the frequencies that we find of individuals that have perform these practices at least once. Among individuals that have practiced, these are that have use these practices at least once. How often do they do them? We're going to what the authors of or what Frazier at all did is they asked the respondents how often they partake in these practices. And they gave them a series of categories, either never, once, occasionally, frequently, or almost always. And so we can see that for this practice of rounding p-values, the most common category of frequency or frequency of use is this category of occasionally. So it seems that a lot among the evolutionary biologists and ecologists who do round p-values, it seems like the most common time category is to do this occasionally. Okay? Whereas their arse, and that looks like around 10 percent of ecologist and evolutionary biologists round p-values occasionally, a smaller percentage will almost always do that. And an intermediate percentage around p-values frequently. What about excluding data? It looks to me again, I'm just kind of eyeballing this. That about 10 percent of ecologist and evolutionary biologists have analyzed your data and then excluded some data. About 10 percent have done this once, where as about a similar percentage. So I'm going to guess around 10 percent of both ecologist and evolutionary biologists do this occasionally. Disturbingly, there is, there are some, but some biologists in both disciplines who almost always undertake this practice. They seem to be rare. What about adding data? Well, again, the occasionally group is most frequent. So it looks like about a quarter of ecologists, evolutionary biologists, a few more evolutionary biologists, will collect data, analyze it, and then continue to collect data after that first analysis. So it looks like a quarter or more of the respondents undertake that practice occasionally. Okay. We can see though, that a lot of them don't think this is a good thing to do, which is interesting. I'm, I'll remind you this light blue area tells us the proportion of the respondents that think that this practice is just fine. And so among the evolutionary biologists here who occasionally add data after first analyzing it, you can see that about half of those think that this is just fine. Whereas the remaining half. So I'm going to guess maybe 10 percent of evolutionary biologists will think that it's bad. To analyze their data and then to continue checking amended, and then to continue collecting data afterwards, despite the fact that they think it's bad. I'm sorry. That was terribly in Arctic in articulate. What I'm trying to say is it looks like around 10 percent of evolutionary biologists will analyze their data and then continue to collect data afterwards. Even though they think that that's a poor practice. What about switching analysis? Well, here we can see that ecologists, evolutionary biologists are pretty similar. It looks like, I don't know, maybe 7% of each have performs this practice once. Whereas it looks like over 25 percent have do this practice occasionally. But again, it looks like most of them think that this is not a wise thing to do, but they will do it anyways, occasionally. Okay. There are some individuals who are rare, but the almost always undertake this practice, okay, and this frequently category is in, in the middle. How does this compare to psychology? I should say that this type of comparison is based on these sorts of data. So data where respondents were asked whether or not they had taken part in these practices at least once. Okay. So these, this comparison does not compare in terms of the frequency of these practices. So we can see, I'm just going to tell you overall before looking at this, that the overall results from psychology and ecology and evolution are, are, are all pretty similar. So if ecologist and evolutionary biologists, it looks like between 35 and 50 percent have at least once collected more data after inspecting whether the results are statistically significant. Okay? So about 40 to 50 percent of evolutionary biologists and ecologists have done that at least once. And for psychology to separate studies, one of the US and one in Italy, suggests that number is a little over 50 percent. Okay? I'll point out that these numbers underneath give us the confidence intervals. So I'm just going to be pointing to the point estimates here. But if you want to know about the confidence intervals, then they're available to you here. It looks like around 20 to 27% of evolutionary biologists and ecologists have at least once rounded their p-values or some other quantity. And you get a very similar measure for the, for a psychologists from those two studies. What about deciding to exclude data? Again, we found that around 25 percent of evolutionary biologists and ecologists have done this at least once. Whereas in psychology, this number is a little bit higher, a little bit closer to 40%, okay, For, That's, that's, that's the mean, okay? Broadly speaking. The results for psychology and ecology and evolution are broadly similar, at least based on this measure. And one of the things that Frazier at all point out is that this should be alarming to ecologist and evolutionary biologists. Because some psychologists have gone to great lengths to assess how reproducible research and psychology is. And their work has shown that it's quite worrying. And that it may be that around 40 percent of sight of work in psychology is not reproducible. Given this overall match, Frazier at all point out that ecology and evolution should be thinking seriously about the reproducibility of work in their field. What are the consequences of these various forms of p-hacking? We'll just talk through them one at a time. What about this? What are the consequences of rounding p-values or other quantities? Well, this is a direct manipulation of your p-value. And it's a direct manipulation that biases the evidence for an effect towards demonstrating the effects with this is going to do, this is going to increase Type I error rate. In other words, this practice is going to increase the frequency of results that we find published that are actually false positives as opposed to being true results. What about the practice of excluding data? It's the exact same consequence. Justice in the first case, this is also a direct manipulation of your dataset in this case, in a way that presumably will tends to decrease the p-value. I'm saying presumably because just based on how the paper is written, I I'm I'm assuming that the that the way in which data were removed was to lead to smaller p-values. That's my reading of the paper. If you want to check your, your, your own interpretation of the paper, then I suggest you read it for yourselves. Okay. It's open access. Everyone can get it. So I'm just pointing at my own interpretation from what I read. And this next, what about this third practice of collecting more data for study after first inspecting whether the results are statistically significant. This is also going to bias results towards demonstrating effects. This is also going to tend to increase a type one error rate. And we're going to discuss this more in the next video. And our last practice will also increase a type one error rate. In other words, this last practice will also tends to increase the frequency of false positives in the literature. So just to remind you, this is the practice of. Changing to another type of statistical analysis. After the first analysis that was used failed to show statistical significance for an effects that they're interested in. Okay. Let's talk about why this will increase a type one error rate. Let's talk about a little bit more detail. What I'd like to point out is that different types of statistical analyses actually represent different types or they actually represent different statistical hypotheses. For example, let's imagine that our goal was, let's imagine that our biological goal was to compare the mean value for two different groups. Ok. I very common way of doing that would be to use a t-test. Where a t-test we'll use the original raw data. And will the mechanics, the t-test, will be such that it will directly compare the mean value of one group to the mean value of another, okay, when it's calculating the test statistic, a part of that calculation involves actually taking the mean of one group and subtracting from that the mean of the other. So that's a t-test has certain statistical mechanics built into it. Often, people might use a Mann-Whitney U test as an alternative for testing whether or not the mean values of two groups are different. So that's what they'll often, people may often choose a Mann-Whitney U test in order to test the same biological hypothesis. But by using a Mann-Whitney U test, they're actually testing a different statistical hypothesis. And that's because a Mann-Whitney-U test works differently. That a t-test, Mann-Whitney U test doesn't work with the original raw data points. It actually works with ranks that are calculated from those original data points. Add Mann-Whitney-U test. Don't directly compare the median value for two different groups. Mann-whitney-u test, actually test the overall, what to actually test whether the overall distributions are different for the two different groups, which is a different question from directly comparing between the two means. Okay? So what I'm trying to point out is, although a biologist might be trying to, or they make, their goal, might be to address the same biological hypothesis using two different tests. The actual statistical hypothesis that they're testing is different. And that will often be true when a, when a biologist decides a switch from one type of analysis to another. And so when a biologist as moving from one type of analysis to another, they are actually increasing the number of statistical hypotheses that they are testing. And we know that the more hypotheses that we test, the more opportunity we have to detect a false positive, okay, or for us to find a false positive. And so the more tests that we perform, the more opportunity there is for us to commit a type one error. Okay, So that's, that's the main reason for why switching from what analysis to another can increase our Type I error rate. Even if our biological hypothesis that we're, that we're looking to address is the same for those two different sets of analyses. This last point really raises an important question, which is, is it ever okay to switch analyses and the answer is absolutely. The important issue at play when deciding whether or not to change it. Analysis is what's your motivation for change for changing the analysis? Or you changed it because you got a result that you weren't happy with? Or are you changing it because you think another type of analysis might be more appropriate To answer your particular question. So it is okay, in fact, that's a good thing to change analyses when an alternative model or an alternative approach to analyze new data is more approach is a more appropriate way of analyzing your data given the biological question. So a simple example would be when you're deciding whether or not to transform your data. Okay? What I tell my students and what you will see in future videos when we talk about data transformation is a data transformation is fine to do. It's a good thing to do if it helps your data meet the assumptions of the analysis that you wish to perform. And one of the things that I suggest to my students is to not look at the results before deciding whether or not you want to transform the data. Instead, I suggest that the students are not just students that everyone, including myself, only assess the assumptions when they're analyze their data. Sorry, I'm getting ahead of myself. Here's my general recommendation. If someone's considering whether or not they should transform their data, what they should do is run a statistical model, check the assumptions of that model. And then just based on those assumptions are just based on the outcome of those checks of the assumptions. Decide whether or not you are going to transform your data. I suggest to not look at your results. Simply is a way to help keep yourself honest so that your unbiased when deciding whether or not to transform your data. And I will give the same general recommendation. When deciding more generally about whether or not a particular analysis is appropriate. Um, it might be that data transformation simply isn't a good enough solution to help your data meets your assumptions. You might want to try another type of analysis. In that case, it might be wise, if possible, to not look at your results before making that type of decision, but instead just look at your assumptions. Alternatively, it might be that while you are doing your experiment or while you are performing an initial analysis of your data. It could be that a new type of analysis was invented because that happens all the time. And it could be that that new type of analysis might be exactly what you need to perform a more appropriate analysis of your data. In that case, change your, change your form of your model, okay? Use a new model type that's a good thing to do. Alternatively, it might simply be that you were unaware of a type of analysis that was available. And when speaking to a colleague, they might point out, Hey, this other type of model might be more suitable. And in that case, That's another great reason to change the model that you're using or change the way in which you're analyzing your data. Oh, it is okay to change your analyses as long as you doing it for the right reason. So I'm just going to get a pre summary at this point. The main message, what we talked about so far is that p-hacking comes in a variety of forms. And overall, they tend to increase the rate of false positives that are found published in literature. So I didn't finish my sentence. What I'd like to do now is talk about these last two forms of of questionable research practices which are kinda stand apart from most of these other question research practices. This one here is undisclosed problems in the last issue is fabrication. And so I just want to quickly read the question that was asked to the respondents in the survey. Want to read it to you? So have this on another computer beside me. So for this undisclosed problems, the participants were asked whether or not they had ever practiced in not disclosing known problems in the Methods and Analysis or problems with the data quality with that could potentially impact conclusions. Okay, so that's we're talking about here. And it looks like roughly 20 to 25 percent of ecologist and evolutionary biologists have at least once failed to disclose issues with the methods or the analysis, or that had to do with the quality of the data that could impact the interpretation. What about fabrication? Fabrication refers to filling in that missing data points without identifying those data as simulate it. Okay? In other words, including data. In an analysis that did not seem to come from this, that were not generated in the way that the authors suggested that would have been made. Kss, we're talking with fabrication. How often do these occur? Well, among individuals who have who have failed to disclose some problems, it looks like about, I don't know, 10 to 15 percent of ecologists, evolutionary biologists do this occasionally. There are some rare individuals who made you this almost always or frequently. Okay. It seems like most individuals that do this, which was about 15 percent of the respondents or 10 percent, will fail to disclose problems occasionally, fabrication should be pretty clear that that's not something that we should be doing. And thankfully, those percentages are all really low. But disturbingly, you can see that there's a non-zero number for the evolutionary biologist to say they almost always do this. That makes me laugh. But that's not funny. That's, that's, that's pretty terrible. Luckily, that is a very rare occurrence, at least according to this survey. Okay. But apparently it does happen. On that note, I want to show you something that appeared in my email this week. Just to end this video, I get Table of Contents alerts for a number of journals. And one of the journals that I get a table of contents alerts for is a journal American naturalist, which is an excellent journal for ecology and evolution, is considered one of the best journals in that field. And it was announced in this month's table of contents that there was a retraction of a paper that was published in 2012. The paper was called iterative evolution of increased behavioral variation, characterizes the transition to sociality in spiders and inspite as proves advantageous. Exactly what that means doesn't matter here. What matters here is why the paper was retracted. So I'll just highlight the relevant text here and I'll, and I'll read it to you or you can just pause the video and read it for yourself. They say that investigation by the journal, which is fully examined by all authors, revealed to duplicate ratios of spider mass to pray mass, which is often an exact value. For example, 0.3 folds up by 0.330 before exactly. So they note this was an anomaly and this a similar anomalies with redundant predator prey mass ratios were also found for other sized categories. These anomalies in data collection by Jonathan and Prewitt could not be explained, okay, undermine the credibility of the articles results. And so on this basis, the authors of this statement regretfully retract the article. The punchline here is basically there was evidence that the data were fabricated in this paper. And so the authors of the paper have decided to to remove it. I'm not really sure what else to say. There are other hand to say, this happens. It should not happen. But it does happen. Luckily, it doesn't seem to happen very often. But it does happen. On that note. I'm going to stop the video. I guess I hope that this has been used for at least interesting. And it'll say, thank you very much.