Okay, In this video, we're going to talk about our third, a questionable research practice. We're going to talk about something called harking. Hurricane is an acronym which stands for hypothesizing after results are known. Harking also comes by a number of other names, including forking paths and researches degrees of freedom. Now that we know the names of the things that we're going to be investigating. What does hearken actually mean? Frazier it all in there 2018 paper and plus one, they define harking as I've listed here, where they draw their definition or this definition from two other papers that I'll highlight in just a moment. So they, they explained that harking includes two very related practices. Want the first is presenting ad hoc and or unexpected findings as if they had been predicted all along. Okay, That's, that's the first way of thinking about harking. The second way is to present exploratory work as though it was confirmatory hypothesis testing. So these two definitions came from these papers here. Highly recommend these papers. They're very readable and they are a different type of paper to read, I think, at least compared to the types of papers I'm used to reading. Because they, they make interesting speculations about the nature of science, about the nature of science and why scientists behave in the way they do. So for example, the second paper here talks a lot about the role of confirmation bias and its role in the process of science and why we need to avoid it as much as possible. So I highly recommend these papers. Before we go any further, I want to talk about the difference between exploratory and confirmatory research. Because these two terms were included in our definition of harking. The main distinction between exploratory research and confirmatory research comes down to whether or not an analysis has been planned in advance. So an exploratory research, the steps for an analysis of a dataset have not been specified in advance. In contrast, with confirmatory research, there's already an analysis plan in place, likely even before the data were collected. So for example, with confirmatory research, we could imagine a scientist having a particular hypothesis in mind. And they might say, well, to test this hypothesis best, I should run this particular experiment and analyze it in this way. So if confirmatory research has already a sense of how to analyze the data before, before looking at the data and potentially even before design the experiment. Despite these differences, I really want to point out that exploratory research and, and confirmatory research are essential for scientific progress. And that's because they have very complimentary roles. We'll talk about exploratory research. First. I kind of think about exploratory research as an opportunity to kind of get to know your biological system. So you can imagine exploratory research occurring where someone might want to understand a general aspect about a system. Hey, something happened, you have maybe sexual selection. And they might collect a whole bunch of data that is all generally relevant to, to a general, a general set of hypotheses that they're interested in. And with that dataset and hand, what people might do is they might look for correlations. For example, among the various things they've measured. If they've measured a number of dependent variables. So for example, let's imagine that the someone's performing a physiology study, they might measure many different aspects of physiology that are all associated with particular system that they're in, that they're studying. And they might ask whether or not those various variables they've measured all responds to particular treatments that they've, that they've implemented, or whether or not to some of the variables they've measured response to the treatments. An exploratory analysis, you might want to know whether there's anything unusual or surprising in your dataset. Vary. Pretty much in general, what you're doing with an exploratory analysis is you're just trying to get a general sense of what's happening. I think the most or the greatest benefit of exploratory research is that it's through this process that we can generate new hypotheses, okay, It's by exploring a Sistema and getting a general sense of how that system appears to work, that we end up getting new ideas about, about how that system might work. And you get new ideas of aspects of the system that we can test. So that's exploratory research. Confirmatory research is very different. Where I've already given an example of confirmatory research on the previous slide where I said, you can imagine a situation where a scientist may have a particular hypothesis in mind that they want to test. And they might imagine that the best way to test that hypothesis is with a particular experiment or a number of complimentary experiments that can be that can all work together. And so with this hypothesis in mind, the scientist will develop these particular experiments. And often when they develop their experiments, they'll already have in mind the best way to analyze those data because the nature of the experiment will determine the kinds of analyses that you can use. So with confirmatory research, a scientist will have a focused hypothesis that they aim to test. And they will obtain a particular dataset in order to test that particular focal hypothesis in a predetermined way. Okay? So hopefully I'm creating a picture here of how these types of analyses are complimentary. With exploratory research, we have an opportunity to generate hypotheses. But with confirmatory research, this is the best way to actually test the hypotheses that were generated by exploratory research. So both of these approaches are essential and the complementary the problem or there is a problem however, that can arise. Where the problems can arise when an exploratory analysis is explained, when it's published as if it had been a confirmatory hypothesis or a confirmatory analysis. And we'll explain more about why that's a problem in just a moment. Okay? Before we do that, I want to point out that the published literature will include whole spectrum of studies that range from being entirely exploratory. Two studies that are entirely confirmatory. But there are many, many studies that will include aspects of both of these. You might have some studies that are larger confirmatory that, but that's still might include some aspects that arose from exploratory research. Or you might have the reciprocal of that, where you might have to work. It's large the exploratory, but it might have prompted the scientist to go on and perform a new experiment as confirmatory analysis of say one aspect or say have one hypothesis that might have been generated from the previous exploratory work. The point is that within the literature we can find or we expect to find a full range along this spectrum from purely exploratory to purely confirmatory. And all of this research is useful. That just useful in different ways. So why is harking a problem? Why is it a problem to perform? An exploratory analysis and then to present the results from an exploratory analysis as if they had resulted from a confirmatory analysis. I'm going to highlight two problems that are not mutually exclusive. I have one problem going to highlight on this slide, another problem on the next slide. Okay? So I've already highlighted that well-performing and blow oratory analysis. As scientists, we are interacting with the data. When we do that. So for example, we might perform one particular analysis which might reveal something surprising. And that might lead us to think, oh, well, if that's true, I wonder if this is true. And then you might go on to test that by looking at some other aspect of the data. Okay? So during the exploratory analysis, our interaction with the data will lead us to generate hypotheses. And I've already said that's useful. It's not only useful, it's essential, we need it for science. But what I want to point out though, is that by, by its very nature, an exploratory analysis is going to generate many more hypotheses, then we would likely have if we're performing a confirmatory analysis. Because remember, confirmatory it with a confirmatory analysis, we go in with a particular hypothesis or small family of hypotheses in mind that we aim to understand and aim to test with an exploratory analysis by analyzing the data and seeing something unexpected that can lead us down the rabbit hole and to explore other aspects of the data and to generate new hypotheses. As I said, that in itself is absolutely fine and that's a good thing. But what we want to appreciate is that by its very nature and exploratory analysis will involve testing many more hypotheses. Maybe shouldn't say many more. But it will tends to involve testing more hypotheses than a confirmatory analysis would. Why is that a problem? So why is it a problem to test more hypotheses? Well, we've mentioned in previous videos that the more hypotheses you test, the more opportunity there is to obtain a false positive or to commit a type I error. Okay? So what this means is that exploratory analyses by their very nature, are going to be more prone to. False positives are type one errors, then a confirmatory analysis would be. And so if we present an exploratory analysis or EDA exploratory results as if they were confirmatory. What we're doing is we're hiding the fact that we actually performed multiple. So I were hiding the fact that we actually tested. Multiple hypotheses and as a result, more hiding the risks that come with analyzing multiple hypotheses. And as a result, we're misleading the reader about how reliable our results are likely to be. That's, that's really the main problem. Where if we go through this process of exploring the data, opening ourselves to more false positives and then presenting them as if being confirmatory. The reader will be interpreting these results as being confirmatory without realizing that they should potentially be more skeptical and aware of the fact that these results are actually more prone to false positives than a true confirmatory analysis would actually be. What's our second problem? The second problem is a little more philosophical. So actually, I enjoy thinking about this one. I want you to appreciate that when we look at a dataset as we're interacting with the data set, we use the patterns that we find in order to generate hypotheses. That's a points we already made on the previous slide. A new, there's a new point that I want you to appreciate about the fact that we generate hypotheses from an exploratory analysis of a particular dataset. That point that I want you to appreciate is that once we generated a set of hypotheses from a particular data set, that dataset is no longer appropriate to test those very hypotheses that date that that dataset generated. Because if you were to use a dataset to generate a hypothesis and then use that same dataset to test that hypothesis that was generated. That logic is circular reasoning. And it's not appropriate. It's not an appropriate way to go about testing hypotheses. What we actually really need to do is to generate a fresh data set in order to test the hypotheses that were generated through the exploratory analysis. Okay, So the problem that I want to highlight on this slide is that if we partake in this form of circular reasoning, where we explore data set and that presents us with a new hypothesis. And then we in essence, use that, that same dataset in order to prove that hypothesis. So if we partake in that circular reasoning, then this process is going to inflate type one errors. In other words, this process is going to tend to increase the rate of false positives. And it also means that the p-values cannot be trusted through this kind of process. Okay? So those are the two main problems that I want to highlight that come from harking. This paper by occur in 998 is a really interesting paper. And curr, things are discussed as a number of kind of philosophic kind of philosophy of science type problems that arise as well that are, that are fun to think about. So I encourage you to go and read this paper to kinda complements the issues that we've already talked about in this video. What I want to address next is how common hurricane is. So this same paper by Kar in 998 highlight some results that arose from a survey of social scientists. It wasn't many social scientists, if I remember right, it was about a 150 social scientists. And they were asked about their impressions for various types of approaches to science. What we have here in white is the classic HD approach. Where this is the approach where that describe as purely confirmatory research. Where we have a particular hypothesis in mind, a priori. We design an experiment in order to test that focal hypothesis. And then we analyze our data and report our data all in a way that's consistent with testing that original hypothesis. That's what's being described here in white. These other three terms here, they all describe various flavors of harking that cur considers, okay? The exact flavor. It doesn't really matter though for these results, because I want you to notice that the overall results are all pretty similar. I'm at least we're going to talk about. So from this survey as social scientists, they were asked how frequently these scientists observe these various forms of hypothesis testing. And you can see that they, they often, they often observe various forms of harking about 40 percent of the time. Okay. They're also asked how often they suspect these forms of harking would occur. And they're also in a similar ballpark, say around 50 percent or so. Okay. So the point here is that harking according to this survey, does not seem to be uncommon. Next we're going to turn to the results of a survey that were described in the Frazier but all paper in 2018 and plus one where they surveyed about 800 ecologist and evolutionary biologists to get a sense of how frequent questionable research practices occur in the fields of ecology and evolution. And the third question that they asked is this one here, and it relates to harking, where they asked the respondents about their experience with reporting an unexpected finding. They asked the respondents about their experience with reporting an unexpected finding a or a result from an exploratory analysis as having been predicted from the start. And here's what we find. So here are all the results from their 10 questions. We're just going to focus on this set of columns here. The dark blue column here refers to data from ecologists. The lighter blue refers to data from evolutionary biologists. But you can see the results are pretty similar. And these results indicate how often evolutionary biologists have partaken in harking at least once. And it looks like around 50 percent of ecologist and evolutionary biologists have partaken in harking at least once. Knowing that people have done it at least once though, isn't that informative? Because it doesn't give us a sense of how frequently harking is practiced by the individuals that have done it at least once. And so we get more insight from this figure here where we can find the percentage of the respondents did have either never heart at once. Do it occasionally freak when the or almost always. In which you can see here is that it looks like about 35%. I would say that's kind of my guests from this around 35 percent, I'll say I've ecologist and evolutionary biologists say they occasionally HARQ. Okay. It's good to that about half of ecologist and evolutionary biologists claim to never HARQ. About 10 percent me, you've done it once. There is a small, small handful that say they almost always do this. Okay? What, there's one other observation like to point out here, and that is if we look at the light blue shading, this light blue shading that I'm pointing here, this refers to the proportion of individuals that actually think it's okay to hark. And you can see that among the people who, who hark or who hypothesise after knowing their results. Most of them hark occasionally, even though they think it's not a good thing to do. And we can speculate about why they might do that. Most of what's, most of what I've come across in the literature in terms of why people think hearken occurs is that it's largely related to pressures related to science and keeping your job. Okay. What I'd like to talk about next is, why is harking a relatively common? I have virtually no data to support the arguments that I'm going to make here. So take what I'm going to say in this slide. With buckets of salt, I was going to say handfuls of salts, maybe buckets assault. Okay. So this one I'm going to say next is highly speculative. But I've pulled out these paragraphs from the start of the curve, 998 paper that I showed you earlier. Because of that paper starts out with a historical perspective, trying to understand when hurricane may have begun and maybe why. I want to highlight kind of what curr is. Curr is pointed out here. Okay. I'm just going to read these. If you find it really fie, fie my reading annoying, then just pause the video and read it for yourself and then you can skip past this bit. Okay? So curse says, practically every modern textbook on scientific research methods teaches some version of logical empiricism, hypothetico deductive approach, okay? Particularly for empirical research. That approach prescribes to do scene or deriving one or more USPS. Sorry, I put the wrong emphasis there. That approach prescribes to do scene or deriving one or more explicit and testable hypotheses for some plausible theories about the phenomenon of interest prior to designing one's be search. The whole process of research design is fundamentally guided and constrained by these hypotheses. What one choose to look at and how one chooses to look depend vitally upon the questions one chooses to pose. So what curr is pointing out is that at that time, remember this was in 998. He claims that every modern textbook is going to talk about this confirmatory approach to, to science, where we start out with a particular hypothesis in mind. We design an experiment or a set of complimentary experiments that are all focused on testing that particular hypothesis. And then we would analyze a data from those folk focused experiments too. To test our focused hypothesis. That's what's being proposed or that's what's being said here. Okay. Curr goes on to say that there's been a change. However, in terms of the messages that we can find in more recent textbooks, excuse me. And when I say recent, remember this was written in 1998. So curd goes on to say, however, in the last few years, several textbooks and articles have suggested a departure from this traditional approach. One of the earliest and probably the clear statement of this non traditional approach can be found in Bem 1987. Where here's somebody has been pulled out of them. 1987. Where Bem says, there are two possible articles you can write. The article you plan to write when you designed your study, which is what we're talking about here. The confirmatory approach or to the article that makes most sense now that you have seen the results, they are rarely the same. And the correct answer is two. That's a BEM is saying. And bam goes on to say, the best journal articles are informed by the actual empirical findings from the opening sentence. Okay? So what curr is pointing out is that there are at least some textbooks that specifically advocate and approach of harking. Okay, so that's probably contributes to why hurricane is relatively rare. I'll just speak very briefly to my own experience. When I was in undergraduate, which was in the early 90s, so 990s. I remember my data analysis Professor highlighting the importance of exploratory research. And what are professor highlighted to us was that when you get a dataset, the first thing you want to do is you want to explore that data set as best you possibly can so you can understand what's going on. You're going on your data to the greatest possible degree. And I agree with that. I agree with that, with that process. And that is the exploratory research process that we were talking about earlier. The problem is that in that course, I do not remember there ever been an emphasis on the need for confirmatory work. In other words, the only memory that's stuck in my head after all these years is the importance of exploratory research and not the importance of confirmatory research. Now I have learned to other lessons from having worked with other scientists since then. But I want to point out that at least part of my training left me with the impression that are left me with a very biased impression about the importance of exploratory versus confirmatory work. And there was no there was no discussion that I can remember about the importance of distinguishing between these two types of approaches. So what I'm speculating here, and as I said, buckets of salt, is that one of the reasons why harking maybe relatively common, is because there are sources out there that teach us that harking is a desirable thing to do. What I'm highlighting this in this video, and I'm just repeating what people have said previously in the published papers that I've brought to your attention and in other papers as well. Is that harking is not an acceptable practice, is not something that we should do. So on that note, we'll end the video with this question. How should we actually do science then? Should we only do confirmatory work? I recognize that the process of this video or the message that you might be getting from this video, might be that you're getting the impression that the only kind of analysis you should be doing is confirmatory research. But no, that's not what I mean to say at all. Because exploratory research is absolutely essential and we need it all the time. Okay? So that's my second here. So confirmatory analyses are essential, but so are exploratory analyses case we need both the vital ingredient we need in order to fix this whole mess that comes from harking is that we need to distinguish in the work that we publish between which analyses came from exploratory or which results came from exploratory analyses and which results came from a priori or confirmatory analyses. If we just did that, if we just declared within our papers which analogy or which results arose from exploratory analyses and which results arose from confirmatory analyses. That will go a long way towards really clarifying the true picture of what we know about the biological systems that we're investigating. Okay. So that's really what I want to end with. The, the main points for this whole video is given right here. What we need to do is we need to distinguish or communicate which types of results that we've obtained came from exploratory analyses and which came from confirmatory analyses. And we can either do that directly in the paper, so we can explain these things in the paper. Or we can go even a step further by doing something called pre-registering our studies. And that's something that I'm going to talk about in the next video. Where what preregistration does is it provides us or it provides an additional source of evidence to support any claims that we might make that an analysis was confirmatory. So I'm going to stop the video there. I'll say I hope this has been helpful and I'll say, thank you very much.