Okay, In this video, we're going to continue our discussion of randomization in our previous video. So in part 1, we talked about why it was essential to include randomization as a component of an experimental design at all stages of that experiment. In this video, we're going to focus on a fairly different topic and infects. We're going to mention randomization very little. Even though it's role is enormous. And we're going to talk about randomization and the context of sampling populations. And specifically, we're going to focus on what can go wrong if we don't sample populations effectively. So that's we're going to do. Just to get us on track for this video, I want to start by just revisiting one example from our previous video where we use this these beautiful puppies to discuss an example where we said we were imagining, we want you to have an experiment in which you would ask whether or not petting a puppy would change its average rate of tail wagging. And we talked about how it was important to use randomization in order to assign our various subjects to our different treatments. And we talked about why that was so important. We said it was important to randomly assign our subjects who are treatments in order to account for potential confounding factors. And as a result, the randomization helps us to be much more confident that the conclusions that we reach are correct for this particular set of subjects will chaos or that our conclusions are robust for this set of subjects. What this type of randomization does however, or, or, sorry, I should say, this randomization of subjects to treatments does not do something that's really important. However, this form of randomization does not ensure that the conclusions that are drawn from this set of subjects will apply to puppies in general. Okay, and why that is, is the focus of this video. Science is most effective when the conclusions that we draw from particular study will apply to some wider population that we're interested in. For example, let's imagine we were ecologists are evolutionary biologists. And we wanted to study the critters and plants that were living in this beautiful valley. And this got high in the Scottish Highlands. If we wanted our results from our study to be able to apply to the population that we're studying here generally in this valley. Then we would want to make sure that the subjects that we use in our study were representative of this population. And one of the best ways, or perhaps I mean, the best way of obtaining a sample that is representative of a population is to use a method that incorporates random sampling. And that's because of random sampling provides a very reliable way of obtaining a representative sample. I want to point out that exactly how we would implement random sampling in this context is not the focus of this video. There are entire courses that are devoted to teaching the skills that are required to sample a population such as this one. In a way that you will get a sample that is representative of the population you wish to study, okay? So those kinds of topics are discussed elsewhere. The importance of finding a sample that's representative of your population does not just apply it. You critters in valleys and to plants living in valleys. They apply to humans as well. So if we wanted to perform a study on humans that say that lived in the city of Edinburgh. Than we would want to make sure that the sample that we obtain was as representative as possible of the population and we wished to make conclusions about. And again, random sampling. All else being equal is going to provide us with the most reliable method of obtaining a representative sample. What we're going to talk about in this video is kinda the flip side to that. We're going to talk about what are the consequences of using samples that are not representative of the population that we wish to understand or that we wish to actually be able to help if we have an applied research project. Before we get into that, I'd like to just point out something that will hopefully be self-evident in a moment. Okay? And that is simply any results that we obtained for an experiment will apply within the context of that experiment. And for individuals that are represented by the sample that was used in that experiment. So for example, if you perform an experiment in the lab, it should be self-evident that the conditions in the lab will not be the same conditions as you had in the field. So by the field, I mean out in nature. And as a result, we can't know whether or not the results we obtain in the lab we'll actually apply in the field. Okay? And that's because the experiments that we performed in the lab does not give us any information about how the experiment would turn out in the field. Because we didn't do the experiment in the field. That should be obvious. But the same thing is true when we're talking about the samples are the types of samples that we're using if we used a sample. So a group of individuals, whether they be humans or mice, or plants, or genotypes of a fungus, whatever we're studying. If we use a sample for an experiment, say in the lab or wherever we do it. The results from our experiment will apply to individuals that are represented by that sample. But We will not know whether or not the results from our experiments will apply outside of that context of our sample. So in other words, we will not know whether or not the conclusions that we draw from an experiment with samples with particular qualities will also apply to samples that have other qualities. And again, that's simply because our experiment will not have any information within it that will allow us to extrapolate to other types of individuals. Okay. That might sound a little bit vague if it does, don't worry, because the rest of this video is devoted to fleshing out the points that were just made on this slide. Here's a first example that's relatively benign. We're going to get to some more examples with greater societal impact in a moment. Okay. This slide should seem familiar. This comes from an experiment that I performed early on in my career at the beginning of my master's degree. And we talked about this in a previous video. And what I was doing in this experiment was I created three different types of floral displays because I was interested in understanding how the arrangements of flowers influenced pollinator behavior. Ultimately because I was interested in, um, how natural selection might have caused different types of floral displays to evolve. Okay, so that's the very quick recap of what this was about. What I'd like to point out here is in the course of performing this experiment, my supervisor, Lawrence harder, and I decided to create treatments with particular qualities. In particular, we decided to set up our treatments so that all of our treatments displayed the same number of flowers. So our three treatments, the unmanipulated low density and sickens treatments, all presented only. Eat flowers. I just checked my fingers to make sure is actually eight. There'll be embarrassing if I did something like that. Okay. So all of our treatments always had eight flowers and we did that. We chose to always display the same number of flowers because we wanted to minimize the amount of variation in the responses of our bumblebees that we're going to be visiting these flowers, are visiting these floral displays. And we thought that having larger number flowers are small in number flowers might influence how the bumblebees responded to our three treatments. So we wanted to minimize the amount of variation in the pollinators behavior so that we would have more power to detect any differences that there might be between these particular treatments. Okay. As a result of this, as a result of this decision-making, I can only use plants from the population where I was working that had certain qualities. In particular, because I needed all the plants to be have at least eight flowers and I was removing half of them. Then I had to make sure that all the plants that I was using it, we're displaying at least 16 flowers. And that's important because that means that I had to leave out many plants in this population that we're displaying fewer this 16 flowers. And those plants might have had less resources. And as a result, they might have done things like produce less nectar. They might have had smaller flowers, et cetera, et cetera. Okay. What I'm it's the reason I'm saying this is because our decision to control our treatments as much as possible meant that we had to omit certain plants within the population that had certain qualities. Okay? As a result of this, what I want to point out is that the results from our experiment where we had eight flowers on each of our treatments might not actually apply for other contexts. Specifically, it might not apply when we use plants or for plants that might have had a smaller number of flowers. Or perhaps for plants that have a large number of flowers. We don't know whether or not these results will apply for larger or smaller floral displays, because this experiment does not have variation in floral displays. Okay? This decision to create this experiment like this, exemplifies a trade off that occurs. I don't want to say everywhere in science I may not be everywhere. But it's extremely common where the trade-off involves having to think about a balance between increase in experimental power by controlling variation versus decreasing the generalize ability of the conclusions that we make from a particular experiment. This is something a trade off that occurs everywhere. Anytime someone decides to work with, say, a particular strain of mice or only with inbred individuals. That means that you might be making that decision in order to increase your experimental power. But you also decrease the generalizability of your experiment. In the context of this experiment. Because we know our results were limited in their scope. But for good reasons. I would say that no one should ever draw firm conclusions from our study alone. Okay? We need further studies in order to check these general conclusions in other contexts. And it turns out for this particular question, there have been other studies performed. And it seems like the main conclusions that arose from this study are more general. So in this case, the results that we obtained contributed to a larger body of work that allows us to reach a robust conclusion. Okay? The main thing I don't want to really emphasize here is because every experiment will be limited in terms of the context in which is performed. We should rarely draw from conclusions from a single study. I want to shift focus now to a much more serious context. Science itself is serious. So in that sense, the previous example was serious. But the next example I want to discuss is when we have serious consequences for human society that are pretty damning, as you'll see in a moment. So I'm going to focus on two papers that came out in 2021. And they are part of a larger body of work. So I didn't want to give the illusion here that what I'm presenting here is all there is on this topic, okay, this is a much larger issue. So this first paper here that I want to point to, it says sex and gender based formula, pharmacological responses to drugs. I'm going to pull, I've pulled out this quote from this paper, cuz this quote sets the stage for we're going to talk about. It says In 2001, a US Government Accountability Office report found that 80 percent of prescription drugs that were withdrawn from the market between 997 and 2000 exhibited greater adverse effects and toxicity in a women than in men. Think about that for a moment. This means that there was, there is evidence that in the not-too-distant past, the vast majority of drugs that had to be removed from the market, well removed because they had really bad consequences. But only primarily for, for one part of the population, those difficult consequences were greater for females than they were for males. Why is that? Well, this paper treads a little bit lately and pointing the blame on for the, for the, for this major problem. And they, what they do is they say perhaps this is due to a bias in the samples that are used in research where biomedical research often focuses on females, rather, sorry, usually focuses on males, pardon me, and has very little focus on females. This next paper makes that same point, but it's much harder hitting. So this paper, considering sex is biological variable will require global shift in science culture. It starts out with this quote where it says the last decade has seen increased public awareness that women are vastly more likely than men to be misdiagnosed in a wide array of medical conditions. That's terrifying and that's horrible. These authors are willing to point the finger much more directly at the cause that they believe is what they believe is the real cause for this. And that's just illustrated here. I'm just going to read this out. They say it much better than I could articulate. They say our incomplete understanding of the aetiology, symptomology, and treatment of mental and neurological disease in women is due in large part to the neglect of female subjects in preclinical neuroscience research. And now landmark 2011 evaluation of biomedical publications found that neuroscience studies used a male animals six times more often than they used females. A more recent analysis of papers published in 2017, sadly suggests that this imbalance has only barely begun to improve despite the more widespread recognition of the disparities of women's health mentioned above. So these are things I talked about earlier in the paper. So they are, these authors are pointing the finger directly at a bias in the samples that are used in research for being the cause of worse medical Merce, worst medical resources for females compared to males. Basically, that's what they're saying. Ok. Would like to point out is that these authors go on and they point out that it's not just, you know, these authors that are taking this seriously, but this is serious enough that funding bodies have acted upon this because this is a serious problem. They go on to say, following similar policies by Canadian Institute, sorry, bike by the Canadian Institutes of Health Research and the European Commission. The US National Institutes, sorry, the US National Institute of Health, the NIH, introduce the considering sex as a biological variable mandate in 2016 as part of a broader initiative to improve the rigor and reproducibility of research funded by the NHS. The policy states the grant applications must include both male and female subjects and or cell lines and experimental design and analysis. With primary objectives. A broadening the general knowledge base and delivery in a more refined understanding of how and in whom the basic science findings will best translate into clinical applications. So the funding bodies are recognizing that using an unrepresentative sample Is leads to, can lead to some terrible outcomes for society. And it needs to change. We need, when we do science, we need our samples would be representative of the populations that we hope to understand and we hope to help. So the main point that I'm making here from this horrible example in this horrible situation is that studies that use mayest by male biased samples are inappropriate to draw conclusions for the population of interests where the population of interest will be humanity in general. Okay. And from this, I'd like to point out that something should be relatively obvious. But sometimes it, I want to say it anyways. Our understanding of biology will always be more narrow when we focus on a subset of the population that we're interested in. Okay? So this is just another way of saying that if you wants to understand a population in general, that we need to use a sample that is representative of that population. So all of these, it really dire consequences for human society aside, I don't mean to sound as callous as, as that might appear. What I'm trying to use and try to recognize that we have some really dire consequences of biased sampling. But what I want to do now is just want to shift gears a little bit and think about these issues in a more general context. Okay? And so by shifting gears, I'm not trying to downplay the seriousness of these, this issue for society. What I'd like to point out is that when research focuses on just one subset of society or of a population. So when you just focus on males, then what that means is that beyond the societal problems that may result, it also means that we are likely to be missing out on some very interesting biology to consider. So what we've learned from these horrible situations is basically females and males are not the same. That's not boring. Fats should, That's fascinating. And we can see that the way in which female and male bodies respond to medical treatments are not the same. That opens a really interesting biological question, which is, why does some medical treatment or some drug or whatever I've just called this phenomenon zed. Why does this phenomenon differ between females and males? That is bound to be a really interesting biological question. That really can use some, really can use some good attention. But that kind of question gotta be missed entirely if we only focus on one aspect of society or one aspect of a population. Let's shift gears here ever so slightly, okay, and we're going to give you a question to think about. Let's imagine that you were studying some particular biological phenomenon, let's say a response to a drug or the effect of a particular gene on some process. Okay. And you were studying this phenomenon in several strains or genotypes. Let's imagine we're looking at 10 different strains, are ten different genotypes. You found a positive response. So in other words, you found something that you considered to be an interesting response. And some of those strains are in some of those genotypes, but not others. You're now face or the question, the question is, should your further research focus only on the strains are the positive results or should your further research consider all of the strain, so including the strains that do not exhibit the phenomenon of interest. So I just want you to stop and think about this for a moment. Okay, Now that you had a good think, perhaps you pause the video to think about this. Before going further, want to point out that there are some genuine incentives to focus only on the individuals, again, the positive results. Okay? I'm not saying that these incentives are good reasons from a scientific and societal point of view. But they are reasons that exist. And that is, if you want to publish a paper, which if you're having a career in science, that's something you need to do. It's often easier to do this if you have a simpler story to tell in your research that you publish. Okay? So if you had a variety of genotypes that displayed inconsistent responses to something, it might be in your interest, your own personal interest, to just focus on the genotypes that had that positive response. And in your subsequent work, you can continue focusing just on those genotypes. It had a positive response because that's the, basically the easiest way forward. Have a group of individuals or a group of genotypes that you know, are giving you the kind of response that you think would be easy to publish. I'll be honest and say, you know, that, that is very tempting. Form of thinking. In the grand scheme of things a high imagine you can guess what I'm going to say next. And that is in the grand scheme of things. That kind of thinking, however, is not good for science and it's not good for society. If we were to only focus on the genotypes are strains that gave us a positive response. Then, first of all, we would be misrepresenting the biology that we claim to be explaining. Sorry, if we, if we only focus on the positive results and did not emphasize that there were some strange genotypes, did not, that did not have this response. We'd be kind of creating an illusion that our results might apply very widely because we've simply left out the cases where the positive result does not occur. And that's not something that should be, it should have a home in science. Science is about finding truth. And we do not want to leave out results simply because it might be inconvenient to us. As you might also gas based on their previous slide. If we just focus on the strains had the positive results. It's also means we're probably going to be missing out on some really interesting biology. If you have one genotype that produces a response and another one that doesn't. What that strongly suggests is that you have interactions between genotypes. Sorry. What was that? I'm sorry, strike that. I was imagining a slightly different situation when I said that. What this obviously would mean if you have a response by one gene, it's height, but I buy not a response by another. That would indicate, give us really strong indication that there's something different about the genes and this positive strain versus the strains that gave a negative response. There's something different about the genes that leaves her that differential response and that that difference in the genetic response, the effect, that would be a really interesting avenue for research. And by looking at trying to answer that question, even if it's hard, that will help us to obtain a much more complete understanding of the biology, the hope to understand. I'm going to close with one last example. I talked about this paper when I was introducing the topic of questionable research habits. Where this is one of the major papers that drew attention to the reproducibility crisis. So that paper is called re standards for preclinical cancer research. And this paper, it takes a very critical view of research done in the cancer field. And they talk about a number of things that they think really need to be improved. And I just want to highlight one of them because it lies in this context of having representative samples. So just read this here, says an engineering challenge and cancer drug development lies in the erroneous use and misinterpretation of preclinical data from cell lines and animal models. Okay? So there's something wrong with the cell lines in animal models that are used. That's what they're claiming. So the limitations of preclinical cancer models have been widely reviewed and are largely acknowledged by the field. They include the use of small numbers, a poorly characterized tumor cell lines that inadequately recapitulate human disease. And then go on to say some more. But that's really the part that I wanted to focus on. What they're largely saying here is that the cell lines that are being used are not appropriate to be able to make conclusions about the population they really wish to understand. First of all, they have very few cell lines, and the cell lines they have are not very well understood or may not may not kinda behave in the way that they hoped they would be representative of the situation they want to help. So here are there, Here's one of their big recommendations. They simply say studies should not be published using a single cell line or model, but should include a number of well-characterized cancer cell lines that are representative of the intended patient population. So this really kind of captures the main point and trying to make from this video. Which is that even if you're working with cell lines, if your goal is to be able to inform biology beyond a particular cell line that you're working with. You need to work with a set of cell lines that are representative of the population that you hope to understand. And that basically brings us to the end of our video. The main two points in this video. First of all, that you want to use samples that represents the population of interest. So whatever population it is, you wish to help her wish to understand, your sample should be representative of that population. And random sampling is an integral part of obtaining representative samples. If it turns out for whatever reason you are unable to obtain a representative sample, then it behooves you when conducting your research and communicating your research, to consider and communicate how any way in which your samples may be unrepresentative of the population may influence your conclusions. In other words, it's important to articulate the limitations of your conclusions because of the nature of the material that you are using in your study. I'm going to stop there. I hope this video has been helpful. And I'll say, thank you very much.