Saturday, April 4, 2015

Why not test your blood every quarter?

Lenny just pointed me to a little internet kerfuffle emerging because of Mark Cuban’s twittering about saying it would be a good thing to run blood tests all the time. Here's what Cuban said:







Essentially, his point is that by using the much larger and more well-controlled dataset that you could get from regular blood testing, you would be able to get a lot more information about your health and thus perhaps be able to earlier action on emerging health issues. Sounds pretty reasonable to me. So I was surprised to see such a strong backlash from the medical community. The counter-argument seems to have a couple of main points:
  1. Mark Cuban is a loudmouth who somehow made billions of dollars and now talks about stuff he doesn’t know anything about.
  2. Just as whole body scans can lead to tons of unnecessary interventions for abnormalities that are ultimately benign, regular blood testing would lead to tons of additional tests and treatments that would be injurious to people.
  3. Performing blood tests on everyone is prohibitively expensive, so we’d end up with “elite” patients and non-elite patients.
I have to say that I find these counterarguments to be essentially anti-scientific. On the face of it, of course Cuban is right. I’ve always been struck by how unscientific medical measurements are. If we wanted to measure something in the lab, we would never be as haphazard and uncontrolled as people are in the clinic. There are of course good reasons why it’s more difficult to do in the clinic, but just because something is hard does not mean that it is fundamentally bad or useless.

I think this feeds into the most interesting aspect of the argument, namely whether it would lead to a huge increase in false positives and thus unnecessary treatment. Well, first off, doing a single measurement is hardly a good way to avoid false positives and negatives. Secondly, yes, in our current medical system, you might end up with more unnecessary treatment–with many noting that getting into the system is the surest way to end up less healthy. That is more of an indictment of the medical system than of Cuban’s suggestion. Sure, it would require a lot more research to fully understand what to do with this data. But without the data, that research cannot happen. And having more information is practically by definition a better way to make decisions than less information, end of story. To argue otherwise sounds a lot like sticking your head in the sand. I'm also not so sure that doctors wouldn't be able to make wise judgements based on even n=1 data without extensive background work. Take a look at Mike Snyder's Narcissome paper (little Nature feature as well). He was able to see the early onset of Type II diabetes and make changes to stave off its effects. Of course, he had a big team of talented people combing over his data. But with time and computational power, I think everyone would have access to the interpretation. What's sad is for people to not have the data.

Leading to another interesting point from the medical research standpoint. If it were really rich people making up the primary dataset, I don’t think that’s a bad thing. Medicine has a pretty long history of doing testing primarily on non-elite patients, after all.

Friday, April 3, 2015

Theorists give great talks

We just had Rob Phillips come visit Penn and give a talk in the chemistry department. It was great! A few months back, we also had Jane Kondev come give a talk in bioengineering that was similarly a lot of fun. Now, Jane and Rob have a lot in common (both are cool, interesting people), but I think one common thread that links them is that they are both theorists by training. (Both do have strong experimental work happening in their group now, by the way.) I think theorists (at least in our sort of systems biology) give some of the most engaging talks, and I think the reasons why are illustrated in some of the best features of both their presentations.

The first departure from business as usual is in the amount of data presented. In Rob’s case, he presented almost no data from his own lab. Jane’s talk also had a lot of background from other people’s work, and the work he did present always came with heavy references to other literature and findings. This allows them to set up the conceptual issues well, as well as their place in the context of science. I think that some people feel like bringing up other work distracts from their own work, and that they don’t have time for it because they have so much of their own to present. I think that Rob and Jane’s talks prove these concerns to be overblown. Rather, I think that their talks feel rich with history and thus significance. Those are good things.

The other main thing I’ve noticed in talks by theorists is that they emphasize the conceptual. Most talks suffer not from a lack of data but an overabundance of data. Here’s a simple rule: if you’re not going to explain a piece of data, don’t show it. If it’s impossible for the audience to truly grasp how the data you show proves your point, then you may as well not show it and just tell them that it all works out. Often times in bad talks, it’s hard to tell that this is happening because people haven’t even set up the question well.

Which brings me to another nice thing about theorists: they aren’t afraid to delve into what might be called philosophy. For some of us, I think there is maybe a fear that people won’t take us seriously if we muse about the big picture in our talks. I think those fears are ill-founded. Overall, I think biomedical science could do with a little more thinking and a little less doing. Another nice thing about this is that for trainees, it can be very inspiring to think about deeper problems. Isn’t that what got us all into this in the first place?

On a related but peripheral note, I was at a conference a couple years ago and was shocked by what I was hearing from the students and postdocs. I asked one student what they thought about some fundamental question about the field, and they responded with a blank stare as though they had never been asked that question before. Another postdoc I met, when asked about some underpinnings of the field, literally responded with “I just want to get an assay that works and get a bunch of data”. If that’s you, go see a talk by a theorist and get back in touch with your inner scientist!

Wednesday, April 1, 2015

Feeling dated

Conversation in the lab:

Andrew: "Isn't there some song that the CIA plays to prisoners to get them to talk?"

Arjun: "Probably some crazy pop songs. There was this one from grad school that used to play on the radio in the lab that drove me nuts. I can't remember it, though."

Olivia: "What was it? I want to know–I feel like it would date you."

Sara: "Was it by John Lennon?"

Ouch.

Thursday, March 26, 2015

Kepler ate with chopsticks?

Check this picture out from Wikipedia:


(Yes, I know it's a compass.)

Friday, March 20, 2015

Priority in science is mostly a mirage

It seems these days that there are a lot of CRISPR priority fights out there, with the biggest being the patent dispute between Doudna and Zhang. Got me thinking about why we place so much emphasis on priority.

As scientists, it is a deeply engrained goal to be first. First to think of an idea, first to work it out, first to publish the idea. And in our culture, the one who does it first typically gets all the credit, sometimes even if they just “win” by a couple of months (although sometimes other factors come into play). It is what separates those perceived as the most shiny stars from the rest of us.

But think about it. If you’re working on something and get there just a few months before someone else, are you really that much more shiny than the next person? If the goal is to win a footrace, sure, you win. But if the goal is to create knowledge, the world would essentially be unchanged if you had never existed. Sobering. And true for the vast majority of us.

I think a pretty rational definition of a really original advance is something that if the researcher had not existed it is likely to have taken a really long time before anyone else would have come up with it. Almost by definition, such an advance is much more likely to come from one person. Like maybe RNAi. Or PCR. Although who knows? Maybe someone would have figured out these same things a few years later. I’m consistently surprised at how often you think you’re working on something completely alone only to find someone else hot on your heels. I think there are two reasons for this. One is that as technology and knowledge develops, the time becomes ripe for some discoveries. Like once sequencing was around, the discovery of several new long non-coding RNAs was essentially an inevitability. The other reason is that there are just so many smart people out there these days. With so many scientists, it’s virtually impossible that nobody out there is thinking about the same things you are. Even in math, which historically has worshipped at the altar of the solitary genius, often has multiple names attached to many new theorems. Which makes the exceptions all the more remarkable, like Yiteng Zhang’s amazing theorem about bounded gap primes or the discovery that primes is in P by a small group in India working largely independently. Kudos to them!

This is not to say that “meat and potatoes” science is not important. In fact, I think the steady, cumulative effects of incremental advances of the entire scientific community mostly outweigh the contributions of those few geniuses, especially in biomedical sciences in the current era. Somehow, I find this very reassuring and in many ways freeing: if you realize that we parcel out winners and losers in this race based on essentially arbitrary factors–and probably our innate desire to create heroic narratives–then it’s okay to just continue doing what you’re doing and not worry about it. In the long run, nobody really wins or loses, and science will continue moving, regardless.

So what is the strategy if you want to do something really original? Well, if the goal is to make a discovery that would take a long time to happen if you did not exist, then you can either do something really original or something that nobody cares about. Often one and the same!

Friday, January 23, 2015

Some thoughts on Tomasetti and Vogelstein (and post-publication review)

Interesting paper from Tomasetti and Vogelstein entitled “Variation in cancer risk among tissues can be explained by the number of stem cell divisions” (screw the paywall). This paper has generated a lot of controversy on Twitter and blogs, which is in many ways a preview of what a post-publication review environment might look like. I worry that it’s been largely negative, so here are my (admittedly relatively uninformed) thoughts.

Here is the abstract:
Some tissue types give rise to human cancers millions of times more often than other tissue types. Although this has been recognized for more than a century, it has never been explained. Here, we show that the lifetime risk of cancers of many different types is strongly correlated (0.81) with the total number of divisions of the normal self-renewing cells maintaining that tissue’s homeostasis. These results suggest that only a third of the variation in cancer risk among tissues is attributable to environmental factors or inherited predispositions. The majority is due to “bad luck,” that is, random mutations arising during DNA replication in normal, noncancerous stem cells. This is important not only for understanding the disease but also for designing strategies to limit the mortality it causes.
Basically, the idea is that part of the reason that some tissues are more prone to cancer is because they have a lot of stem cell divisions–an idea supported by the data they present. I think this is a really important point! In particular, because in some ways it establishes what I consider an important null, which is that in considering cancer incidence, it seems reasonable to consider that the more proliferative tissues will be more prone to cancer just because of the increased number of cell divisions. Darryl Shibata (USC) has a series of really nice papers on this point, focusing on colorectal cancer. In particular, in this paper, he points out that such models would predict that taller (i.e., bigger) people would have more stem cells and thus should have a higher incidence of cancer. And that’s actually what they find! I saw Shibata give an (excellent) talk on this at a Physics of Cancer workshop, and afterwards, a cancer biologist criticised this height result, incredulously saying “Well, but there are so many other factors associated with being tall!” Fair enough. But I think that Darryl’s is an economical model that explains the data, and would be what I would consider an important null that deviations should be measured against. I think this is a nice point that Tomasetti and Vogelstein make as well.

What are the consequences of such a null? Tomasetti and Vogelstein frame their discussion around stochastic, environmental and genetic influences on cancer incidence between tissues. Emphasis on between tissues. What exactly does this mean? Well, what they are saying is that if you compare lung cancer rates in smokers vs. non-smokers (environmental effect), then the rate of getting cancer is around 10-20 times higher, but your chances of getting lung cancer even as a non-smoker is still much higher than getting, say, head osteosarcoma, and a plausible possible reason for this is that there are way more stem cell divisions in lung than in the bones in your head. Similarly, colorectal cancer incidence rates are much higher in people with a genetic predisposition (APC mutation), but again, even without the genetic predisposition, that is still many orders of magnitude higher than in other tissues with much lower rates of stem cell divisions. I think this is pretty interesting! Of course, as with Shibata’s height association, the association with stem cell divisions is not proof that the stem cell divisions are per se the cause of this association, but one of the nice things about Shibata’s work is that he shows that a model of stem cell divisions and number of genetic “hits” required for a particular cancer can match the actual cancer incidence data. So I think this is a plausible null model for a baseline of how much certain tissues will get cancer. Incidentally, this made me realize a perhaps obvious point on the genetic determinants of cancer: if you find an association of a gene with cancer incidence, then it may be that the association is because the gene is associated with, e.g., height, in which case, yes, there is technically a genetic underpinning for that variation, but it is hard to imagine designing any sort of drug based on this finding. Tomasetti and Vogelstein make this point in their paper.

The authors then go on to further analyze their data and separate cancers into ones in which the variance in incidence is dominated by “stochastic” effects vs. “deterministic” effects. I can’t say I’ve gone into the details of this analysis, but it seems interesting–and a natural question to ask with these data. Here are a few thoughts on the ideas this analysis explores. One question that has come up a lot is why is this correlation not so strong, especially on a linear scale? I think that one issue is that the division into stochastic, environmental and genetic is missing a big component, which is the tissue, cell and molecular biology of cancer. Some tissues may require more genetic “hits” than others, or a long series of epigenetic effects, or have structures that enable rapid removal of defective stem cells, and so even tissues with the same number of divisions, in the absence of any genetic or environmental factors, will have different rates of cancer. Another issue is that these data are imperfect, and so you will get some spread no matter what. Still, I think the association is real and interesting.

Anyway, I think this “null model” is pretty cool. I wonder if one of the reasons that we focus so much on environmental and genetic effects is that we can do “experiments” on them, whereas the causal links in the stem cell division hypothesis are hard to prove.

There was a very interesting critique from Yaniv Erlich that said that the authors’ analysis implicitly assumes that there is no interaction between the number of stem cell divisions and genetic and environmental factors. A good point, although I do think that Tomasetti and Vogelstein have thought about this–as I mentioned, they say explicitly:
The total number of stem cells in an organ and their proliferation rate may of course be influenced by genetic and environmental factors such as those that affect height or weight.
Their example about the mouse vs. human incidence of colon vs. small intestine cancer in the case of the APC mutation is I think a nice piece of evidence suggesting that number of divisions is very important factor in determining cancer incidence. Although again, many alternative explanations here.

I think some of the confusion out there about this paper can be summed up as follows:
“You are a smoker and I am not, so I have a lower rate of getting lung cancer.”
“Yeah, but you still have a much higher rate of getting lung cancer than bone cancer.”
“Uhh… okay… sure… don’t think I’m gonna take up smoking anytime soon, though.”
It’s just a weird comparison to make. That said, I don’t think the authors really make this comparison anywhere in their manuscript. What I think they are saying at the end is that for cancers that have strong determinants due to environmental factors, lifestyle changes and other such interventions could be useful (like quitting smoking), whereas for other cancers that arise more randomly, we should just focus on detection. Although I have to admit that perhaps I’m missing something, but this seems like a point one could make even without this analysis.

There has been a lot of discussion out there about how weak the correlation is and whether its appropriate to use log-log or linear scales and so forth. I think the basic point they are trying to make is that more highly proliferative tissues are more prone to cancer. I think the data they present are consistent with this conclusion. Whether the specific amount of variance they quote in the abstract is right or not is an important technical matter that I think other people are already talking about a lot, but I think the fundamental conclusion is sound.

A note about the reaction to this paper: in principle, I like the concept of moving from pre-publication anonymous peer review to a post-publication peer review world. I think that pre-publication anonymous peer review is slow, arbitrary, and (most importantly) demoralizing, especially for trainees. That said, now that I’ve seen a bit of post-publication peer review happen online, I think the sad thing I must report is that in many cases, the culture seems to be one of the hardcore takedown, often in a rather accusatorial tone. And I thought it was hard to get a positive review from a journal! Here are some nice thoughts from Kamoun, who recently responded (admirably) to an issue raised on Pubpeer.

My view is that in any paper with real-world data, there will be points that are solid and points that are weak. In post-publication peer review, we run the risk of reducing a paper to a negative soundbite that propagates very fast, and thus throwing out the baby with the bathwater, not to mention putting the author (often a trainee) under very intense public scrutiny that they might not be equipped to handle. I think we should be very careful in how we approach post-publication review because of its viral nature online. Anyway, those are my two cents.

PS: Apropos of discussions of log-log correlations vs. linear correlations, we have a fairly extensive comparison of RNA-seq data to RNA FISH data. More very soon.

Friday, January 16, 2015

Gordon Conference turns graduate student into crazy reptile lady

Just got back from a cool Gordon conference on Stochastic Physics in Biology with a couple students in the lab. Lots of interesting science, and lots of cool people to talk with as well!

The food was overall really good, but one day, we decided to go get some Mexican food from a local taco shack. Delicious! On the way back, we noticed a little store on the side of the road called "Exotic Emporium". When we went inside, what did we find but a reptile pet store. Olivia fell in love with those little critters, and here's the evidence:

"Hello, strange lizard":


"That is a large snake!"


"I think I like snakes."


"Okay, put the snake around your neck then." "Umm, okay..."


"The colors! The colors!"



"Can I keep it?"