Sunday, April 7, 2013

Thinking vs. doing in science

Lately, I've been thinking some about whether it's always best to think about experiments before doing them.  I know, of course, you have to think about experiments you do on some level.  I guess I mean more big picture thinking about "well, is this experiment informative, what will I learn from this?" kinda stuff.  I know that on the face of it, every experiment should make sense, and that the abstract notion of the scientific method we learn in grade school, every experiment should be designed to test some scientific hypothesis.  But on the other hand, most of the interesting findings that we have come across did not really arise in this way.  If we really thought through those experiments rationally, I think we would have ended up with considerably more boring results.  I'm guessing a lot of people have had this experience, where a little side-experiment ends up being super interesting and turns into a whole line of work unto itself.  And sometimes it doesn't make sense, but you just do it and see what happens.

That said, it's important to know where the line is between interesting side-experiment and hopeless distraction.  Also, one of the most important skills to develop as a scientist is the ability to take the seed of something interesting and turn it into a coherent and logical argument through carefully planned and executed experiments.  But I guess what I'm saying is that it's important to let the mind wander sometimes.  A question: how do people find those seeds?  Is it luck?  Hard work?  A "nose" for good problems?  Or is it a skill that one can acquire and hone with experience?  I think that last one is actually more true than people think...

Centromeric DNA

I was just listening to a couple of interviews (1,2) about ENCODE, which were interesting (link from Jan Skotheim's cool blog).  But aside from the arguments about ENCODE, there was some talk about comparisons to the human genome project, and how the human genome project had a nice end goal, but that goal was perhaps a bit more subtle than most realize.  Specifically, they didn't sequence all of the genome, because they skipped some of the repeats and the centromeric regions.  Then, in discussing ENCODE, they talk about how one could define the entire genome as being "functional" by ENCODE's definition because DNA polymerase interacts with every base.

Anyway, I thought it was somehow weird/fun to think about the fact that we don't know the sequence of the centromeric regions, but we know that DNA polymerase gets in there and replicates it.  I don't know why this seems strange to me, but I guess I like to think about these situations where something goes through some unexplored region and comes out the other end without really revealing much about its journey.  It reminds of how we have laid down these huge transatlantic cables connecting the US to Europe (in 1858!).  They basically just put down a big wire running the entire distance.  The deep of the ocean is still largely a mystery, but these cables run right through it, with our information none the wiser for their deep-sea voyage.  Somehow, I find this cool...

Sunday, March 17, 2013

Some thoughts on ENCODE and functionality

Gautham recently found a paper from some evolutionary biologists in which they discuss (or, more accurately, offer a scathing criticism) of ENCODE, specifically its claims about how much of the genome is functional:


It's an interesting read.  A bit of a polemic, but a witty one, with lines like:

"ENCODE chose to bias its results by excessively favoring sensitivity over specificity. In fact, they could have saved millions of dollars and many thousands of research hours by ignoring selectivity altogether, and proclaiming a priori that 100% of the genome is functional. Not one functional element would have been missed by using this procedure."


and

"For example, according to ENCODE, the putative function of the H4K20me1 modification is 'preference for 5’ end of genes.' This is akin to asserting that the function of the White House is to occupy the lot of land at the 1600 block of Pennsylvania Avenue in Washington, D.C."

I think the latter is really getting at the issue of functionality.  The authors claim that ENCODE makes a logical error in assigning function.  For instance, they say that ENCODE asserts something like the following: "1. Some DNA property (say, histone modification) serves in some situation to cause a functional change.  2. We observe that another piece of DNA has the same modification.  3. Therefore, that other piece of DNA must be functional as well."  Fair point.  Although I'm no expert, I believe transcription factor binding has to be one of the prime examples: many people have shown that there are tons of binding sites for transcription factors, identified both by sequence similarity to a motif and by actual ChIP experiments, that are not actually functional in the sense that they don't alter expression of some known gene.  Such "interesting looking places" are perhaps a reasonable place to start looking for function, but the fact that it may look interesting does not inherently prove that it is functional.

Of course, that begs the question of what it means to be functional...

The authors of this paper believe that evolutionary conservation is the only way to characterize something as functional.  I think I disagree.  I think, depending on your definition of functionality, you can have functional elements that are not conserved, and you can have conserved things that aren't functional.

Take a recent example from our lab.  We (meaning Hedia) have characterized a long non-coding RNA that, when you knock it down in mouse ES cells, changes the expression of a nearby protein coding gene (Hoxa1).  It seems that this non-coding RNA doesn't display any conservation.  Is this non-coding RNA functional or not?  Well, that depends.  We have no mouse knockout data showing any phenotype for this lncRNA.  Let's say we did have a mouse knockout and there was no overt mouse phenotype.  Does that mean this lncRNA is not functional?  Some would say no.  But such a strong definition of functionality would of course eliminate most genes, including the majority of those displaying purifying selection!  What do we then mean by "overt phenotype"?  Perhaps functionality from the phenotype point of view could mean a fitness defect?  Well, this would help you find genes that have obvious phenotypes, like embryonic lethality or heart development defects and so forth.  But would it help you find, for instance, hair color?  Are those genes functional by this definition?  Maybe, maybe not.  But then you're really evaluating just so stories about whether hair color actually affects fitness, etc.  And it may be that hair color in the end doesn't affect fitness at all.  I think most people would still want to consider genes affecting hair color to be functional.  Then, however, you've started going down the slope.  At what point is the phenotype deemed "unimportant" in a functional sense?  Isn't changing gene expression of Hoxa1 a phenotype just as much as changing hair color?  And on that basis, isn't our lncRNA functional?  I bet the authors of the paper would say so.  ENCODE (I think) takes this one step further and say that the mere transcription of a lncRNA is "functional" in the sense that it changes the biochemistry of the cell.  I'm not sure about that, but I can see the logic.  After writing most of this post, I came across this interesting post from one of the ENCODE folks that touches on the functionality issue.

So our lncRNA could be considered functional, but doesn't show any evolutionary conservation (at the level we have looked).  One could counter that this might just be a consequence of weakness in our metrics of evolutionary constraint–perhaps a more sophisticated view of conservation would reveal that this lncRNA is actually conserved, along with many other "important" DNA elements.  That would be great, and if that were the case, then ENCODE would be very useful as a dataset that one could use to evaluate such methods.  There is also an alternative: that our lncRNA is something specific to mouse, and does something specific in the mouse, and so is not detectable via conservation methods (presumably there are many such things that make a mouse different than a rat, etc.).  It is, I think, good to at least be open to this possibility.

I guess to sum up my ramblings up to this point, I would say that a strict definition of functionality based on strong phenotype would eliminate a bunch of conserved genes, and a looser definition of functionality (even not ENCODE-level loose) would include a bunch of non-conserved stuff.  So I don't know, but I think conservation is useful but imperfect measure of functionality.  Of course, this sort of very abstract thinking is probably of little use when evaluating whether to study something in the lab.  At that point, I think the definition of functionality largely comes down to a matter of taste.

Another issue is the value of ENCODE as a data set for others.  Michael Eisen, who I don't know but I think I would like to meet one day, had this thoughtful post about how ENCODE data doesn't fit the bill for everyone's scientific question.  I definitely understand this point.  My advisor at Courant (Charlie Peskin) posed to me the question "But at some point, wouldn't we have just sequenced everything there is to sequence?".  My answer was that there's always something else to sequence, because every new scientific question may require some new data, like adding some different drug to some different cell line at some different time point, etc.  Eisen's main point seems to be that the ENCODE data cannot cover all the different needs of different researchers for their specific questions.

My response to that is, well, whatever.  There's nothing I can do about whether or not it was a good idea to generate this or that dataset for ENCODE, and I honestly don't know whether any particular set of data was a good or bad idea to begin with.  Nobody's listening to my opinion, and frankly, they probably shouldn't anyway.  The fact is that the data is here.  The way we're using it in the lab is as reference data, data that we can use as a basis for some of our scientific questions.  It is undeniable, for instance, that it is very useful for our science to have RNA-seq and ChIP-seq data for a large number of cell lines, and it's also likely that we'll probably need to generate some additional genomic data sets of our own, which is probably in line with what the ENCODE people would have expected.  For example, we're using the GM12878 cell line for some allelic expression stuff these days. It's not the ideal cell line for imaging (we wouldn't have picked it), but we can make it work, and it would be really hard, both cost and time-wise, for us to generate all the background data and analysis we need in another cell line.  Another bonus is that the data we generate is not just "in some weird cell line" but in one of the ENCODE cell lines, which always helps when you try to argue "relevance" (whatever that means).  Anyway, I think our job in the lab is to spend our time trying to think of creative ways to use the ENCODE data to do cool science.  Speaking of which, I've spent a lot of time on this blog post...

Friday, February 8, 2013

Single molecule RNA FISH FAQ

UPDATE: I've expanded this post to an entire website.  Check it out, including its extensive FAQ section!

So we often get questions about how we know single molecule RNA FISH is working the way we think it is.  A LOT of questions.  Seriously, a lot.  At this point, we've got a fairly extensive list of canned answers, and so I thought it might be useful to post them all in one place for people.

Q: How do you know that the spots you are detecting are single RNA molecules?  Couldn't they be conglomerates?

A: This is a good question, and one that has a variety of answers.  Many of the control experiments that  Sanjay did are in his excellent Vargas et al. PNAS 2005 paper.  One (beautiful, in my mind) experiment that Sanjay did was the following.  He in vitro synthesized a bunch of target RNA and put it in two different tubes.  In these tubes, he labeled the RNA with probes, with the RNA in each tube labeled with a different dye (say, red or green).  Then he combined the two tubes, so he had one tube with RNA that was either labeled with red probe or green probe, but not both.  He then injected these into the cell and observed.  If the RNA were forming conglomerates, then you would expect yellow blobs containing both red and green RNA.  If they were single molecules, though, you would expect the spots to be either red or green but never both.  The latter is what he observed.  You might question whether this holds for endogenous RNA, but he expressed that RNA and compared intensities, and it was the same.  This means that the endogenous RNA was also single particles.  Nice!  Definitely caveats to this, and technically it applies only to this RNA, but whatever, I think this is pretty solid.

There are other things you can do.  One is to measure the fluorescent intensity of the spots and show that you get a unimodal distribution of intensities.  Pretty weak in my mind, because if you had some spots with two RNA and some with one RNA, these peaks would overlap so much that it would probably look like a unimodal peak anyway.  But what do I know.

To me, one of the strongest experiments are some new results from Eric Lubeck and Long Cai (Lubeck and Cai, Nat Meth 2012).  They use super-resolution microscopy to actually read out a barcode of different colors along a single RNA molecule.  Think about how cool that is for a minute!  Anyway, it's very hard to imagine that conglomerates of RNA would show anything like that sort of thing.  I think Sanjay has some other similar experiments that corroborate this.

Q: How do you know you're getting all the RNA in the cell?

A: Honest answer: no idea.  What we have done to get at this is compare to qRT-PCR data, for whatever that's worth.  I think Singer first did this in Femino et al. Science 1998, and Vargas et al. PNAS 2005 has a nice demonstration as well.  In those cases, you can try and use absolute standard curves to get an actual average number of RNA molecules per cell via RT-PCR and compare to what you get by molecule counting via RNA FISH.  In Vargas et al. PNAS 2005, we got a pretty close correspondence, with the numbers coming within 30% of each other.  But given all the vagaries associated with RT-PCR (RT efficiency, PCR efficiency measurement error, etc.), I'm sort of amazed this number came out so close.  I think others have shown the same thing with RT-PCR, and so I guess that's pretty good evidence.  Many have shown (e.g., Raj et al. Nat Meth 2008) that fold changes in RNA counts are similar when comparing RNA FISH to RT-PCR, but I'm not sure what that really tells you about detection efficiency except that it's the same (maybe good, maybe bad) in both conditions.

Some will tell you that you can detect the same transcript with two different probe sets and look for colocalization between the colors.  The idea is that if you detect with both colors, that means that your efficiency is high.  I don't think that actually makes sense–if you have an RNA that is inaccessible for whatever reason, this control tells you nothing, and if you have an RNA that is accessible, then a single color will probably detect it.  This two color colocalization approach is good for specificity, though...

Q: How do you know your probes are detecting the right RNA?

A: This is where the two color test comes in handy.  What you can do is label every other oligo with a different fluorophore (i.e., R,G,R,G,R,G,R,G...).  If the signals colocalize, that is pretty good evidence that you're detecting the right RNA, since it is very unlikely that a whole bunch of different oligos are all binding to the same incorrect target.  Usually, you don't need to do this, because if you get good signal in a single color, you are almost certainly detecting the right thing.  However, if you are seeing bright transcription sites, they could potentially be off targets because even a single oligo can light those up.  If you are doing analysis of those sites, you will probably want to check things out this way.  Also, lincRNA are very prone to these sorts of issues and you should really check those out with this "odds and evens" approach (we'll have a paper on this soon).

Q: What is the hybridization efficiency of each oligo?

A: Lubeck and Cai estimated a hybridization efficiency of around 60-70%, and we have seen similar numbers.  Hard to know for sure why it's not 100%, but whatever, if you get enough oligos, you'll be fine.

Q: How do you know ribosomes are not preventing RNA detection?

A: In the Raj et al. Nat Meth 2008 paper, we simultaneously targeted both the open reading frame (ORF) and the 3' untranslated region (UTR) with differently colored probes and saw good colocalization.  Ribosomes should bind to the ORF but not the 3' UTR, so if the ribosomes were causing a problem, we would have noticed many more spots with the 3' UTR probes.

Q: How do you know secondary structure is not a problem?

A: In some of our early experiments (Raj et al. PLoS Bio 2006), we targeted oligos to the PP7 RNA hairpin, which is a very strong secondary structure, and saw great signal.  Same for targeting MS2 RNA hairpins.  So I'm not so worried about it.

Well, hope this helps someone somewhere.  If you're a Ph.D. student doing RNA FISH, you should definitely memorize the answers to these questions–could really help you out in your quals!

Tuesday, February 5, 2013

"Tidy data"

http://vita.had.co.nz/papers/tidy-data.pdf

The article linked above talks about a typical but undiagnosed source of unnecessary effort in data analysis, untidy data, explains what 'tidy data' looks like, and illustrates some tools that help you make the change.

Keeping data tidy saves a lot of effort. "Tidy data" is not a table format that is visually pleasing for a presentation. It is the format you'd most like data to be in for manipulations. In fact, storing data in formats that make for visually pleasing tables usually makes them especially difficult for other folks to use within programming-style analysis tools like R and Matlab. I was reminded of this when a coworker asked for help turning his manual Excel workflow into an automated Matlab workflow.

After trying to get all kinds of different types of data incorporated into an analysis related to my current project, often from the Supplementary Info in scientific papers, I've found that the less creative the authors are with their data presentation, the easier the job is.

Excel unintentionally encourages the basic problem. Since you constantly see the data, and there are all kinds of features to make borders, change fonts, join cells and pretty things up, it is hard to resist the temptation to make it into a pretty table. So your workflow looks like this:

data  -->   presentable table   ( usually stored in an Excel file and given as Supplementary Info.)
presentable table  -->   Analysis and Graphics

That last step is hard because most presentable data is not readily amenable to downstream analysis. If instead you program your data analysis (or make use of Excel's more advanced features like pivot tables), your workflow can look like this:

data  -->   presentable table
data  -->   Analysis and Graphics

As it turns out, liberating yourself from the need to have your data look presentable on its own, lets you structure it in a way that makes for rapid and painless plotting and analysis. Optimize the data format for manipulability, and save your time and others'.

- Gautham