Saturday, April 6, 2013

How well can we predict future criminal acts from fMRI data?


paper recently published in PNAS by Aharoni et al. entitled "Neuroprediction of future arrest" has claimed to demonstrate that future criminal acts can be predicted using fMRI data.  In the study, the group performed fMRI on 96 individuals who had previously been incarcerated, using a go/no-go task.  They then followed up the individuals (up to four years after release) and recorded whether they had been rearrested.  A survival model was used to model the likelihood of being re-arrested, which showed that activation in the dorsal anterior cingulate cortex (dACC) during the go/no-go task was associated with rearrest, such that individuals with higher levels of dACC activity during the task were less likely to be rearrested.  This fits with the idea that the dACC is involved in cognitive control, and that cognitive control is important for controlling impulses that might land one back in jail.  For example, using a median split of dACC activity, they found that the upper half had a rearrest rate of 46% while the lower half had a rearrest rate of 60%.  Survival models also showed that dACC was the only variable amongst a number tested that had a significant relation to rearrest.

This is a very impressive study, made even more so by the fact that the authors released the data for the tested variables (in spreadsheet form) with the paper.  However, there is one critical shortcoming to the analyses reported in the paper, which is that they do not examine out-of-sample predictive accuracy.  As I have pointed out recently, statistical relationships within a sample generally provide an overly optimistic estimate of the ability to generalize to new samples.  In order to be able to claim that one can "predict" in a real-world sense, one has to validate the predictive accuracy of the technique on out-of-sample data.

With the help of Jeanette Mumford (my local statistical guru), I took the data from the Aharoni paper and examined the ability to predict rearrest on out-of-sample data using crossvalidation; the code and data for this analysis are available at https://github.com/poldrack/criminalprediction.  The proper way to model the data is using a survival model that can deal with censored observations (since subjects differed in how long they were followed).  We did this in R using the Cox regression model from the R rms library.  We replicated the reported finding of a significant effect of dACC activation on rearrest in the Cox model, with parameter estimates matching those reported in the paper, suggesting to me that we had correctly replicated their analysis.  

We examined predictive accuracy using the pec library for R, which generates out-of-sample prediction error curves for survival models.  We used 10-fold crossvalidation to estimate the prediction error, and ran this 100 times to assess the variability of the prediction error estimates. The figure below shows the prediction error as a function of time for the reference model (which simply estimates a single survival curve for the whole group) in black, and the model including dACC activation as a predictor in green; the thick lines represent the mean prediction error across the 100 crossvalidation runs, and the light lines represent the curve for each individual run.  



This analysis shows that there is a slight benefit to out-of-sample prediction of future rearrest using dACC activation, particularly in the period from 20 to 48 months after release.  However, this added prediction ability is exceedingly small; if we take the integrated Brier score across the period of 0-48 months, which is a metric for assessment of probabilistic predictions (taking the value of 0 for perfect predictions and 1 for completely inaccurate predictions), we see that the score for the reference model is 0.214 and the score for the model with dACC as a predictor is 0.207. We found slightly improved prediction (integrated Brier score of 0.203) if we also added Age alongside dACC as a predictor.  

The take-away message from this analysis is that fMRI can indeed provide information relevant to whether an individual will be rearrested for a crime.  However, this added predictability is exceedingly small, and we don't know whether there are other (unmeasured) demographic or behavioral measures that might provide similar predictive power.  In addition, these analyses highlight the importance of using out-of-sample prediction analyses whenever one makes a claim about the predictive ability of neuroimaging data for any outcome.  We are currently preparing a manuscript that will address the issue of "neuroprediction" in greater detail.  

Wednesday, March 13, 2013

My adventures in self-quantification


In September 2012 I began a project to characterize how my own brain function and metabolism fluctuate over the course of an entire year.  This has involved MRI scans three times a week along with blood draws once a week and daily tracking of a large set of potentially interesting variables.  This post is the first installment in my story about the experience.

First, the motivation.  In the last couple of years I have become very interested in understanding the dynamics of brain function over a days-months timescale and how they relate to cognitive function and bodily metabolism.  This interest has been spurred by a number of influences, such as my growing interest in nutrition and its relation to brain function as well as my ongoing interest in better understanding psychiatric disorders.  Once I started thinking about the issue, it became very clear that there were basically no data in existence that provide any insight into how the function of an individual's brain fluctuates over such a relatively long time course. This is probably not surprising, because doing studies with volunteers that require repeated testing over a long period of time is very challenging.  

At some point in 2011 it dawned upon me that I should try to bootstrap such a study by collecting data from myself.  There were several inspirations for this idea.  First was Michael Snyder's "integrated personal omics" study, published in Cell in 2011, in which he repeatedly collected blood from himself and performed a broad set of "omics" measures on his samples, which provided some interesting insights into the temporal dynamics of metabolic function.  Second was my interaction with Laurie Frick, who is the artist-in-residence at the UT Imaging Research Center.  Laurie's work is based on patterns that she finds in data obtained by self-tracking, and she is deeply enmeshed in the Quantified Self movement (see her excellent TEDx talk).  Talking to her got me increasingly interested in tracking a broader set of data about myself. The very fun book Smoking Ears and Screaming Teeth also convinced me that self-experimentation is not (completely) crazy, and actually has long been an important tool for scientific discovery.

In early 2012 I began hatching a plan to collect a broad set of data on myself.  It was essential that the all aspects of data collection were as consistent as possible in order to minimize extraneous variability in the data (such as time of day effects).  I ended up settling on a schedule of three MRI scanning sessions a week, at consistent times of day and days of the week (one afternoon and two mornings every week).  Each of the MRI scanning sessions includes a resting state fMRI scan, which will allow us to assess how functional connectivity between brain regions fluctuates over time.  In addition, once a week I perform other scans, including structural MRI (T1- and T2-weighted), diffusion tensor MRI (to assess white matter connectivity), and task fMRI (using a working memory task with faces, scenes, and chinese characters).  

I also wanted to collect biological samples in order to measure the relation between bodily metabolism and brain function.  Working with some molecular biologists here at UT (along with helpful input from the Snyder lab at Stanford), we developed a protocol in which I have 20 ml of blood drawn once a week (while fasting, immediately after one of the morning MRI scans). This sample is then processed to extract RNA, white blood cells, and plasma, all of which are frozen for later analysis.  This will let us examine many different aspects of metabolism, including gene expression (via RNA sequencing), metabolomic and proteomic analyses, and other potential analyses to relate metabolism to brain function.

Finally, I realized that the dataset would be most useful if I also collected as much data as possible about my daily life activities.  Working with Zack Simpson, we developed a self-tracking app using the Appsoma framework, which allows me to easily complete surveys every morning and evening and after every MRI scan.  These data are automatically fed into a web database which is the central repository for all of the self-tracking data in the study other than MRI and biological analyses.  Some of the things that I track daily include:
- blood pressure and weight (using a FitBit Aria wireless scale)
- foods eaten, alcohol intake, and supplements/medicines taken
- exercise, time spent outdoors, and physical soreness
- a free-text log of daily events
- sleep quality (assessed both by subjective report and using a ZEO sleep monitor)
After every scan I also complete a mood questionnaire and also provide a structured report of what I was thinking about during the resting state fMRI scan.  I should note that other than the addition of all of these tracking activities, I have done my best to keep my life as consistent as possible and have avoided any other major lifestyle changes.

With this plan in place, we began data collection on September 25, 2012.  We treated the first month as a pilot period, and made some changes to the imaging protocol to optimize data collection, beginning the production period on October 22, 2012.  In total so far we have collected 20 blood samples and 55 MRI scanning sessions.  Members of the research team have started analyzing the data, though I have made every effort not to expose myself to the results of any analyses that examine changes over time, because I don't want the results to feed back and change my behavior.  

When I describe this study, many people ask if I am worried about being exposed to MRI scanning so often.  My answer has been "no", at least not with regard to the magnetic fields involved in MRI; there is no evidence of lasting effects of MRI exposure (though of course we can't ever prove that something is safe).  However, soon after the study began it became clear that there was a side effect that I had to worry about, which is the intense noise of the MRI scanner.  I have long suffered from tinnitus (which I attribute to too many loud rock shows as a youngster without ear plugs), and within the two weeks of scanning I noticed that my tinnitus was increasing.  For this reason, I went to the UT Speech and Hearing Center and had my hearing tested.  I had never had my hearing tested as an adult, but I was not terribly shocked to find out that I had quite significant high frequency hearing loss.  Because I don't want to damage my hearing any further, I have continued to get tested each month.  The results had been fairly stable until early March, when they showed about a slight worsening at 6000 Hz (consistent with a subjective increase in tinnitus around the same time).  For this reason, I am taking the month of March off of scanning, and will have my hearing re-tested at the beginning of April before resuming the scans.

We will also make some changes to the MRI protocol to reduce scanner noise.  We have an OptoAcoustics noise canceling headphone system in place at our imaging center that works quite well to reduce the noise of the functional MRI scans, so those can continue without much danger.  However, we will likely discontinue some of the other scans (such as gradient field maps, which are useful but not necessary) and greatly reduce the frequency of others that we don't expect to change much over time (including the anatomical and diffusion scans), because those scans are not compatible with the noise cancellation system.  I am hopeful that with these changes I can continue scanning without danger of further hearing damage, while still collecting a very useful dataset.  

Assuming that I am able to continue scanning, I plan to collect 50 weeks worth of usable data, which should provide sufficient power for an initial set of analyses.  This will likely take though the end of 2013 due to travel and other events that will interfere with data collection during some weeks.  Once we have completed our initial set of analyses, nearly all of the data will be made available to other researchers, which I hope will help spur new analyses.  I'll keep you all posted as the study moves along.

Wednesday, February 20, 2013

Anatomy of a coding error

A few days ago, one of the students who I collaborate with found a very serious mistake in some code that I had written.  The code (which is openly available through my github repo) performed a classification analysis using the data from a number of studies from the openfmri project, and the results are included in a paper that is currently under review.  None of us likes to admit mistakes, but it's clear that they happen often, and the only way to learn from them is to talk about them. This is why I strongly encourage my students to tell me about their mistakes and discuss them in our lab meeting.  This particular mistake highlights several important points:
  1. Sharing code is good, but only if someone else actually looks at it very closely.
  2. You can't rely on tools to fail when you make a mistake.
  3. Classifiers are very good at finding information, even if it's not the information you had in mind.

The code in question is 4_classify_wholebrain.py which reads in the processed data (saved in a numpy file) and classifies each dataset (with about 184K features and 400 observations) into one of 23 different classes (representing different tasks). The code was made publicly available before submitting the paper; while I have no way of knowing whether the reviewers have examined it, it's fair to say that even if they did, they would most likely not have caught this particular bug unless they were very eagle-eyed.  As it happens, a student here was trying to reproduce my analyses independently, and was finding much lower classification accuracies than the ones I had reported.  As he dug into my code, it became clear that this difference was driven by a (lazy, in hindsight) coding mistake on my part.

The original code can be viewed here - the snippet in question (cleaned up a bit) is:

skf=StratifiedKFold(labels,8)

if trainsvm:
    pred=N.zeros(len(labels))
    for train,test in skf:
        clf=LinearSVC()
        clf.fit(data[train],labels[train])
        pred[test]=clf.predict(data[test])


Pretty simple - it creates a crossvalidation object using sklearn, then loops through, fitting to the train folds and computing the predicted class for the test fold.  Running this, I got about 93% test accuracy on the multiclass problem; had I gotten 100% accuracy I would have been sure that there was a problem, but given that we have previously gotten around 80% for similar problems, I was not terribly shocked by the high accuracy. Here is the problem:

In [9]: data.shape
Out[9]: (182609, 400)

When I put the data into the numpy object, I had voxels as the first dimension, whereas for classification analysis one would usually put the observations in rows rather than columns.  Now, numpy is smart enough that when I give it the train list as an array index, it uses it as an index on the first dimension.  However, because of the transposition of the dimensions in the data, the effect was to classify voxels, rather than subjects:

In [10]: data[train].shape
Out[10]: (350, 400)

In [11]: data[test].shape
Out[11]: (50, 400)

When I fix this by using the proper data reference (as in the current revision of the code on the repo), then it looks as it should (i.e. all voxels included for the subjects in the train or test folds):

In [12]: data[:,train].T.shape
Out[12]: (350, 182609)


In [14]: data[:,test].T.shape
Out[14]: (50, 182609)

When I run this with the fixed code I get about 53% accuracy; still well above chance (remember that it's a 23-class problem), but much less than the 93% we had gotten previously.

It's worth noting that randomization tests with the flawed code showed the expected null distribution; the source of the information being used by the classifier is a bit of a mystery, but likely reflects the fact that the distance of the voxels in the matrix is related to their distance in space in the brain, and the labels were grouped together sequentially in the label file, such that they were correlated with physical distance in the brain and thus provided information that could drive the classification.

This is clearly a worst-case scenario for anyone who codes up their own analyses; the paper has already been submitted and you find an error that greatly changes the results. Fortunately, the exact level of classification accuracy is not central to the paper in question, but it's worrisome nonetheless.

What are the lessons to be learned here?  Most concretely, it's important to check the size of data structures whenever you are slicing arrays.  I was lazy in my coding of the crossvalidation loop, and I should have checked that the size of the dataset being fed into the classifier was what I expected it to be (the difference between 400 and 182609 would be pretty obvious).  It might have added an extra 30 seconds to my initial coding time but would have saved me from a huge headache and hours of time needed to rerun all of the analyses.

Second, sharing code is necessary but not sufficient for finding problems.  Someone could have grabbed my code and gotten exactly the same results that I got; only if they looked at the shape of the sliced arrays would they have noticed a problem.  I am becoming increasingly convinced that if you really want to believe a computational result, the strongest way to do that is to have an independent person try to replicate it without using your shared code.  Failing that, one really wants to have a validation dataset that one can feed into the program where you know exactly what the output should be; randomization of labels is one way of doing this (i.e., where the outcome should be chance) but you also want to do this with real signal as well.  Unfortunately this is not trivial for the kinds of analyses that we do, but perhaps some better data simulators would help make it easier.

Finally, there is a meta-point about talking openly about these kinds of errors. We know that they happen all the time, yet few people ever talk openly about their errors.  I hope that others will take my lead in talking openly about errors they have made so that people can learn from them and be more motivated to spend the extra time to write robust code.



Wednesday, January 16, 2013

Is reverse inference a fallacy? A comment on Hutzler

A new paper by Florian Hutzler has been published online at Neuroimage which claims to show that reverse inference is not as problematic as has been claimed in my previous publications (TICS, 2006; Neuron, 2010).  I had previously reviewed this paper for another journal (I signed my review so this is not a surprise), and I'm happy to see that some of my concerns about the paper were addressed in the version that was published at Neuroimage.  However, I still have one major concern about the general framing of the paper.

I would first like to be clear about what I said about reverse inference in my 2006 paper:

"It is crucial to note that this kind of ‘reverse inference’ is not deductively valid, but rather reflects the logical fallacy of affirming the consequent...However, cognitive neuroscience is generally interested in a mechanistic understanding of the neural processes that support cognition rather than the formulation of deductive laws. To this end, reverse inference might be useful in the discovery of interesting new facts about the underlying mechanisms. Indeed, philosophers have argued that this kind of reasoning (termed ‘abductive inference’ by Pierce [8]), is an essential tool for scientific discovery [9]."
Thus, while I did point out the degree to which reverse inference reflects a fallacy under deductive logic, I also pointed out that it could be potentially useful under other forms of reasoning; it's a bit of a stretch to go from this statement to using the term "reverse inference fallacy" which has started to pervade peer reviews.  This is unfortunate in my view, if only because authors must often think that I am the culprit!  (I assure you all that I would never use this phrase in a review.) The potential utility of reverse inference has been further cashed out in the Neurosynth project (Yarkoni et al, 2011).  I say all of this just to highlight the fact that I have never painted reverse inference as wholly fallacious, but rather have tried to highlight ways in which its limited utility can be quantified (e.g. through meta-analysis) or its power improved (e.g., through the use of machine learning methods).

The Hutzler paper applies reverse inference in a much more restrictive sense than it has usually been discussed, which he calls "task-specific functional specificity."  The idea is that given some task, one can compute (e.g., using meta-analysis) the reverse inference conditional on that task (which I had noted but not further explored in my 2006 paper).  I have no quibbles with the paper's analysis, and I think it nicely shows how reverse inference can be useful within a limited domain (in fact, Anthony Wagner and I made this point in 2004 in regard to left prefrontal function). My general concern is that the situation described in the Hutzler paper is fairly different from the one in which most reverse inference is performed.  Here is what I said in my initial review of his paper, which still holds for the published version:
If it is true that reverse inference is helpful within the context of a specific task, then that’s perfectly fine, except that in the wild reverse inference is rarely used within the same task.  In fact, it’s almost always used in task domains where one doesn’t know what to expect!  See my recent Neuron paper for examples of these kinds of reverse inferences; rarely does one see a reverse inference based on prior data from very similar tasks.  Thus, the paper basically makes my point for me by showing that the procedure is only effective in very specific cases which are outside of the standard way it is used.

In summary, while I agree with the analysis presented by Hutzler, I hope that readers will go beyond the title (which I think oversells the result) to see that it really shows the success of reverse inference in a very limited domain.

Sunday, December 16, 2012

The perils of leave-one-out crossvalidation for individual difference analyses

There is a common tendency of researchers in the neuroimaging field to use the term "prediction" to describe observed correlations.  This is problematic because the strength of a correlation does not necessarily imply that one can accurately predict the outcome of new (out-of-sample) observations.  Even if the underlying distributions are normally distributed, the observed correlation will generally overestimate the accuracy of predictions on out-of-sample observations due to overfitting (i.e., fitting the noise in addition to the signal).  In degenerate cases (e.g., when the observed correlation is driven by a single outlier), it is possible to observe a very strong in-sample correlation with almost zero predictive accuracy for out-of-sample observations.

The concept of crossvalidation provides a way out of this mess; by fitting the model to subsets of the data and examining predictive accuracy on the held-out samples, it's possible to directly assess the predictive accuracy of a particular statistical model.  This approach has become increasingly popular in the neuroimaging literature, which should be a welcome development.  However, crossvalidation in the context of regression analyses turns out to be very tricky, and some of the methods that are being used in the literature appear to be problematic.

One of the most common forms of crossvalidation is "leave-one-out" (LOO) in which the model is repeatedly refit leaving out a single observation and then used to derive a prediction for the left-out observation.  Within the machine learning literature, it is widely appreciated that LOO is a suboptimal method for cross-validation, as it gives estimates of the prediction error that are more variable than other forms of crossvalidation such as K-fold (in which the training/testing is performed after breaking the data into K groups, usually 5 to 10) or bootstrap; see Chapter 7 in Hastie et al. for a thorough discussion of this issue).

There is another problem with LOO that is specific to its use in regression.  We discovered this several years ago when we started trying to predict quantitative variables (such as individual learning rates) from fMRI data.  One thing that we always do when running any predictive analysis is to perform a randomization test to determine the distribution of performance when the relation between the data and the outcome is broken (e.g., by randomly shuffling the outcomes). In a regression analysis, what one expects to see in this randomized-label case is a zero correlation between the predicted and actual values.  However, what we actually saw was a very substantial negative bias in the correlation between predicted and actual values.  After lots of head-scratching this began to make sense as an effect of overfitting.  Imagine that you have a two-dimensional dataset where there is no true relation between X and Y, and you first fit a regression line to all of the data; the slope of this line should on average be zero, but will likely deviate from zero on each sample due to noise in the data (an effect that is bigger for smaller sample sizes).  Then, you drop out one of the datapoints and fit the line again; let's say you drop out one of the observations at the extreme end of the X range.  On average, this is going to have the effect of bringing the estimated regression line closer to zero than the full estimate (unless your left-out point was right at the mean of the Y distribution).  If you do this for all points, you can then see how this would result in a negative correlation between predicted and actual values when the true slope is zero, since this procedure will tend to pull the line towards zero for the extreme points.

To see an example of this fleshed out in an ipython notebook, visit http://nbviewer.ipython.org/4221361/ - the full code is also available at https://github.com/poldrack/regressioncv.  This code creates random X and Y values and then tests several different crossvalidation schemes to examine their effects on the resulting correlations between predicted and true values.  I ran this for a number of different samples sizes, and the results are shown in the figure below (NB: figure legend colors fixed from original post).



The best (i.e. least biased) performance is shown by the split half method.  Note that this is a special instance of split half crossvalidation with perfectly matched X distributions, because it has exactly the same values on the X variable for both halves.  The worst performance is seen for leave-one-out, which is highly biased for small N's but shows substantial bias even for very large N's.  Intermediate performance is seen when a balanced 8-fold crossvalidation scheme is used; the "hi" and "lo" versions of this are for two different balancing thresholds, where "hi" ensures fairly close matching of both X and Y distributions across folds whereas "lo" does not.  We have previously used balanced crossvalidation schemes on real fMRI data (Cohen et al., 2010) and found them to do a fairly good job of removing bias in the null distributions, but it's clear from these simulations that bias can still remain.

As an example using real data, I took the data from a paper entitled "Individual differences in nucleus accumbens activity to food and sexual images predict weight gain and sexual behavior" by Demos et al.  The title makes a strong predictive claim, but what the paper actually found was an observed correlation of r=0.37 between neural response and future weight gain.  I ripped the data points from their scatterplot using PlotDigitizer and performed a crossvalidation analysis using leave-one-out and 4-fold crossvalidation either with or without balancing of the X and Y distributions across folds (see code and data in the github repo).   The table below shows the predictive correlations obtained using each of the measures (the empirical null was obtained by resampling the data 500 times with random labels; the 95th percentile of this distribution is used as the significance cutoff):

CV methodr(pred,actual)r(pred,actual) with random labels95%ile
LOO 0.176395 -0.276977 0.154975
4-fold 0.183655 -0.148019 0.160695
Balanced 4-fold 0.194456 -0.066261 0.212925

This analysis shows that, rather than the 14% of variance implied by the observed correlation, one can at best predict about 4% of the variance in weight gain from brain activity; this result is not significant by a resampling test, though the LOO and 4-fold results, while weaker numerically, are significant compared to their respective empirical null distributions. (Note that a small amount of noise was likely introduced by the data-grabbing method, so the results with the real data might be a bit better.)

UPDATE: As noted in the comments, the negative bias can be largely overcome by fixing the intercept to zero in the linear regression used for prediction.  Here are the results obtained using the zero-intercept model on the Demos et al. data:

CV methodr(pred,actual)r(pred,actual) with random labels95%ile
LOO0.258-0.0670.189
4-fold0.263-0.0550.192
Balanced 4-fold0.256-0.0370.213



This gets us up to about 7% variance accounted for by the predicted model, or about half of that implied by the in-sample correlation, and now the correlation is significant by all CV methods.

Take-home messages:
  • Observed correlation is generally larger than predictive accuracy for out-of-sample observations, such that one should not use the term "predict" in the context of correlations.
  • Cross-validation for predicting individual differences in fMRI analysis is tricky.  
  • Leave-one-out should probably be avoided in favor of balanced k-fold schemes
  • One should always run simulations of any classifier analysis stream using randomized labels in order to assess the potential bias of the classifier.  This means running the entire data processing stream on each random draw, since bias can occur at various points in the stream.


PS: One point raised in discussing this result with some statisticians is that it may reflect the fact that correlation is not the best measure of the match between predicted and actual outcomes.  If someone has a chance to take my code and play with some alternative measures, please post the results to the comments section here, as I don't have time to try it out right now.

PPS: I would be very interested to see how this extends to high-dimensional data like those generally used in fMRI.  I know that the bias effect occurs in that context given that this is how we discovered it, but I have not had a chance to simulate its effects.




Tuesday, May 15, 2012

Obesity, health, and Gary Taubes

I recently posted a link on Facebook to Gary Taubes' article about why the campaign to stop America's obesity crisis keeps failing, and my friend Scott raised the following issue:

Taubes may or may not be on to something. But, he comes off like an Intelligent Design guy, "The experts are wrong, and they can't handle the Truth that I'm bringing."

I agree that Taubes' writings sometimes have the feel of a crazy outsider fighting against the establishment.  However, both my personal experience (as well as those of a number of friends) and my (non-expert) reading of the literature both suggest that Taubes is largely right on in terms of his critique of the standard dogma regarding weight loss, food, and health.

First, the testimonial.  As I noted in a previous post, it was Taubes' writings that were largely responsible for pushing me towards the low-carb way of eating that I have followed for more than a year now.  After cutting carbs way down (except for my daily dose of dark chocolate, which is non-negotiable), I lost 20 pounds of fat and have kept it off, without any sense that I am being deprived; I basically eat whatever I want whenever I want, as long as it's real food and paleo-friendly (i.e., avoiding refined sugar, grains, and seed oils).  Most important, I feel great eating this way; in particular, whereas I used to get serious hunger pangs and energy dips 3-4 hours after eating, I can now easily fast for 24 hours without feeling particularly hungry.  My wife Jen has also had an interesting experience on this diet.  She has been able to maintain her weight or lose weight while never feeling hungry, whereas on our old carb-heavy vegetarian diet she was only able to lose weight through radical caloric restriction that left her constantly famished.  Similarly, a number of friends and family members have found that they were able to lose a substantial amount of weight after reducing carbs, while still feeling like they were able to eat to satiety.  I have *never* heard anyone say that they went on a low-carb diet and gained weight; more often, I have heard from people that they went on a low-carb diet and lost weight but then were afraid that all of the saturated fat was going to cause them to have a heart attack any day.

This brings us to one of the central points of Taubes' writings, which is that the standard story about what comprises a healthy diet, namely the link between heart disease, cholesterol, and saturated fat, is just plain wrong.  If you want a good overview of his general narrative without reading the books, I would suggest three NY Times pieces: one on the relation between dietary fat and disease from 2002, a blistering critique of epidemiological research from 2007, and his piece "Is Sugar Toxic" from 2011.

There are a lot of claims in the Taubes books, and I have not looked into all of them.  However, to the degree that I have looked into the claims that I found most important and relevant to my own diet, I have found them to all have fairly compelling scientific bases. The most important regards the relation between heart disease, cholesterol, and saturated fat.  It is amazing how the supposed unhealthiness of dietary cholesterol and saturated fat has become a "fact" that is repeated almost reflexively (e.g., most recently I encountered it in Tyler Cowen's "An Economist Gets Lunch").  I think that in part it is due to the visual similarity of saturated fat in meat and the plaques that are seen in atherosclerosis; it's just too easy to believe that the saturated fat that we eat is "clogging our arteries."  The data appear to say otherwise.  First, it has been known since 1950 that serum cholesterol bears little relation to dietary cholesterol; the mechanisms behind this are laid out nicely in Peter Attia's recent series on cholesterol.  Second, a large recent meta-analysis (including data from more than 347,000 individuals across 21 published studies) found no relation between saturated fat intake and heart disease or stroke.  Similarly, a recent Cochrane Collaborative meta-analysis of intervention studies showed that there was no significant reduction of total or cardiovascular mortality due to changes in dietary fat.  Although I think Taubes is correct in his arguments that epidemiological studies are hugely problematic (which I will discuss some other time), I trust these large meta-analyses much more than I trust any individual study (e.g., the China Study), especially when they show no effect (given all of the biases towards publishing positive effects).  I have also put my money where my mouth is: I now eat a high-fat diet including full-fat yogurt, eggs, and bacon almost every morning. It look me several months to stop craving sugar, but I'm now perfectly happy to a dinner without dessert, and in fact I no longer have a taste for foods that are extremely sugary.

Another of Taubes' main assertions is that obesity is caused primarily by carbohydrates (fleshed out in his book Why We Get Fat).  My feeling here is that obesity is an incredibly complex problem that involves both the body and the brain, and any story that tries to simplify it to a single component of our environment is bound to be wrong.  That said, it's clear to me from the person experience described above that the "calories in, calories out" story is wrong, and that it does matter a lot what one eats, not just how much.  There has recently been a big argument in the blogosphere recently between Stephan Guyenet and Taubes over the relative importance of peripheral factors (e.g., insulin's effects on fat storage) versus neural factors (e.g., the role of satiety hormones and reward pathways); both of them have staked out strong positions, and I think that the truth is likely to fall somewhere in the middle (as usual).  As a neuroscientist I clearly think that the brain plays an important role; I won't talk more about that here, maybe some time soon.  I have not dug very deeply into the science of feeding trials in animals, in part because my reading of several summaries of this work suggests that the details can be very tricky but very important. In particular, without a great deal of control over exactly what kinds of nutrients are in the food, it can be very difficult to make any conclusions from the data. There are however some individual studies in humans that do provide support for Taubes' claims.  For example, the A to Z Weight Loss Study showed that subjects assigned to the Atkins diet lost more weight and had better metabolic outcomes than people assigned to low-fat/high-carb diets like the Ornish Diet.   It's just one study, and it would be good to see more, but this along with my personal experience is enough to convince me.

Later in the discussion that I mentioned above, Scott made the following additional observation:

But the question remains, why doesn't the science on the cellular level ever reach the public health scientific community? I'm usually very skeptical when some sort of conspiracy is trotted out to explain the lack of uptake.

There are a lot of answers to this, which Taubes goes into great detail to discuss in his books.  But the general problem is that science itself can be a slow-moving ship; when it gets drawn well off course (as it appears to have been by the anti-fat arguments of Keys and others), it can take a long time to get back, and even longer for that new scientific knowledge to get translated into medical education and practice.  What is most striking to me is how studies whose data seem clearly inconsistent with the standard view are often presented in a way that suggests that they support the view, often by picking and choosing specific conditions.  Taubes gives numerous examples of this, as does Denise Minger's detailed analysis of the China Study.  I've also noticed it on a number of occasions when reading papers in this literature.  Thus, unless one is reading the papers closely (or following others who do this), it can be easy to continue to think that the standard model remains valid; you just can't take the abstract (or even the results section) at face value.  However, the fact that the rise in obesity has occurred alongside declining fat intake (coupled with increasing carb intake) over the last 40 years makes it pretty clear to me that the standard theory is just plain wrong, and that the carb theory is a viable alternative that needs to be studied more intently.  Unfortunately, many of the thought leaders in this area continue to expound solutions based on the "calories-in/calories-out" and low-fat ideas that got us here in the first place.

Wednesday, April 25, 2012

Things I like to do in Beijing

Several friends have asked for suggestions about things to do while in Beijing for the OHBM meeting this June.  I don't have any particular wisdom, but having visited several times I thought that I would share a few of my favorite things (along with photos from some of our past trips).  I'm not including many of the obvious attractions (Forbidden City, Summer Palace, Olympic Park) because I figure those will be in every guide book.



798 Art Complex: an amazing art complex created from an old factory complex.  If you like modern art you can easily spend more than 1/2 a day there. There is a great noodle shop tucked away in the middle of the complex, and also at least one really nice cafe.



 






Hutongs: These are the old neighborhoods in the center of the city.  Some of them are very touristy, but if you walk a few blocks off the main street you can find some streets that feel pretty far from touristy. I would particularly recommend the area around Heizhima Hutong, where these photos were taken.

 
 








Drum Tower: This was built in 1272 and served as the official timepiece of the Chinese government until 1924.  If you show up at the right time, you can see an awesome drumming show. 





Great Wall:  We visited the Great Wall at Badaling in 2005, which is apparently the most touristy place to go but also relatively close to Beijing.  Go early in the day, to avoid both crowds and heat.  Perhaps the best part of visiting at Badaling is that there is a roller coaster that can take you down from the top.  


 





Eating

Roasted duck hearts at Quanjude
We have had a lot of wonderful meals in Beijing, both as fish-etarians on our first two visits and as omnivores on our most recent visit. Prepare to eat well, but also be prepared to have your sensibilities challenged.  A few highlights are:

Roast duck:  As recently reformed vegetarians, we spent our first two visits to China without trying "Peking Duck" (or, as they call it in Beijing, "roast duck").  On our last visit we had it at Quanjude and it was pretty awesome.  We had the full on roast duck experience, including "duck breast" (which is basically just fat and skin) and roasted duck hearts.  A must-have. (NB: If you order the duck hearts, they come with a bowl of flaming liquid.  Apparently you are not supposed to actually dip the heart into the liquid, as I did.)
Spicy snails at Spicy Grandma restaurant

Sichuan food: Our friends in China are largely from Sichuan province, and thus we often end up eating at Sichuan restaurants.  The Sichuan peppercorn has an amazing numbing quality.  Also be sure to try the Sichuan hot pot, which is like a very spicy version of shabu shabu.  I would suggest bringing a significant ration of Pepto Bismol, as the western gut starts to ache after a few days of this kind of spicy food.  But it is so worth the burn. 






Yunnan food:  One of the  most amazing meals we had was at the Rainbow Restaurant in the Beijing Sun Palace Hotel. The greeters are dressed in traditional Yunnan dress, and the food is absolutely amazing with a heavy focus on mushrooms.  
A dish that contained "smelly tofu" - actually really tasty


Grilled matsutake musrooms at Rainbow