Tuesday, March 6, 2012

Skeletons in the closet

 As someone who has thrown lots of stones in recent years, it's easy to forget that anyone who publishes enough will end up with some skeletons in their closet.  I was reminded of that fact today, when Dorothy Bishop posted a detailed analysis of a paper that was published in 2003 on which I am a coauthor.

This paper studied a set of children diagnosed with dyslexia who were scanned before and after treatment with the Fast ForWord training program.  The results showed improved language and reading function, which were associated with changes in brain activation. 

Dorothy notes four major problems with the study:
  • There was no dyslexic control group; thus, we don't know whether any improvements over time were specific to the treatment, or would have occurred with a control treatment or even without any treatment.
  • The brain imaging data were thresholded using an uncorrected threshold.
  • One of the main conclusions (the "normalization" of activation following training") is not supported by the necessary interaction statistic, but rather by a visual comparison of maps.
  • The correlation between changes in language scores and activation was reported for only one of the many measures, and it appeared to have been driven by outliers.
Looking back at the paper, I see that Dorothy is absolutely right on each of these points.  In defense of my coauthors, I would note that points 2-4 were basically standard practice in fMRI analysis 10 years ago (and still crop up fairly often today).  Ironically,  I raised two of of these issues in my recent paper for the special issue of Neuroimage celebrating the 20th anniversary of fMRI, in talking about the need for increased methodological rigor:

Foremost, I hope that in the next 20 years the field of cognitive neuroscience will increase the rigor with which it applies neuroimaging methods. The recent debates about circularity and “voodoo correlations” ( [Kriegeskorte et al., 2009] and [Vul et al., 2009]) have highlighted the need for increased care regarding analytic methods. Consideration of similar debates in genetics and clinical trials led (Ioannidis, 2005) to outline a number of factors that may contribute to increased levels of spurious results in any scientific field, and the degree to which many of these apply to fMRI research is rather sobering:
•small sample sizes
•small effect sizes
•large number of tested effects
•flexibilty in designs, definitions, outcomes, and analysis methods
•being a “hot” scientific field
Some simple methodological improvements could make a big difference. First, the field needs to agree that inference based on uncorrected statistical results is not acceptable (cf. Bennett et al., 2009). Many researchers have digested this important fact, but it is still common to see results presented at thresholds such as uncorrected p < .005. Because such uncorrected thresholds do not adapt to the data (e.g., the number of voxels tests or their spatial smoothness), they are certain to be invalid in almost every situation (potentially being either overly liberal or overly conservative). As an example, I took the fMRI data from Tom et al. (2007), and created a random “individual difference” variable. Thus, there should be no correlations observed other than Type I errors. However, thresholding at uncorrected p < .001 and a minimum cluster size of 25 voxels (a common heuristic threshold) showed a significant region near the amygdala; Fig. 1 shows this region along with a plot of the “beautiful” (but artifactual) correlation between activation and the random behavioral variable. This activation was not present when using a corrected statistic. A similar point was made in a more humorous way by Bennett et al. (2010), who scanned a dead salmon being presented with a social cognition task and found activation when using an uncorrected threshold. There are now a number of well-established methods for multiple comparisons correction (Poldrack et al., 2011), such that there is absolutely no excuse to present results at uncorrected thresholds. The most common reason for failing to use rigorous corrections for multiple tests is that with smaller samples these methods are highly conservative, and thus result in a high rate of false negatives. This is certainly a problem, but I don't think that the answer is to present uncorrected results; rather, the answer is to ensure that one's sample is large enough to provide sufficient statistical power to find the effects of interest.
Second, I have become increasingly concerned about the use of “small volume corrections” to address the multiple testing problem. The use of a priori masks to constrain statistical testing is perfectly legitimate, but one often gets the feeling that the masks used for small volume correction were chosen after seeing the initial results (perhaps after a whole-brain corrected analysis was not significant). In such a case, any inferences based on these corrections are circular and the statistics are useless. Researchers who plan to use small volume corrections in their analysis should formulate a specific analysis plan prior to any analyses, and only use small volume corrections that were explicitly planned a priori. This sounds like a remedial lesson in basic statistics, but unfortunately it seems to be regularly forgotten by researchers in the field.
Third, the field needs to move toward the use of more robust methods for statistical inference (e.g., Huber, 2004). In particular, analyses of correlations between activation and behavior across subjects are highly susceptible to the influence of outlier subjects, especially with small sample sizes. Robust statistical methods can ensure that the results are not overly influenced by these outliers, either by reducing the effect of outlier datapoints (e.g., robust regression using iteratively reweighted least squares) or by separately modeling data points that fall too far outside of the rest of the sample (e.g., mixture modeling). Robust tools for fMRI group analysis are increasingly available, both as part of standard software packages (such as the “outlier detection” technique implemented in FSL: Woolrich, 2008) and as add-on toolboxes (Wager et al., 2005). Given the frequency with which outliers are observed in group fMRI data, these methods should become standard in the field. However, it's also important to remember that they are not a panacea, and that it remains important to apply sufficient quality control to statistical results, in order to understand the degree to which one's results reflect generalizeable patterns versus statistical figments.
It should be clear from these comments that my faith in the results of any study that uses such problematic methods (as the Temple et al. study did) is relatively weak.  I personally have learned my lesson and our lab now does its best to adhere to these more rigorous standards, even when they mean that a study sometimes ends up being unpublishable.   I can only hope that others will join me.

      Thursday, February 9, 2012

      Quitting cable

      I was inspired by Nathan's post at Flowing Data to say a bit about how our experiment with giving up cable TV is going. Back in September, we turned off our U-Verse subscription, sent back the DVR, and started getting our TV solely from the computer (we use a Mac Mini as our media center PC).  Here is our experience so far:

      We watch a lot less TV.  Our TV routine has now morphed from watching 2-3 hours per night into watching a single show every night (recent favorites are The Layover and Top Chef, with Colbert Report as our fallback).  In its place we are reading a lot more; in fact, much of the money that we are saving on cable is probably flowing to the Kindle store at Amazon.  However, we have also recently started using the Austin Public Library's ebook lending service which is a great way to save on ebooks.  I've also been playing the guitar more often.

      Sometimes you really want live TV.  The one problem with getting everything from the web is that it's often hard to find a good live stream; we had this problem on new year's eve.  To solve this, I recently installed a solution to allow us to view live broadcast TV from the computer, using an
      Elgato EyeTV One Computer TV Tuner with a Mohu Leaf HDTV antenna.  With this slick combination we are able to get 13 channels of over-the-air HDTV for free.  The EyeTV software is really nice; it has good DVR functionality and an integrated TV Guide.  It's very much like having cable with 13 channels, except that the DVR functions are much better than any set-top DVR we ever had.

      Hulu Plus  > Netflix.  We have found that Hulu Plus meets our TV viewing needs quite well.  Sure we have to watch some commercials, but we are usually able to get new shows the next day after they air, and the selection is pretty good.  I tried a free trial of Netflix online, but we have not found that it has much to offer us, except for an occasional movie.  However, we watch movies pretty rarely, and so it probably makes more sense for us to just buy them from iTunes.  For shows that are not available on the web or via Hulu (e.g., The Layover), we buy them from iTunes as well.  It's not cheap but we still come out ahead in the long run.

      Media center software sucks.  We tried using both Plex and Boxee on the mac mini, but gave up on both after too many things just didn't work; in particular, the Hulu integration on Plex was really frustrating, as it seems like it should work but then it never quite does.  Now we just watch Hulu content through a web browser, live/recorded TV through EyeTV, and iTunes content through iTunes.  The main drawback of this setup is that we can't get remote functionality that works seamlessly across all these different interfaces, but that's not been a problem.

      Overall I would rate this experiment as a success and would definitely recommend giving up cable.

      Saturday, December 31, 2011

      2011 in review

      In the spirit of Chris Guillebeau's Annual Review, I decided to take a few minutes this morning to review what worked well and what didn't 2011 and look forward to 2012.

      Goals:

      Last year I set three goals for 2011.  Here's how they fared.

      FAIL: Work toward a travel moratorium for 2012.  This started with me responding to all requests with "I'm sorry, but I'm not traveling at all for work in 2012."  At some point that became untenable, and the floodgates opened.  At this point, I am still planning to travel less in 2012 than in 2011, but will still probably take 8-10 trips.

      SUCCESS: Improve climbing skills well enough to lead climb. My climbing skills have improved enormously over the last year, and in the summer I began lead climbing at the rock gym, where I can now lead a number of routes in the 5.9 range.  I was not able to lead outside, mostly because the weather in December did not cooperate, but I plan to do so very soon.

      SUCCESS: No new web projects. I only purchased one new domain name this year, and that was for a project that had been hatched in 2010.  We have instead focused heavily on our existing projects, particularly openfmri.org.

      Here are my goals for 2012:

      Improve my posture.  Some nagging neck and back issues this year have highlighted the need to improve my posture.  Who knows, I might even get some mental benefits from it as well.

      Improve my code management.  In the last year I have started integrating source code management (using git) into my workflow (see my github repo for a tour of some of my adventures during the last year).  However, it still has not become a habit for me during everyday coding.

      Exercise on every trip.  One of the reasons that I find travel so disruptive is that it interferes with my fitness routine.  I carry my yoga mat on nearly every trip, and this year I did a fairly good job of exercising while on the road, but I was not very consistent.  Next year I plan to make sure that I get some exercise on every trip, even if it's just some burpees and squats in the hotel room.

      I hope it's not true that making these goals public will make them harder to achieve!

      Stats

      Countries visited: 7
      Miles flown: 76,162
      Talks given: 12
      Papers published: 14
      Grants funded: 2

      Property crimes (committed against me, not by me): 2


      Best meals:
      1. Tasting menu at Congress
      2. Lunch at Les Arcenuax, Marseille
      3. Tie between Franklin BBQ and JMueller BBQ
      4. Tasting menu at Uchi (the meal that sealed our transition to full-blown carnivores)

      Tuesday, October 4, 2011

      NYT Letter to the Editor: The uncut version

      The NY Times has now printed our letter to the editor regarding the Lindstrom article.  However, the published version is an edited and shortened version of our original letter, which I am posting here for the record.


      Dear Editor,
      The Op-Ed “You Love Your iPhone, Literally” by Martin Lindstrom purports to show, using brain imaging, that our attachment to digital devices, reflects not addiction but instead the same kind of emotion that we feel for human loved ones. However, the evidence the author presents does not show this.  The region that he points to as being “associated with feelings of love and compassion” (the insular cortex) is a brain region that is active in as many as one third of all brain imaging studies.  Further, in studies of decision making the insula is more often associated with negative than positive emotions.  The kind of reasoning that Lindstrom uses is well known to be flawed, because there is rarely a one-to-one mapping between any brain region and a single mental state; insula activity could reflect one or more of several psychological processes. This same point was made by some of us regarding a similar Op-Ed piece in 2007.
      We are disappointed that the Times has published extravagant claims based on scientific data that have not been subjected to the standard scientific review process, especially considering how often its pages exhort policy makers to pay more attention to peer-reviewed scientific evidence and disregard specious claims.
      Sincerely,

      Russell Poldrack, Ph.D., University of Texas at Austin
      Geoffrey K Aguirre, M.D., Ph.D., University of Pennsylvania
      Adam Aron, Ph.D., University of California at San Diego
      Lisa Feldman Barrett, Ph.D., Northeastern University
      Mark G. Baxter, Ph.D., Mount Sinai School of Medicine
      Susan Bookheimer, Ph.D., University of California at Los Angeles
      Colin Camerer, Ph.D., California Institute of Technology
      McKell Carter, Ph.D., Duke University
      Christopher Chabris, Ph.D., Union College
      Molly Crockett, Ph.D., University of Zurich, Switzerland
      Nathaniel Daw, Ph.D., New York University
      Paul Downing, Ph.D., University of Bangor, Wales, UK
      Russell Epstein, Ph.D., University of Pennsylvania
      Michael Frank, Ph.D., Brown University
      Janet Frick, Ph.D., University of Georgia
      Paul Glimcher, Ph.D., New York University
      Tom Hartley, Ph.D., University of York, UK
      Benjamin Hayden, Ph.D., University of Rochester
      Hauke R. Heekeren, M.D., Freie Universität Berlin, Germany
      Simon Hjerrild, M.D., University of Aarhus, Denmark
      Scott Huettel, Ph.D., Duke University
      Nancy Kanwisher, Ph.D., Massachusetts Institute of Technology
      Brian Knutson, Ph.D., Stanford University
      John Kubie, Ph.D., SUNY Downstate Medical Center
      Michael V. Lombardo, Ph.D., University of Cambridge, UK
      Ken Norman, Ph.D., Princeton University
      Olivier Oullier, Ph.D., Aix-Marseille University, France
      Steven Petersen, Ph.D., Washington University
      Elizabeth Phelps, Ph.D., New York University
      Rajeev Raizada, Ph.D., Cornell University
      Antonio Rangel, Ph.D., California Institute of Technology
      Peter B. Reiner, Ph.D., University of British Columbia, Canada
      Gregory Samanez-Larkin, Ph.D., Vanderbilt University
      Geoff Schoenbaum, M.D., Ph.D., University of Maryland
      Daphna Shohamy, Ph.D., Columbia University
      Jon Simons, Ph.D., University of Cambridge, UK
      Peter Sokol-Hessner, Ph.D., California Institute of Technology
      David Somers, Ph.D., Boston University
      Damian Stanley, Ph.D., California Institute of Technology
      John Van Horn, Ph.D., University of California at Los Angeles
      Bradley Voytek, Ph.D., University of California, San Francisco
      Anthony Wagner, PhD, Stanford University.
      Daniel Willingham, Ph.D., University of Virginia
      Tal Yarkoni, Ph.D., University of Colorado Boulder
      Jeff Zacks, Ph.D., Washington University
      Jamil Zaki, Ph.D., Harvard University

      Monday, October 3, 2011

      Signers of letter to the editor of the New York Times

      A letter has been submitted to the editor of the NY Times regarding the outrageous Op-Ed piece by Martin Lindstrom.  (Once it is published I will add a link here.)  Because the NY Times will not allow a long list of signers on a letter, I am attaching here a list of all of the individuals who contributed to and signed the letter.  If you would like to add your name in support of the letter, please do so in the comments section.

      Russell Poldrack, Ph.D., University of Texas at Austin
      Geoffrey K Aguirre, M.D., Ph.D., University of Pennsylvania
      Adam Aron, Ph.D., University of California at San Diego
      Lisa Feldman Barrett, Ph.D., Northeastern University
      Mark G. Baxter, Ph.D., Mount Sinai School of Medicine
      Susan Bookheimer, Ph.D., University of California at Los Angeles
      Colin Camerer, Ph.D., California Institute of Technology
      McKell Carter, Ph.D., Duke University
      Christopher Chabris, Ph.D., Union College
      Molly Crockett, Ph.D., University of Zurich, Switzerland
      Nathaniel Daw, Ph.D., New York University
      Paul Downing, Ph.D., University of Bangor, Wales, UK
      Russell Epstein, Ph.D., University of Pennsylvania
      Michael Frank, Ph.D., Brown University
      Janet Frick, Ph.D., University of Georgia
      Paul Glimcher, Ph.D., New York University
      Tom Hartley, Ph.D., University of York, UK
      Benjamin Hayden, Ph.D., University of Rochester
      Hauke R. Heekeren, M.D., Freie Universität Berlin, Germany
      Simon Hjerrild, M.D., University of Aarhus, Denmark
      Scott Huettel, Ph.D., Duke University
      Nancy Kanwisher, Ph.D., Massachusetts Institute of Technology
      Brian Knutson, Ph.D., Stanford University
      John Kubie, Ph.D., SUNY Downstate Medical Center
      Michael V. Lombardo, Ph.D., University of Cambridge, UK
      Ken Norman, Ph.D., Princeton University
      Olivier Oullier, Ph.D., Aix-Marseille University, France
      Steven Petersen, Ph.D., Washington University
      Elizabeth Phelps, Ph.D., New York University
      Rajeev Raizada, Ph.D., Cornell University
      Antonio Rangel, Ph.D., California Institute of Technology
      Peter B. Reiner, Ph.D., University of British Columbia, Canada
      Gregory Samanez-Larkin, Ph.D., Vanderbilt University
      Geoff Schoenbaum, M.D., Ph.D., University of Maryland
      Daphna Shohamy, Ph.D., Columbia University
      Jon Simons, Ph.D., University of Cambridge, UK
      Peter Sokol-Hessner, Ph.D., California Institute of Technology
      David Somers, Ph.D., Boston University
      Damian Stanley, Ph.D., California Institute of Technology
      John Van Horn, Ph.D., University of California at Los Angeles
      Bradley Voytek, Ph.D., University of California, San Francisco
      Anthony Wagner, PhD, Stanford University.
      Daniel Willingham, Ph.D., University of Virginia
      Tal Yarkoni, Ph.D., University of Colorado Boulder
      Jeff Zacks, Ph.D., Washington University
      Jamil Zaki, Ph.D., Harvard University

      Saturday, October 1, 2011

      NYT Op-Ed + fMRI = complete crap

      Many of you may remember the controversy that arose a few years back when the NY Times published an Op-Ed titled "This is your brain on politics" by Marco Iacoboni and colleagues.  This steaming pile of shoddy reverse inferences inspired a group of us to write a letter to the editor, published online.  Well, the NYT editorial page is at it again, this time with a piece by self-proclaimed neuromarketer Martin Lindstrom, titled "You love your iPhone, literally" (h/t Raj Raizada for pointing me to it).  The argument of the article is that rather than our feelings about iphones reflecting something like an addiction driven by dopamine (which I have argued for in the past), our feelings about our digital devices instead reflect true love, based on fMRI:

      But most striking of all was the flurry of activation in the insular cortex of the brain, which is associated with feelings of love and compassion. The subjects’ brains responded to the sound of their phones as they would respond to the presence or proximity of a girlfriend, boyfriend or family member.
      In short, the subjects didn’t demonstrate the classic brain-based signs of addiction. Instead, they loved their iPhones.
      Insular cortex may well be associated with feelings of love and compassion, but this hardly proves that we are in love with our iPhones.  In Tal Yarkoni's recent paper in Nature Methods, we found that the anterior insula was one of the most highly activated part of the brain, showing activation in nearly 1/3 of all imaging studies!  Further, the well-known studies of love by Helen Fisher and colleagues don't even show activation in the insula related to love, but instead in classic reward system areas.  So far as I can tell, this particular reverse inference was simply fabricated from whole cloth.  I would have hoped that the NY Times would have learned its lesson from the last episode.

      Friday, July 1, 2011

      My analysis of OHBM 2011 abstracts

      As the past chair of the Organization for Human Brain Mapping (which just ended its 2011 meeting in Quebec City), I was tasked with giving the "Meeting Highlights" talk which traditionally closes the meeting.  It's a pretty daunting challenge to summarize an entire meeting with such little time for preparation, so I took the tack of doing lots of mining on the full text of the abstracts prior to the meeting.  My entire wrap-up talk is available here; below I present the main results of my text mining, along with some additional analyses that didn't make it into the talk.

      Overall meeting stats:

      Number of Abstracts 2230
      Number of Unique Authors 7622
      Mean # of abstracts/author 1.64 (max=33)
      Mean # of authors/abstract 5.6 (max=41)

      Authorship distribution:

      The OHBM is known for being an international organization, and the authorship data confirm this.  In order to visualize the authorship data, I used the Google Maps API to identify the latitude/longitude for each affiliation in the authorship list. This was successful for more than 90% of the abstracts.  These latitude/longitude values were uploaded into Google Fusion Tables, from which I exported a KML file  (available here) which I then opened in Google Earth.  (That's a lot of Google!)

      Using Google Earth I then created a tour that circled the globe, showing all of the author locations on a path from Quebec City to Beijing (location of the 2012 meeting).  Here is the video:



      Each red pin represents the location of an author at the meeting.

      Authorship networks:

      Using the abstracts I created a coauthorship network and did some basic analyses on this network (using the Networkx toolbox in Python and the Network Workbench).  The code and an anonymized version of the graph (in graphml format) are available via github.  Here is an overall view of the network:


      This shows one giant connected component with 4600 authors (60.3%), along with a large number of much smaller components (the second largest component had 103 authors).  Focusing in on the giant component, here is the spring-embedded visualization:



      Here are the network statistics:

      Clustering coefficient 0.88
      Average degree 8.10
      Average shortest path length (giant component only) 6.96
      Maximum shortest path length 18
      Modularity (giant component only) 0.92


      Here is the degree distribution plotted in log space, with a degree distribution for a matched random graph for comparison:


      The degree distribution has a long tail compared to the random network, which is what one would expect from this kind of network (for background on this kind of analysis, see Mark Newman's paper The structure of scientific collaboration networks).

      Using PageRank centrality, I identified the 10 most central authors in this network (listed with number of abstracts and centrality value):
      1. Paul Thompson (33 abstracts: 0.002020)
      2. Vince Calhoun (21 abstracts: 0.001816)
      3. Arno Villringer (23 abstracts: 0.001756)
      4. Arthur Toga (30 abstracts: 0.001625)
      5. Yong He (19 abstracts: 0.001416)
      6. Peter Fox (26 abstracts: 0.001381)
      7. Michael Milham (24 abstracts: 0.001340)
      8. Alan Evans (16 abstracts: 0.001318)
      9. Robert Turner (23 abstracts: 0.001292)
      10. Daniel Margulies (13 abstracts: 0.001194)
      Content analysis:


      Using the full text from the articles, I created several tag clouds (using Wordle) to show different aspects of the content.  The first was created from the entire abstract text after filtering out standard stop words along with anatomical regions and author names.


      The second was created using a count of all anatomical terms (from the PubBrain anatomical lexicon):


      The third was created using a count of all of the terms in the Cognitive Atlas lexicon of mental concepts:


      These tag clouds give a good overview of the major topics at the meeting.


      If you have other ideas for mining of these data, let me know and I'll give it a try. I have also done topic modeling using latent Dirichlet allocation, and may get around to writing about that in the future.