Data and Blogging

Ah Friday, a great day to do final edits on my and Julie's paper for Midwest – also a good time for a quick blog post with a somewhat misleading title as the two subjects refer to two seperate links. First, via Freakonomics, is a competition for the Fraser Institute to have them collect data.  … Read more

End of the week blogging

The summer is rapidly approaching its end as many of us are preparing for the classes we are teaching, quickly completing those papers for the upcoming APSA conference, and finishing up any summer projects (or finding ways to push back those deadlines).  As such, the blogging here has slowed down a bit while the Dark … Read more

Data collection using Web-based Forms

Thanks to a comment on an earlier post, Stephen Haptonstahl answered some of my questions and technical misgivings I had about setting up a larger user interface for collecting data via a webpage.  Specifically, he has an article in the Political Methodologist‘s from 2008 (the specific issue can be found here, starts on page 12) … Read more

Manual Data Collection in the age of Computers

I am beginning a new data collection project that requires the manual coding of data collected from various sources in print and online.  As I start this project, I am tasked with how to build a master record of all the data I collect in the process.  I have worked on projects that used extensive paper coding forms that were later filed away only to be retrieved when appropriate.  This serves as a safeguard to both checking original coding decision, errors in the database, and any other information the coders found while researching the topic at hand.  Alternatively, other projects had an evolving excel spreadsheet that itself was the master record – duplicate copies served as a safeguard against accidents while the final form only existed when the researchers stopped coding.  Finally, projects that are entirely automated tend to generate their own database that the researchers can then use to extract useful information into a final dataset for examination.

For this project, I decided to make a separate, master database that then will be used to generate data sets as needed.  Any changes will be documented in the master set and stripping out coding variables (last updated, side notes, etc.) that are otherwise irrelevant to people using the data.  Thus, the final versions will be tab delimited for general consumption, produce in Stata, and a R version while the master remains in a different format for official changes.

More after the jump; those interested, there are screenshots and I attempt to elicit ideas for more efficient mechanisms to collect and store data…

Read more

A Few Non-Connected Thoughts and Links

The three of us, along with Ray Carman, traveled to New York City for the weekend to enjoy a few hours of Eddie Izzard performing at Radio City Music Hall for this current "Stripped" tour.  As such, the trip is still fresh in my mind as I return to work on a few projects involving … Read more

Did Data Kill Theory?

Thanks to Geoff McGovern for pointing us toward a fascinating essay in Wired.  Chris Anderson posits that the accessibility of information has vaulted us into what he calls the Petrabyte Age, in which

information is not a matter of simple three- and four-dimensional
taxonomy and order but of dimensionally agnostic statistics. It calls
for an entirely different approach, one that requires us to lose the
tether of data as something that can be visualized in its totality. It
forces us to view data mathematically first and establish a context for
it later.

Given how much data is readily available, Anderson continues, "[w]e can
analyze the data without hypotheses about what it might show."  The
scientific method encourages us to explain what we know about the world
and make greater generalizations about the rest of it that we have not
observed; but if we can observe everything, essentially, it seems that
generalizations are no longer necessary.  We don’t need to guess about
what the world might look like, because an hour in front of the
computer can tell us. 

More after the jump.

Read more

Something to keep an eye out for: Google Data

Wired reports that Google plans to release, soon, a framework for hosting, storing, and distributing large or frequently used data.  The Project, Palimpsest, will pay the fees to both ship the data (by sending users a 3TB hard drive to download the data) and for hosting.  This, if applicable for political science scholars, not only … Read more

When Form can Overwhelm Content

I am not a visually oriented person or, more appropriately, I am less than stellar at design.  This may not be a surprise to anyone that has seen my attempts to assemble a wardrobe, but this is also true in the sense of organizing information – whether it is in a paper, on a poster, … Read more

Presenting Results over at ELS

Christopher Zorn over at the Empirical Legal Studies blog is doing a series on presenting results in an understandable and useful manner.  Given my own sins in generating tables, this should be an enlightening series to follow.  Michael A. AllenMichael A. Allen is an Associate Professor of Political Science at Boise State University. His research … Read more