Manual Data Collection in the age of Computers

I am beginning a new data collection project that requires the manual coding of data collected from various sources in print and online.  As I start this project, I am tasked with how to build a master record of all the data I collect in the process.  I have worked on projects that used extensive paper coding forms that were later filed away only to be retrieved when appropriate.  This serves as a safeguard to both checking original coding decision, errors in the database, and any other information the coders found while researching the topic at hand.  Alternatively, other projects had an evolving excel spreadsheet that itself was the master record – duplicate copies served as a safeguard against accidents while the final form only existed when the researchers stopped coding.  Finally, projects that are entirely automated tend to generate their own database that the researchers can then use to extract useful information into a final dataset for examination.

For this project, I decided to make a separate, master database that then will be used to generate data sets as needed.  Any changes will be documented in the master set and stripping out coding variables (last updated, side notes, etc.) that are otherwise irrelevant to people using the data.  Thus, the final versions will be tab delimited for general consumption, produce in Stata, and a R version while the master remains in a different format for official changes.

More after the jump; those interested, there are screenshots and I attempt to elicit ideas for more efficient mechanisms to collect and store data…

Read more

I am Easily Distracted by Databases

Jonathan Dingel on Friday stumbled upon a Preferential Trade Agreements Database hosted by the McGill University Faculty of Law which contains the text, or link to the text, of multiple PTAs. Given the abundance of studies that use trade activity as a direct (or proxy) measure for openness, this is an incredible collection that makes … Read more