Paper 3

MEDCollector: Multisource Epidemic Data Collector

Authors: João Zamite, Fabrício A. B. Silva, Francisco Couto and Mário J. Silva

Volume 4 (2011)

Abstract

We present a novel approach for epidemic data collection and integration based on the principles of interoperability and modularity. Accurate and timely epidemic models require large, fresh datasets. The World Wide Web, due to its explosion in data availability, represents a valuable source for epidemiological datasets. From an e-science perspective, collected data can be shared across multiple applications to enable the creation of dynamic platforms to extract knowledge from these datasets. Our approach, MEDCollector, addresses this problem by enabling data collection from multiple sources and its upload to the repository of an epidemic research information platform. Enabling the flexible use and configuration of services through workflow definition, MEDCollector is adaptable to multiple Web sources. Identified disease and location entities are mapped to ontologies, not only guaranteeing the consistency within gathered datasets but also allowing the exploration of relations between the mapped entities. MEDCollector retrieves data from the web and enables its packaging for later use in epidemic modeling tools.