CFP: 2nd International Workshop on Public Data about Software Development
"Jesus M. Gonzalez-Barahona" <jgb-eneoqThHf3GDmu24MmG3/[email protected]> Thu, 01 Mar 2007 01:07:54 +0100
| Newsgroups | gmane.science.opensource.community |
|---|---|
| Message-ID | <1172707674.3515.146.camel@localhost> |
CALL FOR CONTRIBUTIONS
2nd International Workshop on Public Data about Software Development
====================================================================
(WoPDaSD 2007)
co-located with The Third International Conference on
Open Source Systems, June 14th 2007, Limerick, Ireland
Motivation
----------
In the latest years, and specially thanks to the huge availability of
data about software development that can be obtained from libre (free,
open source) projects, the research community is starting to produce,
use and exchange large data sets of information. These data sets have
to be retrieved, purged, described, and can be published for public
consumption by other groups. Their availability allows for the
decoupling of research activities (some groups can focus on data
retrieval and preliminary analysis, which others can devote to more
in-dept analysis without bothering with data retrieval), the
reproducibility of research results, and even the collaboration (and
competition) in the analysis of data.
All this activity is being presented in several workshops and
conferences, but a single place to exchange experiences does not exist
yet. We propose this workshop as such a place, where researchers in
the field can discuss specifically about this kind of data sets, how
they are retrieved, how can they be analyzed and mined, how they can
be exchanged and complemented, etc.
Main goals
----------
The goal of this workshop is to foster the analysis of public
available data sources about software development and the exchange of
data between different research groups.
The workshop is aimed specifically at two different target studies:
1. Analysis of some data collections about software development
(provided by the organizers, see below).
The analysis should show a methodology for exploring any of
those data sets (or better, to relate both) searching for some
specific result in the area of software development, and its
applications to the actual data sets. The study can be in the
field of software engineering, economics, sociology, human
resources, and others.
2. Retrieval process and exchange formats of public available data
collections about software development.
The data collections presented should be publicly available,
based themselves on public data (so that other groups could
reproduce the data collection process), and be related to the
field of software development. This includes, but is not
limited to, data from source control systems, but tracking
systems, mailing lists, websites, source and binary code,
quality assurance systems, etc. Although any kind of
data collection can be considered, those including information
about a large amount of projects will be considered especially
appropriate.
Detailed description
--------------------
Following the goals described above, the workshop will accept papers
about two specific issues:
1. Analysis of two data collections about libre software
development: FLOSSMole and CVSAnaly-SF.
These collections, already available to any researcher, are
offered for the analysis. The studies submitted should detail
how they have been used, which part of the information has been
considered, how they have been validated or filtered and/or
post-processed (if that is the case). The description should
be detailed enough to let any other research group reproduce the
study.
2. Studies about the data retrieval and preparation for public
consumption of data sets in the same realm, which could be proposed
for analysis in future editions of the workshop.
FLOSSMole
---------
FLOSSmole (formerly OSSmole) is a set of tools for gathering data
(metrics) about the development of free/libre/open source
projects. The FLOSSMole project also publishes the resulting analysis
about FLOSS projects, and accepts data donations from other research
groups. It offers this workshop a complete set of data gathered from
the SourceForge development platform and the Freshmeat announcement
systems.
More information can be obtained from http://ossmole.sourceforge.net
CVSAnaly-SF
-----------
CVSAnalY is a tool created by the Libre Software Engineering Group at
the Universidad Rey Juan Carlos that extracts statistical information
out of CVS (and recently Subversion) repository logs and transforms it
in database SQL formats. It has been used to retrieve information for
all projects that have an active CVS system at SourceForge. This data
set is publicly offered to be analyzed in this workshop.
More information can be obtained from http://libresoft.urjc.es/Data
Target audience
---------------
The target audience is composed by the research groups interested in
empirical software engineering and quantitative studies of the
software development processes and methods. This includes not only
software engineers, but also researchers from other fields that might
use the data for economic, social and other studies.
Submissions
-----------
We solicit short position papers (3 pages) and research papers (6
pages). Short papers will be expected to discuss controversial issues
in the field, or describe interesting or thought-provoking ideas that
are not yet fully developed, while full papers will be expected to
describe new research results, and have a higher degree of technical
rigor than short papers. The papers must be in ACM 2-column
format. Authors may indicate their intent to submit a paper by April
16th 2007 (the title of the paper and abstract will need to be
submitted online via e-mail to grex-eneoqThHf3GDmu24MmG3/[email protected]). The full
paper should be sent to grex-eneoqThHf3GDmu24MmG3/[email protected] by April 25th 2007 as
PDF. Notification of acceptance will be sent by May 10th 2007. The
final version of the paper is due on May 28th 2007. Accepted papers
will be published as part of the WoPDaSD proceedings (which will be
available on the Internet).
Challenge
---------
This edition an specific challenge is proposed to contributors, in
addition to regular papers. The topic of the challenge is "data
visualization". For participating in it, contributors should send a
short paper (2 pages of text, up to 4 pages of figures) visualizing
the data in any of the datasets offered (FLOSSMole, CVSAnaly-SF, or
both).
The text in the paper should explain the visualization technique used,
and its possible applications. The images in the paper should be the
visualization images themselves, or snapshots of them. Visualization
techniques that help to answer interesting questions, to better
understand the data, or to find relationships in it (including
relating data in both datasets) are encouraged.
Each paper will undergo a thorough review, and accepted challenge
papers will be published as part of the workshop proceedings. Authors
of selected challenge papers will be invited to give a presentation at
a special session at the workshop, Deadlines for challenge papers will
be the same than for regular papers.
Important Dates
---------------
* Intent to submit: 16th April 2007 (not mandatory, only for
organizational purposes)
* Deadline for submission: 25th April 2007
* Paper notification: 10th May 2007
* Camera-ready paper due: 28th May 2007
* Workshop date: 14th June 2007
Organizing Committee
--------------------
* Jess M. Gonzlez-Barahona (Universidad Rey Juan Carlos, Spain)
* Megan Conklin (Elon University, USA)
* Gregorio Robles (Universidad Rey Juan Carlos, Spain)
Program Committee
-----------------
* Kevin Crowston (Syracuse University, USA)
* Jean-Christophe Deprez (CETIC, Belgium)
* Daniel M. Germn (University of Victoria, Canada)
* Stefan Koch (Wirtschaftsuniversitt Vienna, Austria)
* Bart Massey (Portland State University, USA)
* Sandro Morasca (Universit dell'Insubria, Italy)
* Walt Scacchi (University of California at Irvine, USA)
* Diomidis Spinellis (Athens University of Economics and Business,
Greece)
* Giancarlo Succi (Free University of Bozen-Bolzano, Italy)
* Tony Wassermann (Carnegie Mellon West, USA)
* Dawid Weiss (Poznan University of Technology, Poland)
* Jim Whitehead (University of California at Santa Cruz, USA)
* Thomas Zimmermann (Universitt des Saarlandes, Germany)
Sponsoring projects
-------------------
Some research projects sponsor this workshop (although it is open to
anyone who registers):
* FLOSSMole
* FLOSSMETRICS
* QUALOSS
* SQO-OSS
* QUALIPSO
* FLOSSWORLD
* EDOS
Some of these projects are funded in part by the European Commission,
under the Information Society Technologies (IST) research programme of
the Sixth Framework Program. A list of the IST projects in the area of
Software Technologies is available from
http://cordis.europa.eu/ist/st/projects.htm
Further information
-------------------
Should you need further information, you can email
jgb @ gsyc.escet.urjc.es.