log file data source (aDisk perhaps)

[email protected] Wed, 10 Aug 2005 22:34:51 +0000
Newsgroups gmane.comp.java.seda.user
Message-ID <081020052234.10427.42FA810B000C21F8000028BB22007610649B9C9D0108090A99@comcast.net>
Ok I know this is a little bit off the roadmap, but I am involved in a project right now where I need to mine a tremendous amount of XML log files on a daily basis.  I currently have written a parser in Java that does the job but it is very slow.  Some of my analysis steps are much slower than other steps, so I want to try to accelerate the process by using the sandstorm engine to parallelize the slower steps and improve throughput overall.

At some point in the future, I will modify the production systems to send a copy of the SOAP request/response to a SEDA listener in real time.  But in the short term I am restricted to the parsing of log files.

Does anyone know of any examples that they can point me to of an aDisk first stage?  Or any general guidance on the best way to get records from a text file into the data flow?  Is it as simple as creating a stage that reads the file, and enqueues the records (so long as I restrict it to a single thread)?  Or is there something already implemented that I didn't notice from the docs.

Thanks for your help.

-Mark



-------------------------------------------------------
SF.Net email is Sponsored by the Better Software Conference & EXPO
September 19-22, 2005 * San Francisco, CA * Development Lifecycle Practices
Agile & Plan-Driven Development * Managing Projects & Teams * Testing & QA
Security * Process Improvement & Measurement * http://www.sqe.com/bsce5sf