Re: HTML processing
"Ger Hobbelt" <[email protected]>
| Newsgroups | gmane.mail.spam.crm114 |
|---|---|
| Message-ID | <[email protected]> |
On Thu, May 22, 2008 at 9:29 PM, Trever L. Adams <[email protected]> wrote: > I am thinking of using CRM114 for a project. I have a few questions, > only one of which do I think I know the answer to. > > 1) Can CRM114 be used as a library or do I have to have it run as a > daemon running my code and connect via sockets? (I believe the answer is > the latter.) There's no libcrm114 yet - I've something planned there, but don't wait up for me there - and one small extra bit: crm114 is not a daemon, nor does it understand sockets. crm114 can be executed using spawn(), system() or other 'external application invocation' calls available in your programming language. In that way, it's more like sed, grep or awk - to name a few tools which I use regularly on UNIX. > 2) Is it possible to process HTML such that I remove all HTML tags BUT > keep the content of the CONTENT part of META tags and ALT and TITLE > content from <IMG>? If so, how? First thing that comes up in my mind is using the regular expression parsing available in crm114. 'match', 'alter' and other commands in the crm114 language accept regexes, including support for subexpressions, which can be used to pick the data you want out of the input. Of course, other ways are feasible too. -- Met vriendelijke groeten / Best regards, Ger Hobbelt -------------------------------------------------- web: http://www.hobbelt.com/ http://www.hebbut.net/ mail: [email protected] mobile: +31-6-11 120 978 -------------------------------------------------- ------------------------------------------------------------------------- This SF.net email is sponsored by: Microsoft Defy all challenges. Microsoft(R) Visual Studio 2008. http://clk.atdmt.com/MRT/go/vse0120000070mrt/direct/01/