Re: Re: major HTML scubbing needed
Sascha Nemecek <[email protected]>
| Newsgroups | gmane.comp.windows.shells.litestep |
|---|---|
| Message-ID | <[email protected]> |
Maybe a simple grep under linux will do. Could you post an example of
the HTML code?
Sascha
Brian Wolven wrote:
> Paul wrote:
>> Brian Wolven wrote:
>>> Paul wrote:
>>>> OK, all this talk of Bash got me all hot and bothered. I am
>>>> downloading the entire website now. I will weed out all the tiny
>>>> nothing pages and only keep the pages that have quotes on them. Once
>>>> that is done, we need to find a way to remove all the extraneous
>>>> HTML junk.
>>>>
>>>> Brian, this sounds right up your alley. I am assuming that all the
>>>> quote pages formatting will be identical, otherwise it will be for
>>>> naught. If they are, could a VBS script be written to parse out
>>>> plain text that we could then convert into one line quotes line our
>>>> quotes file?
>>>
>>> You mean you can't just do it with wazup.dll? ;)
>>
>> I didn't think so. Otherwise it would have been done by now right?
>
> I bet rabidcow can do it in DOS.
>
>
> ---------------------------------------------------------------------
> Can't Unsubscribe? Check http://desktopian.org/listunsub.html
> LS List Homepage: http://wuzzle.org/list/litestep.php
> Get the LS FAQ: http://lsfaq.shellfront.org
---------------------------------------------------------------------
Can't Unsubscribe? Check http://desktopian.org/listunsub.html
LS List Homepage: http://wuzzle.org/list/litestep.php
Get the LS FAQ: http://lsfaq.shellfront.org