Re: Yahoo! Groups: Welcome to sed-users. Visit today!
Luigi Assom <[email protected]>
| Newsgroups | gmane.editors.sed.user |
|---|---|
| Message-ID | <CAEiLF=tE6qkm-h0J0R5Mqy4tUA3j3EtREiM2fk51Nikt5sbgCg@mail.gmail.com> |
thank you Cameron! very useful! On Tue, Apr 24, 2012 at 1:05 PM, Joe Gain <[email protected]> wrote: > Hey Luigi, > > one of sed's problems is dealing with newline characters and you have > your html anchor tags split over new lines. There's probably a smart > way to overcome this using sed (and sed's line buffer), but it's not > an easy alternative. > > I think the best way for you is to: > > 1. remove all the new line characters,or at least don't split tags > with new lines (there are many ways to do this, depending on which > version of sed you are using, could be as easy as # sed -n -e > 's/\n//g; p' data.html) > > 2. remove the surrounding tags from the text that you want in your > anchor attributes with something like: > sed -n -e 's/\(<a[^h]*href="\)\([^"]*\)"\([^>]*>\)/\2/g; p' > > 3. remove all the other tags, using the sed script you have already. > (You can pipe the output of 2 into 3.) > > Hope this helps. > > Joe > > PS. In general, you can look at some other tools like perl or ruby to > make your life easier. > > On Mon, Apr 23, 2012 at 11:21 AM, Luigi Assom <[email protected]> > wrote: > > Hi community! > > I subscribed to sed group because I need to do some parsing of txt files > > with html. > > I am totally new to sed, I 've just discovered it yesterday!! > > I tried something (strip all html tag) and I succeed, but I need > something > > more complex and I cannot dig it - could you please help with the syntax? > > > > I need to extract URL contained in <a> tags, only the part like > /RED/BLUE/ > > , and to strip out every other tag, so that I have clean text and only > > THOSE urls. > > > > I made this attempt: > > > > txt structure: > > > >> </b> (Displaying X Results)<table><tr> <td valign="top"><a > >> href="/RED/BLUE/" onClick="(new Image()).src='/url/';"><img > src="image.jpg" > >> width="30" height="30" border="0"></a> </td><td align="right" > >> valign="top"><img src="/images/c.gif" width="1" > height="6"><br>1.</td><td > >> valign="top"><img src="/images/c.gif" width="1" height="6"><br><a > >> href="/RED/BLUE/" onclick="(new Image()).src='url/';">TEXT TO > EXTRACT</a> > >> TEXT TO EXTRACT <p class="find-alike">alike "OTHER TEXT > TO > >> EXTRACT" - BLABLA <em>(blabla)</em></p> <p > class="find-alike">alike > >> "OTHER TEXT TO EXTRACT" - BALABLA</p> </td></tr><tr> <td > >> valign="top"><img src="/images/b.gif" alt="" width="23" > height="1"></td><td > >> align="right" valign="top">2.</td><td valign="top"><a href="/RED/BLUE/" > >> onclick="(new Image()).src='url/';">TEXT TO EXTRACT</a> TEXT TO EXTRACT > >> </td></tr> > >> > > > > command: > > sed -e 's/<[^>]*>//g;/</N;//b' > > > > it works, but strip out all tags. > > > > I've tried to customize it, but nothing happened: > > sed -e '/RED/!d;' > > > > I also tried some script on the web but I cannot make them work... > > I even new in using Unix :) > > I am working with macbook sed based. > > > > > > Could you please help out in this task? > > Thank you! > > Luigi > > > > > > [Non-text portions of this message have been removed] > > > > > > > > ------------------------------------ > > > > -- > > Yahoo! Groups Links > > > > > > > > > > -- > joe gain > > jacob-burckhardt-str. 16 > 78464 konstanz > germany > > +49 (0)7531 60389 > > (...otherwise in ???) > -- Luigi Assom Skype contact: oggigigi [Non-text portions of this message have been removed]