Best way to extract all the links from a HTML page

Santiago Basulto <[email protected]>
Newsgroups gmane.comp.parsers.htmlparser.user
Message-ID <[email protected]>
Hello people.

I'm starting with HTMLParser. It seems a great library. I've doing
some benchmarking and runs really fast.

Now i'm trying to improve it a little bit.

In my software, i use something like this to extract all links:

public class LinkVisitor extends NodeVisitor {
        private Set<String> links = new HashSet<String>(100);
	public LinkVisitor(){
	}
	public void visitTag(Tag tag) {
		String name = tag.getTagName();
		if ("a".equalsIgnoreCase(name)){
			String hrefValue = tag.getAttribute("href");
			links.add(tag.getAttribute("href"));
		}
	}
	public Set<String> getLinks(){
		return this.urls;
	}
	
}

But, reading a little bit i found other classes that may help, but
don't know how to use them. Can anyone help me out?

The idea is to extract all the links from a String (that contains an
HTML page already read from an URLConnection). Is there anyway to
"Canonize" them? I mean, if the href says "/food/fruits/2" convert it
to "http://www.foodsite.com/home/fruits/2"?


Thanks a lot!

-- 
Santiago Basulto.-

------------------------------------------------------------------------------
Beautiful is writing same markup. Internet Explorer 9 supports
standards for HTML5, CSS3, SVG 1.1,  ECMAScript5, and DOM L2 & L3.
Spend less time writing and  rewriting code and more time creating great
experiences on the web. Be a part of the beta today.
http://p.sf.net/sfu/beautyoftheweb
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.