I want to make a program that will retrieve some information a url.
For example i give the url below, from
librarything
How can i retrieve all the words below the “TAGS” tab, like
Black Library fantasy Thanquol & Boneripper Thanquol and Bone Ripper Warhammer ?
I am thinking of using java, and design a data mining wrapper, but i am not sure how to start. Can anyone give me some advice?
EDIT:
You gave me excellent help, but I want to ask something else.
For every tag we can see how many times each tag has been used, when we press the “number” button. How can I retrieve that number also?
You could use a HTML parser like Jsoup. It allows you to select HTML elements of interest using simple CSS selectors:
E.g.
which prints
Please note that you should read website’s
robots.txt-if any- and read the website’s terms of service -if any- or your server might be IP-banned sooner or later.