I’m writing a simple crawler, and ideally to save bandwidth, I’d only like to download the text and links on the page. Can I do that using HTTP Headers? I’m confused about how they work.
Share
Sign Up to our social questions and Answers Engine to ask questions, answer people’s questions, and connect with other people.
Login to our social questions & Answers Engine to ask questions answer people’s questions & connect with other people.
Lost your password? Please enter your email address. You will receive a link and will create a new password via email.
Please briefly explain why you feel this question should be reported.
Please briefly explain why you feel this answer should be reported.
Please briefly explain why you feel this user should be reported.
You’re on the right track to solving the problem.
I’m not sure how much you already know about HTTP headers, but basically an HTTP header is just a string formatting for a web server – it follows a protocol – and is pretty straightforward in that aspect. You write a request, and receive a response. The requests look like the things you see in the Firefox plugin LiveHTTPHeaders at https://addons.mozilla.org/en-US/firefox/addon/3829/.
I wrote a small post at my site http://blog.gnucom.cc/2010/write-http-request-to-web-server-with-php/ that shows you how you can write a request to a web server and then later read the response. If you only accept text/html you’ll only accept a subset of what is available on the web (so yes, it will “optimize” your script to an extent). Note this example is really low level, and if you’re going to write a spider you may want to use an existing library like cURL or whatever other tools your implementation language offers.