I’m trying to write my first crawler by using PHP with cURL library. My

Question

0

Asked: June 16, 20262026-06-16T16:45:41+00:00 2026-06-16T16:45:41+00:00

I’m trying to write my first crawler by using PHP with cURL library. My

0

I’m trying to write my first crawler by using PHP with cURL library. My aim is to fetch data from one site systematically, which means that the code doesn’t follow all hyperlinks on the given site but only specific links.

Logic of my code is to go to the main page and get links for several categories and store those in an array. Once it’s done the crawler goes to those category sites on the page and looks if the category has more than one pages. If so, it stores subpages also in another array. Finally I merge the arrays to get all the links for sites that needs to be crawled and start to fetch required data.

I call the below function to start a cURL session and fetch data to a variable, which I pass to a DOM object later and parse it with Xpath. I store cURL total_time and http_code in a log file.

The problem is that the crawler runs for 5-6 minutes then stops and doesn’t fetch all required links for sub-pages. I print content of arrays to check result. I can’t see any http error in my log, all sites give a http 200 status code. I can’t see any PHP related error even if I turn on PHP debug on my localhost.

I assume that the site blocks my crawler after few minutes because of too many requests but I’m not sure. Is there any way to get a more detailed debug? Do you think that PHP is adequate for this type of activity because I wan’t to use the same mechanism to fetch content from more than 100 other sites later on?

My cURL code is as follows:

function get_url($url)
{
    $ch = curl_init();
    curl_setopt($ch, CURLOPT_HEADER, 0);
    curl_setopt($ch, CURLOPT_RETURNTRANSFER, 1);
    curl_setopt($ch, CURLOPT_CONNECTTIMEOUT, 30);
    curl_setopt($ch, CURLOPT_URL, $url);
    $data = curl_exec($ch);
    $info = curl_getinfo($ch);  
    $logfile = fopen("crawler.log","a");
    echo fwrite($logfile,'Page ' . $info['url'] . ' fetched in ' . $info['total_time'] . ' seconds. Http status code: ' . $info['http_code'] . "\n");
    fclose($logfile);
    curl_close($ch);

    return $data;
}

// Start to crawle main page.

$site2crawl = 'http://www.site.com/';

$dom = new DOMDocument();
@$dom->loadHTML(get_url($site2crawl));
$xpath = new DomXpath($dom);

Report

Leave an answer
Cancel reply

You must login to add an answer.

Need An Account,

1 Answer

Editorial Team · Answer 1 · 2026-06-16T16:45:42+00:00

Editorial Team

2026-06-16T16:45:42+00:00Added an answer on June 16, 2026 at 4:45 pm

Use set_time_limit to extend the amount of time your script can run for. That is why you are getting Fatal error: Maximum execution time of 30 seconds exceeded in your error log.

0

Reply
Share
Share

- Report

Sign Up

Sign In

Forgot Password

The Archive Base Latest Questions

I’m trying to write my first crawler by using PHP with cURL library. My

Leave an answerCancel reply

1 Answer

Leave an answer
Cancel reply