php curl获得header检测GZip压缩的源代码
获得网页header信息,是网站开发人员和维护人员常用的技术。网页的header信息,非常丰富,非专业人士一般较难读懂和理解各个项目的含义。
获取网页header信息,方法多种多样,就php语言来说,我作为一个菜鸟,知道的方法也有4种那么多。下面逐一献上。
方法一:使用get_headers()函数
这个方法很多人使用,也很简单便捷,只需要两行代码即可搞定。如下:
$thisurl = “http://www.webkaka.com/”;
print_r(get_headers($thisurl, 1));
得到的结果为:
Array
(
[0] => HTTP/1.1 200 OK
[Cache-Control] => max-age=86400
[Content-Length] => 76102
[Content-Type] => text/html
[Content-Location] => http://www.webkaka.com/index.html
[Last-Modified] => Fri, 19 Jul 2013 03:52:30 GMT
[Accept-Ranges] => bytes
[ETag] => “50bc48643384ce1:5cb3”
[Server] => Microsoft-IIS/6.0
[X-Powered-By] => ASP.NET
[Date] => Fri, 19 Jul 2013 09:06:39 GMT
[Connection] => close
)
方法二:使用http_response_header
代码也很简单,仅需三行:
$thisurl = “http://www.webkaka.com/”;
$html = file_get_contents($thisurl );
print_r($http_response_header);
得到的结果为:
Array
(
[0] => HTTP/1.1 200 OK
[1] => Cache-Control: max-age=86400
[2] => Content-Length: 76102
[3] => Content-Type: text/html
[4] => Content-Location: http://www.webkaka.com/index.html
[5] => Last-Modified: Fri, 19 Jul 2013 03:52:30 GMT
[6] => Accept-Ranges: bytes
[7] => ETag: “50bc48643384ce1:5cb3”
[8] => Server: Microsoft-IIS/6.0
[9] => X-Powered-By: ASP.NET
[10] => Date: Fri, 19 Jul 2013 09:06:41 GMT
[11] => Connection: close
)
方法三:使用stream_get_meta_data()函数
代码也只有三行:
$thisurl = “http://www.webkaka.com/”;
$fp = fopen($thisurl, ‘r’);
print_r(stream_get_meta_data($fp));
得到的结果为:
Array
(
[wrapper_data] => Array
(
[0] => HTTP/1.1 200 OK
[1] => Cache-Control: max-age=86400
[2] => Content-Length: 76102
[3] => Content-Type: text/html
[4] => Content-Location: http://www.webkaka.com/index.html
[5] => Last-Modified: Fri, 19 Jul 2013 03:52:30 GMT
[6] => Accept-Ranges: bytes
[7] => ETag: “50bc48643384ce1:5cb3”
[8] => Server: Microsoft-IIS/6.0
[9] => X-Powered-By: ASP.NET
[10] => Date: Fri, 19 Jul 2013 09:06:41 GMT
[11] => Connection: close
)
[wrapper_type] => http
[stream_type] => tcp_socket
[mode] => r+
[unread_bytes] => 1086
[seekable] =>
[uri] => http://www.webkaka.com/
[timed_out] =>
[blocked] => 1
[eof] =>
)
上述三种方法都可以轻松获得网页header信息,且包含的信息都已经相当丰富,满足一般要求,不过比较遗憾的是,上述三种方法都不能用来检测网页是否启用了GZip压缩。要检测GZip压缩,还需其他的方法才行。这里介绍的是用curl()函数来检测。
使用curl获得header可以检测GZip压缩
先贴出代码:
<?php
$szUrl = ‘http://www.webkaka.com/’;
$curl = curl_init();
curl_setopt($curl, CURLOPT_URL, $szUrl);
curl_setopt($curl, CURLOPT_HEADER, 1); //输出header信息
curl_setopt($curl, CURLOPT_RETURNTRANSFER, 1); //不显示网页内容
curl_setopt($curl, CURLOPT_ENCODING, ”); //允许执行gzip
$data=curl_exec($curl);
if(!curl_errno($curl))
{
$info = curl_getinfo($curl);
$httpHeaderSize = $info[‘header_size’]; //header字符串体积
$pHeader = substr($data, 0, $httpHeaderSize); //获得header字符串
$split = array(“\r\n”, “\n”, “\r”); //需要格式化header字符串
$pHeader = str_replace($split, ‘<br>’, $pHeader); //使用<br>换行符格式化输出到网页上
echo $pHeader;
}
?>
输出结果如下:
HTTP/1.1 200 OK
Cache-Control: max-age=86400
Content-Length: 15189
Content-Type: text/html
Content-Encoding: gzip
Content-Location: http://www.webkaka.com/index.html
Last-Modified: Fri, 19 Jul 2013 03:52:28 GMT
Accept-Ranges: bytes
ETag: “0268633384ce1:5cb3”
Vary: Accept-Encoding
Server: Microsoft-IIS/6.0
X-Powered-By: ASP.NET
Date: Fri, 19 Jul 2013 09:27:21 GMT
上面输出结果里可以看到一个项目:Content-Encoding: gzip,这个正是我们用来判断网页是否启用GZip压缩的项目。
另外,需要认真注意下本实例里的注释部分,不能少了任何一项,否则可能获取header信息有误。