欢迎您访问程序员文章站本站旨在为大家提供分享程序员计算机编程知识!
您现在的位置是: 首页  >  IT编程

PHP实现抓取百度搜索结果页面【相关搜索词】并存储到txt文件示例

程序员文章站 2022-05-02 23:06:04
本文实例讲述了php实现抓取百度搜索结果页面【相关搜索词】并存储到txt文件。分享给大家供大家参考,具体如下: 一、百度搜索关键词【萬仟网】 【萬仟网】搜索链接...

本文实例讲述了php实现抓取百度搜索结果页面【相关搜索词】并存储到txt文件。分享给大家供大家参考,具体如下:

一、百度搜索关键词【】

PHP实现抓取百度搜索结果页面【相关搜索词】并存储到txt文件示例

【】搜索链接

https://www.baidu.com/s?ie=utf-8&f=8&rsv_bp=0&rsv_idx=1&tn=baidu&wd=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&rsv_pq=ab33cfeb000086a2&rsv_t=7c65vt3kzhcnfgyoin%2fdss%2boquicycaspxwzsobfkhypgripkmi74wii8k8&rqlang=cn&rsv_enter=1&rsv_sug3=1

PHP实现抓取百度搜索结果页面【相关搜索词】并存储到txt文件示例

搜索结果部分源代码:

<div id="rs"><div class="tt">相关搜索</div><table cellpadding="0"><tbody><tr><th><a href="/s?wd=%e6%b8%b8%e6%88%8f%e8%84%9a%e6%9c%ac%e4%b8%80%e8%88%ac%e9%83%bd%e5%9c%a8%e5%93%aa%e6%89%be&rsf=4562&rsp=0&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >游戏脚本一般都在哪找</a></th><td></td><th><a href="/s?wd=%e8%84%9a%e6%9c%ac%e6%80%8e%e4%b9%88%e5%86%99&rsf=4562&rsp=1&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >脚本怎么写</a></th><td></td><th><a href="/s?wd=%e8%84%9a%e6%9c%ac%e6%98%af%e4%bb%80%e4%b9%88%e6%84%8f%e6%80%9d&rsf=4562&rsp=2&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >脚本是什么意思</a></th></tr><tr><th><a href="/s?wd=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6app&rsf=4562&rsp=3&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >app</a></th><td></td><th><a href="/s?wd=%e6%89%8b%e6%9c%ba%e8%84%9a%e6%9c%ac%e5%88%b6%e4%bd%9c&rsf=4562&rsp=4&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >手机脚本制作</a></th><td></td><th><a href="/s?wd=%e6%89%8b%e6%9c%ba%e8%84%9a%e6%9c%ac%e5%a4%a7%e5%85%a8&rsf=4562&rsp=5&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >手机脚本大全</a></th></tr><tr><th><a href="/s?wd=%e8%84%9a%e6%9c%ac%e6%b8%b8%e6%88%8f%e5%88%b6%e4%bd%9c%e5%a4%a7%e5%b8%88&rsf=4562&rsp=6&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >脚本游戏制作大师</a></th><td></td><th><a href="/s?wd=%e6%b8%b8%e6%88%8f%e8%84%9a%e6%9c%ac%e5%88%b6%e4%bd%9c%e6%95%99%e7%a8%8b&rsf=4562&rsp=7&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >游戏脚本制作教程</a></th><td></td><th><a href="/s?wd=%e8%84%9a%e6%9c%ac%e7%b2%be%e7%81%b5&rsf=4562&rsp=8&f=1&oq=%e8%84%9a%e6%9c%ac%e4%b9%8b%e5%ae%b6&ie=utf-8&rsv_idx=1&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam&rqlang=cn&rs_src=0&rsv_pq=c1ff4bdb000208b4&rsv_t=a1f2ocsgs6vkkbcxsdqfbfehkxor65%2ftflpsi30%2f%2fmmk6jqjeukzbv30xam" rel="external nofollow" >脚本精灵</a></th></tr></tbody></table></div>

二、抓取并保存本地

 PHP实现抓取百度搜索结果页面【相关搜索词】并存储到txt文件示例

源代码

index.php:

<form action="index.php" method="post">
<input name="q" type="text" />
<input type="submit" value="get keywords" />
</form>
<?php
header('content-type:text/html;charset=gbk');
class combaike{
  private $o_string=null;
  public function __construct(){
    include('cls.stringex.php');
    $this->o_string=new stringex();
  }
  public function getitem($word){
    $url = "http://www.baidu.com/s?wd=".$word;
    // 构造包头,模拟浏览器请求
    $header = array (
      "host:www.baidu.com",
      "content-type:application/x-www-form-urlencoded",//post请求
      "connection: keep-alive",
      'referer:http://www.baidu.com',
      'user-agent: mozilla/5.0 (compatible; msie 9.0; windows nt 6.1; wow64; trident/5.0; bidubrowser 2.6)'
    );
    $ch = curl_init ();
    curl_setopt ( $ch, curlopt_url, $url );
    curl_setopt ( $ch, curlopt_httpheader, $header );
    curl_setopt ( $ch, curlopt_returntransfer, 1 );
    $content = curl_exec ( $ch );
    if ($content == false) {
    echo "error:" . curl_error ( $ch );
    }
    curl_close ( $ch );
    //输出结果echo $content;
    $this->o_string->string=$content;
    $s_begin='<div id="rs">';
    $s_end='</div>';
    $summary=$this->o_string->getpart($s_begin,$s_end);
    $s_begin='<div class="tt">相关搜索</div><table cellpadding="0"><tr><th>';
    $s_end='</th></tr></table></div>';
    $content=$this->o_string->getpart($s_begin,$s_end);
    return $content;
  }
  public function __destruct(){
    unset($this->o_string);
  }
}
if($_post){
  $com = new combaike();
  $q = $_post['q'];
  $str = $com->getitem($q); //获取搜索内容
  $pat = '/<a(.*?)href="(.*?)" rel="external nofollow" (.*?)>(.*?)<\/a>/i';
  preg_match_all($pat, $str, $m);
  //print_r($m[4]); 链接文字
  $con = implode(",", $m[4]);
  //生成文件夹
  $dates = date("ymd");
  $path="./search/".$dates."/";
  if(!is_dir($path)){
    mkdir($path,0777,true);
  }
  //生成文件
  $file = fopen($path.iconv("utf-8","gbk",$q).".txt",'w');
  if(fwrite($file,$con)){
    echo $con;
    echo '<script>alert("success")</script>';
  }else{
    echo '<script>alert("error")</script>';
  }
  fclose($file);
}
?>

cls.stringex.php:

<?php
header('content-type: text/html; charset=utf-8');
class stringex{
  public $string='';
  public function __construct($string=''){
    $this->string=$string;
  }
  public function preggetpart($s_begin,$s_end){
    $s_begin==preg_quote($s_begin);
    $s_begin=str_replace('/','\/',$s_begin);
    $s_end=preg_quote($s_end);
    $s_end=str_replace('/','\/',$s_end);
    $pattern='/'.$s_begin.'(.*?)'.$s_end.'/';
    $result=preg_match($pattern,$this->string,$a_match);
    if(!$result){
      return $result;
    }else{
      return isset($a_match[1])?$a_match[1]:'';
    }
  }
  public function strstrgetpart($s_begin,$s_end){
    $string=strstr($this->string,$s_begin);
    $string=strstr($string,$s_end,true);
    $string=str_replace($s_begin,'',$string);
    $string=str_replace($s_end,'',$string);
    return $string;
  }
  public function getpart($s_begin,$s_end){
    $result=$this->preggetpart($s_begin,$s_end);
    if(!$result){
      $result=$this->strstrgetpart($s_begin,$s_end);
    }
    return $result;
  }
}
?>

更多关于php相关内容感兴趣的读者可查看本站专题:《php curl用法总结》、《php网络编程技巧总结》、《php数组(array)操作技巧大全》、《php字符串(string)用法总结》、《php数据结构与算法教程》及《php中json格式数据操作技巧汇总

希望本文所述对大家php程序设计有所帮助。