php解析字符串里所有URL地址的方法_PHP教程

程序员文章站 2022-05-14 13:11:05

...

php解析字符串里所有URL地址的方法

具体如下：

// $html = the html on the page

// $current_url = the full url that the html came from

//(only needed for $repath)

// $repath = converts ../ and / and // urls to full valid urls

function pageLinks($html, $current_url = "", $repath = false){

preg_match_all("/\

$links = array();

if(isset($matches[2])){

$links = $matches[2];

}

if($repath && count($links) > 0 && strlen($current_url) > 0){

$pathi = pathinfo($current_url);

$dir = $pathi["dirname"];

$base = parse_url($current_url);

$split_path = explode("/", $dir);

$url = "";

foreach($links as $k => $link){

if(preg_match("/^\.\./", $link)){

$total = substr_count($link, "../");

for($i = 0; $i

array_pop($split_path);

}

$url = implode("/", $split_path) . "/" . str_replace("../", "", $link);

}elseif(preg_match("/^\/\//", $link)){

$url = $base["scheme"] . ":" . $link;

}elseif(preg_match("/^\/|^.\//", $link)){

$url = $base["scheme"] . "://" . $base["host"] . $link;

}elseif(preg_match("/^[a-zA-Z0-9]/", $link)){

if(preg_match("/^http/", $link)){

$url = $link;

}else{

$url = $dir . "/" . $link;

}

$links[$k] = $url;

}

return $links;

}

header("content-type: text/plain");

$url = "http://www.jb51.net";

$html = file_get_contents($url);

// Gets links from the page:

print_r(pageLinks($html));

// Gets links from the page and formats them to a full valid url:

print_r(pageLinks($html, $url, true));

上一篇： ZendOptimizer配置指南_PHP

下一篇： PHP 后端很难返回规范的 JSON 数据吗？

php解析字符串里所有URL地址的方法_PHP教程

php解析字符串里所有URL地址的方法

php解析字符串里所有URL地址的方法_PHP

PHP为表单获取的URL 地址预设 http 字符串函数代码_PHP教程

php提取字符串中网站url地址的方法

php提取字符串中网站url地址的方法

解析关于java,php以及html的所有文件编码与乱码的处理方法汇总_PHP教程

php解析字符串里所有URL地址的方法_PHP教程

php防止伪造数据从地址栏URL提交的方法，伪造url_PHP教程

PHP字符串mbstring处理中文字符串的具体方法解析_PHP教程

php正则表达式解析字符串里的所有URL地址

php正则表达式解析字符串里的所有URL地址