python抓取网页数据( Python中获取整个页面的代码：运行结果实例扩展详解)

优采云发布时间: 2022-04-09 22:04

　　python抓取网页数据(

Python中获取整个页面的代码：运行结果实例扩展详解)

　　如何在python中获取整个网页的源代码

　　更新时间：2020年8月3日07:54:00 转载：Ly

　　在这篇文章中，小编整理了python获取整个网页源代码的方法。需要的朋友可以参考一下。

　　1、Python中获取整个页面的代码：

import requests

res = requests.get('https://blog.csdn.net/yirexiao/article/details/79092355')

res.encoding = 'utf-8'

print(res.text)

　　2、运行结果

　　实例扩展：

from bs4 import BeautifulSoup

import time,re,urllib2

t=time.time()

websiteurls={}

def scanpage(url):

websiteurl=url

t=time.time()

n=0

html=urllib2.urlopen(websiteurl).read()

soup=BeautifulSoup(html)

pageurls=[]

Upageurls={}

pageurls=soup.find_all("a",href=True)

for links in pageurls:

if websiteurl in links.get("href") and links.get("href") not in Upageurls and links.get("href") not in websiteurls:

Upageurls[links.get("href")]=0

for links in Upageurls.keys():

try:

urllib2.urlopen(links).getcode()

except:

print "connect failed"

else:

t2=time.time()

Upageurls[links]=urllib2.urlopen(links).getcode()

print n,

print links,

print Upageurls[links]

t1=time.time()

print t1-t2

n+=1

print ("total is "+repr(n)+" links")

print time.time()-t

scanpage(http://news.163.com/)

　　这是关于python获取整个网页源代码的方法文章的介绍。更多关于python如何获取整个页面的内容，请搜索德牛网之前的文章或继续浏览下方的相关文章，希望大家多多支持德牛网未来！

0

2022-04-09

python抓取网页数据

0 个评论

要回复文章请先登录或注册

AI时代内容工厂

python抓取网页数据( Python中获取整个页面的代码：运行结果实例扩展详解)

0 个评论

发起人

AI时代内容工厂

python抓取网页数据( Python中获取整个页面的代码：运行结果实例扩展详解)

0 个评论

发起人

相关问题