php多线程抓取多个网页(行行网电子书多线程-撸代码代码非常简单-写在前面 )

优采云发布时间: 2022-03-19 08:05

　　php多线程抓取多个网页(行行网电子书多线程-撸代码代码非常简单-写在前面

)

　　星星网电子书多线程-写在前面

　　最近，我正在寻找一些电子书来阅读，所以我翻阅了它们。然后，我找到了一个叫周渡的网站，网站很好，简单清爽，书很多，而且都是开放的，都可以直接从百度网盘下载，而且更新速度也还行，就爬了上去。这篇文章可以文章学习，这么好的分享网站，尽量不要爬，会影响别人访问速度。需要数据的可以在我的博客下评论，我会发给你，QQ，邮箱什么的。

　　这个网站页面的逻辑很简单。我翻阅了这本书的详细信息页面，它看起来像下面这样。我们只需要循环生成这些页面的链接，然后就可以爬取了。为了速度，我使用如果你想爬取以后的数据，就在这个博客下面评论，不要破坏别人的服务器。

　　http://www.ireadweek.com/index.php/bookInfo/11393.html

http://www.ireadweek.com/index.php/bookInfo/11.html

....

　　星星网电子书多线程-代码代码

　　代码很简单。以我们之前的教程做铺垫，用很少的代码就可以实现完整的功能。最后将采集的内容写入csv文件，（什么是csv，你百度一下就知道了）这段代码是IO密集型操作，我们使用aiohttp模块来写。

　　第一步

　　拼接 URL 以启动线程。

　　import requests

# 导入协程模块

import asyncio

import aiohttp

headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/68.0.3440.106 Safari/537.36",

"Host": "www.ireadweek.com",

"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,image/apng,*/*;q=0.8"}

async def get_content(url):

print("正在操作:{}".format(url))

# 创建一个session 去获取数据

async with aiohttp.ClientSession() as session:

async with session.get(url,headers=headers,timeout=3) as res:

if res.status == 200:

source = await res.text() # 等待获取文本

print(source)

if __name__ == '__main__':

url_format = "http://www.ireadweek.com/index.php/bookInfo/{}.html"

full_urllist = [url_format.format(i) for i in range(1,11394)] # 11394

loop = asyncio.get_event_loop()

tasks = [get_content(url) for url in full_urllist]

results = loop.run_until_complete(asyncio.wait(tasks))

　　上面的代码可以同步开启N个多线程，但是这样很容易导致别人的服务器瘫痪。因此，我们必须限制并发数。下面的代码，你可以试着把它放在指定的位置。

　　sema = asyncio.Semaphore(5)

# 为避免爬虫一次性请求次数太多，控制一下

async def x_get_source(url):

with(await sema):

await get_content(url)

　　第 2 步

　　处理抓取网页的源代码，提取我们想要的元素。我添加了一种使用 lxml 提取数据的方法。

　　def async_content(tree):

title = tree.xpath("//div[@class='hanghang-za-title']")[0].text

# 如果页面没有信息，直接返回即可

if title == '':

return

else:

try:

description = tree.xpath("//div[@class='hanghang-shu-content-font']")

author = description[0].xpath("p[1]/text()")[0].replace("作者：","") if description[0].xpath("p[1]/text()")[0] is not None else None

cate = description[0].xpath("p[2]/text()")[0].replace("分类：","") if description[0].xpath("p[2]/text()")[0] is not None else None

douban = description[0].xpath("p[3]/text()")[0].replace("豆瓣评分：","") if description[0].xpath("p[3]/text()")[0] is not None else None

# 这部分内容不明确，不做记录

#des = description[0].xpath("p[5]/text()")[0] if description[0].xpath("p[5]/text()")[0] is not None else None

download = tree.xpath("//a[@class='downloads']")

except Exception as e:

print(title)

return

ls = [

title,author,cate,douban,download[0].get('href')

]

return ls

　　第 3 步

　　格式化数据后，将其保存为 csv 文件并收工！

　　 print(data)

with open('hang.csv', 'a+', encoding='utf-8') as fw:

writer = csv.writer(fw)

writer.writerow(data)

print("插入成功！")

　　星星网电子书多线程——运行代码，查看结果

　　因为这可能涉及到从其他人的服务器获取重要数据，所以代码不会上传到github。需要的可以留言，我单独发给你

0

2022-03-19

php多线程抓取多个网页

0 个评论

要回复文章请先登录或注册

AI时代内容工厂

php多线程抓取多个网页(行行网电子书多线程-撸代码代码非常简单-写在前面 )

0 个评论

发起人

AI时代内容工厂

php多线程抓取多个网页(行行网电子书多线程-撸代码代码非常简单-写在前面 )

0 个评论

发起人

相关问题