我浏览了网页,在网上阅读和观看了几个关于如何解决我的问题的指南,但我被卡住了,希望能得到一些意见。我试图建立一个网络刮板,将从Reuters刮并购交易部分,并已成功地编写了一个程序,可以刮标题,摘要,日期和链接的文章。然而,我试图解决的问题是,我希望程序仅从标题/文章中抓取摘要,这些标题/文章位于合并和收购列的正下方。当前的程序正在抓取它看到的所有用标签“文章”和属性/类“故事”表示的标题,因此不仅从合并和收购栏目中抓取标题,而且还从市场新闻栏目中抓取标题。
一旦机器人开始从市场新闻栏目中抓取标题,我就一直收到属性错误,因为市场新闻栏目没有任何摘要,因此没有文本可拉,导致我的代码终止。我试图用try/except逻辑路径来解决这个问题,我认为它不会从市场新闻专栏中拉出标题,但代码却一直拉着标题。
我试着写了一行新的代码,告诉程序不要寻找所有的标签和文章,而是寻找所有的标签,如果我给机器人一条更直接的路径,它将从自上而下的方法中抓取文章。然而,这失败了,现在我的头就疼了。提前感谢大家!
到目前为止,我的代码如下:
from bs4 import BeautifulSoup
import requests
website = 'https://www.reuters.com/finance/deals/mergers'
source = requests.get(website).text
soup = BeautifulSoup(source, 'lxml')
for article in soup.find_all('article'):
headline = article.div.a.h3.text.strip()
#threw in strip() to fix the issue of a bunch of space being printed before the headline title.
print(headline+ "\n")
date = article.find("span",class_ = 'timestamp').text
print(date)
try: #Put in Try/Except logic to keep the code going
summary = article.find("div", class_="story-content").p.text
print(summary + "\n")
link = article.find('div', class_='story-content').a['href']
#this bit [href] is the syntax needed for me to pull out the URL from the html code
origin = "https://www.reuters.com/finance/deals/mergers"
print(origin + link + "\n")
except Exception as e:
summary = None
link = None
#This section here is another part I'm working on to get the scraper to go to
#the next page and continue scraping for headlines, dates, summaries, and links
next_page = soup.find('a', class_='control-nav-next')["href"]
source = requests.get(website + next_page).text
soup = BeautifulSoup(source, 'lxml')发布于 2020-03-17 09:54:11
仅更改此行:
for article in soup.select('div[class="column1 col col-10"] article'):使用此语法,.select()将查找<div class="column1 col col-10">下的所有article标记,其中包含您感兴趣的标头,而不是其他标头。
这里是文档:https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.html?highlight=select#css-selectors
https://stackoverflow.com/questions/60713889
复制相似问题