首页
学习
活动
专区
圈层
工具
发布
社区首页 >问答首页 >Python Web Scraper尝试让程序抓取某个特定位置的数据,而不是整个页面

问Python Web Scraper尝试让程序抓取某个特定位置的数据,而不是整个页面
EN

Stack Overflow用户
提问于 2020-03-17 06:14:21
回答 1查看 139关注 0票数 1

我浏览了网页,在网上阅读和观看了几个关于如何解决我的问题的指南,但我被卡住了,希望能得到一些意见。我试图建立一个网络刮板,将从Reuters刮并购交易部分,并已成功地编写了一个程序,可以刮标题,摘要,日期和链接的文章。然而,我试图解决的问题是,我希望程序仅从标题/文章中抓取摘要,这些标题/文章位于合并和收购列的正下方。当前的程序正在抓取它看到的所有用标签“文章”和属性/类“故事”表示的标题,因此不仅从合并和收购栏目中抓取标题,而且还从市场新闻栏目中抓取标题。

一旦机器人开始从市场新闻栏目中抓取标题,我就一直收到属性错误,因为市场新闻栏目没有任何摘要,因此没有文本可拉,导致我的代码终止。我试图用try/except逻辑路径来解决这个问题,我认为它不会从市场新闻专栏中拉出标题,但代码却一直拉着标题。

我试着写了一行新的代码,告诉程序不要寻找所有的标签和文章,而是寻找所有的标签,如果我给机器人一条更直接的路径,它将从自上而下的方法中抓取文章。然而,这失败了,现在我的头就疼了。提前感谢大家!

到目前为止,我的代码如下:

代码语言:javascript
复制
from bs4 import BeautifulSoup
import requests

website = 'https://www.reuters.com/finance/deals/mergers'
source = requests.get(website).text
soup = BeautifulSoup(source, 'lxml')

for article in soup.find_all('article'):
    headline = article.div.a.h3.text.strip()
    #threw in strip() to fix the issue of a bunch of space being printed before the headline title.
    print(headline+ "\n")

    date = article.find("span",class_ = 'timestamp').text
    print(date)

    try: #Put in Try/Except logic to keep the code going
        summary = article.find("div", class_="story-content").p.text
        print(summary + "\n")
        link = article.find('div', class_='story-content').a['href']
        #this bit [href] is the syntax needed for me to pull out the URL from the html code
        origin = "https://www.reuters.com/finance/deals/mergers"
        print(origin + link + "\n")
    except Exception as e:
        summary = None
        link = None

    #This section here is another part I'm working on to get the scraper to go to
    #the next page and continue scraping for headlines, dates, summaries, and links
    next_page = soup.find('a', class_='control-nav-next')["href"]
    source = requests.get(website + next_page).text
    soup = BeautifulSoup(source, 'lxml')
EN

回答 1

Stack Overflow用户

回答已采纳

发布于 2020-03-17 09:54:11

仅更改此行:

代码语言:javascript
复制
for article in soup.select('div[class="column1 col col-10"] article'):

使用此语法,.select()将查找<div class="column1 col col-10">下的所有article标记,其中包含您感兴趣的标头,而不是其他标头。

这里是文档:https://www.crummy.com/software/BeautifulSoup/bs4/doc/index.html?highlight=select#css-selectors

票数 0
EN
页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持
原文链接:

https://stackoverflow.com/questions/60713889

复制
相关文章

相似问题

领券
问题归档专栏文章快讯文章归档关键词归档开发者手册归档开发者手册 Section 归档