文章/答案/技术大牛

发布

社区首页 >问答首页 >bs4.FeatureNotFound:无法找到具有所需特性的树构建器:html-解析器。您需要安装解析器库吗？

问bs4.FeatureNotFound:无法找到具有所需特性的树构建器:html-解析器。您需要安装解析器库吗？
EN

Stack Overflow用户

提问于 2020-03-29 11:13:47

回答 3查看 1K关注 0票数 0

我试图通过以下代码在Web上刮取：

from bs4 import BeautifulSoup
import requests
import pandas as pd

page = requests.get('https://www.google.com/search?q=phagwara+weather')
soup = BeautifulSoup(page.content, 'html-parser')
day = soup.find(id='wob_wc')

print(day.find_all('span'))

但经常会出现以下错误：

 File "C:\Users\myname\Desktop\webscraping.py", line 6, in <module>
    soup = BeautifulSoup(page.content, 'html-parser')
  File "C:\Users\myname\AppData\Local\Programs\Python\Python38-32\lib\site-packages\bs4\__init__.py", line 225, in __init__
    raise FeatureNotFound(
bs4.FeatureNotFound: Couldn't find a tree builder with the features you requested: html-parser. Do you need to install a parser library?

我安装了lxml和html5lib，但这个问题仍然存在。

beautifulsoup

html-parser

html-treebuilder

python

python-3.x

回答 3

Stack Overflow用户

回答已采纳

发布于 2020-03-29 11:55:52

您需要提到这个标记，所以它应该是soup.find(id="wob_wc")，而不是soup.find("div", id="wob_wc"))

解析器名是html.parser，而不是html-parser，区别是点。

此外，在默认情况下，Google通常会给出200的响应，以防止您了解是否阻塞。通常你必须检查r.content。

我已经包括了headers，现在它开始工作了。

import requests
from bs4 import BeautifulSoup

headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:74.0) Gecko/20100101 Firefox/74.0'}
r = requests.get(
    "https://www.google.com/search?q=phagwara+weather", headers=headers)
soup = BeautifulSoup(r.content, 'html.parser')

print(soup.find("div", id="wob_wc"))

票数 0

Stack Overflow用户

发布于 2020-03-29 11:18:58

您需要将‘html-解析器’转换为soup = BeautifulSoup(page.content, 'html.parser')。

票数 2

Stack Overflow用户

发布于 2021-08-31 11:00:03

实际上，您不需要对整个过程进行迭代："div #wob_wc"，因为当前的位置、天气、日期、温度、降水、湿度和风都是由一个元素组成的，其他地方不需要重复，您可以使用select()或find()。

如果您想要迭代某些内容，那么迭代温度预测是一个好主意，例如：

for forecast in soup.select('.wob_df'):
  high_temp = forecast.select_one('.vk_gy .wob_t:nth-child(1)').text
  low_temp = forecast.select_one('.QrNVmd .wob_t:nth-child(1)').text
  print(f'High: {high_temp}, Low: {low_temp}')

'''
High: 67, Low: 55
High: 65, Low: 56
High: 68, Low: 55
'''

查看SelectorGadget Chrome扩展，在这里您可以通过单击浏览器中所需的元素来获取CSS选择器。CSS选择器参考文献.

代码与在线IDE中的完整示例

from bs4 import BeautifulSoup
import requests, lxml

headers = {
  "User-Agent":
  "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/70.0.3538.102 Safari/537.36 Edge/18.19582"
}

params = {
  "q": "phagwara weather",
  "hl": "en",
  "gl": "us"
}

response = requests.get('https://www.google.com/search', headers=headers, params=params)
soup = BeautifulSoup(response.text, 'lxml')

weather_condition = soup.select_one('#wob_dc').text
tempature = soup.select_one('#wob_tm').text
precipitation = soup.select_one('#wob_pp').text
humidity = soup.select_one('#wob_hm').text
wind = soup.select_one('#wob_ws').text
current_time = soup.select_one('#wob_dts').text

print(f'Weather condition: {weather_condition}\n'
      f'Tempature: {tempature}°F\n'
      f'Precipitation: {precipitation}\n'
      f'Humidity: {humidity}\n'
      f'Wind speed: {wind}\n'
      f'Current time: {current_time}\n')

for forecast in soup.select('.wob_df'):
  day = forecast.select_one('.QrNVmd').text
  weather = forecast.select_one('img.uW5pk')['alt']
  high_temp = forecast.select_one('.vk_gy .wob_t:nth-child(1)').text
  low_temp = forecast.select_one('.QrNVmd .wob_t:nth-child(1)').text
  print(f'Day: {day}\nWeather: {weather}\nHigh: {high_temp}, Low: {low_temp}\n')

---------
'''
Weather condition: Partly cloudy
Temperature: 87°F
Precipitation: 5%
Humidity: 70%
Wind speed: 4 mph
Current time: Tuesday 4:00 PM

Forcast temperature:
Day: Tue
Weather: Partly cloudy
High: 90, Low: 76
...
'''

或者，您也可以通过使用来自Google直接答疑盒API的SerpApi来实现同样的目标。这是一个有免费计划的付费API。

您的示例的主要区别在于，您只需要迭代已经提取的数据，而不是从头开始做所有事情，或者找出如何绕过Google的块。

合并守则：

params = {
  "engine": "google",
  "q": "phagwara weather",
  "api_key": os.getenv("API_KEY"),
  "hl": "en",
  "gl": "us",
}

search = GoogleSearch(params)
results = search.get_dict()

loc = results['answer_box']['location']
weather_date = results['answer_box']['date']
weather = results['answer_box']['weather']
temp = results['answer_box']['temperature']
precipitation = results['answer_box']['precipitation']
humidity = results['answer_box']['humidity']
wind = results['answer_box']['wind']

forecast = results['answer_box']['forecast']

print(f'{loc}\n{weather_date}\n{weather}\n{temp}°F\n{precipitation}\n{humidity}\n{wind}\n')

print(json.dumps(forecast, indent=2))



---------
'''
Phagwara, Punjab, India
Tuesday 4:00 PM
Partly cloudy
87°F
5%
70%
4 mph

[
  {
    "day": "Tuesday",
    "weather": "Partly cloudy",
    "temperature": {
      "high": "90",
      "low": "76"
    },
    "thumbnail": "https://ssl.gstatic.com/onebox/weather/48/partly_cloudy.png"
  }
...
]
'''

免责声明，我为SerpApi工作。

票数 0

页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持

原文链接：

https://stackoverflow.com/questions/60913369

复制

相似问题

问bs4.FeatureNotFound:无法找到具有所需特性的树构建器:html-解析器。您需要安装解析器库吗？
EN

回答 3

Stack Overflow用户

Stack Overflow用户

Stack Overflow用户

社区

活动

圈层

关于

腾讯云开发者

热门产品

热门推荐

更多推荐

问bs4.FeatureNotFound:无法找到具有所需特性的树构建器:html-解析器。您需要安装解析器库吗？EN

回答 3

Stack Overflow用户

Stack Overflow用户

Stack Overflow用户

社区

活动

圈层

关于

腾讯云开发者

热门产品

热门推荐

更多推荐

问bs4.FeatureNotFound:无法找到具有所需特性的树构建器:html-解析器。您需要安装解析器库吗？
EN