文章/答案/技术大牛

发布

社区首页 >问答首页 >如何使用BeautifulSoup从网页中获取整个正文文本？

问如何使用BeautifulSoup从网页中获取整个正文文本？
EN

Stack Overflow用户

提问于 2019-07-14 03:08:37

回答 1查看 212关注 0票数 1

我想从一个自然语言处理项目的医学文档的网页上获取一些文本，并且有问题使用BeautifulSoup提取必要的信息。我正在浏览的网站可以在以下地址找到：https://www.mtsamples.com/site/pages/sample.asp?Type=24-Gastroenterology&Sample=2332-Abdominal%20Abscess%20I&D

我想要做的是从这个页面抓取整个文本正文，然后用我的光标这样做，简单地应用一个副本/粘贴就可以给我我感兴趣的合适的文本：

Sample Type / Medical Specialty: Gastroenterology
Sample Name: Abdominal Abscess I&D
Description: Incision and drainage (I&D) of abdominal abscess, excisional debridement of nonviable and viable skin, subcutaneous tissue and muscle, then removal of foreign body.
(Medical Transcription Sample Report)
PREOPERATIVE DIAGNOSIS: Abdominal wall abscess.

... (body text) ...

The finished wound size was 9.0 x 5.3 x 5.2 cm in size. Patient tolerated the procedure well. Dressing was applied, and he was taken to recovery room in stable condition.

但是，我想使用BeautifulSoup实现这一点，因为我想执行一个循环从同一个网站获取多个医学文档。

import requests  
r = requests.get('https://www.mtsamples.com/site/pages/sample.asp?Type=24-Gastroenterology&Sample=2332-Abdominal%20Abscess%20I&D')

from bs4 import BeautifulSoup  
soup = BeautifulSoup(r.text, 'html.parser')  
results = soup.find_all('div', attrs={'id':'sampletext'})

# Here I am able to specify the <h1> tag to get 'Sample Type / Medical Specialty' as well as 'Sample Name' text fields

record.find('h1').text.replace('\n', ' ')

然而，我不能对其余的文本重复这一点(即描述、术前诊断、术后诊断、程序等)。因为没有唯一的标记来标识这些文本字段

如果有人熟悉使用BeautifulSoup进行网络抓取的概念，我将非常感谢您的反馈！同样，我的目标是从网页中获得全文，我最终想要添加到Pandas Dataframe中。谢谢!

python

html

web-scraping

beautifulsoup

回答 1

Stack Overflow用户

回答已采纳

发布于 2019-07-14 11:15:38

好的，我花了一段时间，但是除非手动遍历所有元素，否则提取可用文本的方法并不简单：

import requests
import re
from bs4 import BeautifulSoup, Tag, NavigableString, Comment

url = 'https://www.mtsamples.com/site/pages/sample.asp?Type=24-Gastroenterology&Sample=2332-Abdominal%20Abscess%20I&D'
res = requests.get(url)
res.raise_for_status()
html = res.text
soup = BeautifulSoup(html, 'html.parser')

到目前为止没什么特别的。

title_el = soup.find('h1')
page_title = title_el.text.strip()
first_hr = title_el.find_next_sibling('hr')

description_title = title_el.find_next_sibling('b', text=re.compile('description', flags=re.I))
description_text_parts = []
for s in description_title.next_siblings:
    if s is first_hr:
        break
    if isinstance(s, Tag):
        description_text_parts.append(s.text.strip())
    elif isinstance(s, NavigableString):
        description_text_parts.append(str(s).strip())
description_text = '\n'.join(p for p in description_text_parts if p.strip())

这里我们从page_title得到了<h1>

'Sample Type / Medical Specialty:  Gastroenterology\nSample Name: Abdominal Abscess I&D'

和description，在我们看到文本Description:之后，通过遍历元素。

'Incision and drainage (I&D) of abdominal abscess, excisional debridement of nonviable and viable skin, subcutaneous tissue and muscle, then removal of foreign body.\n(Medical Transcription Sample Report)'

现在，所有标题都放在横向规则下：

# titles are all bold and uppercase
titles = [b for b in first_hr.find_next_siblings('b') if b.text.strip().isupper()]

我们在标题之间找到文本，并将其分配给前面看到的标题。

docs = []
for t in titles:
    text_parts = []
    for s in t.next_siblings:
        # go until next title
        if s in titles:
            break
        if isinstance(s, Comment):
            continue
        if isinstance(s, Tag):
            if s.name == 'div':
                break
            text_parts.append(s.text.strip())
        elif isinstance(s, NavigableString):
            text_parts.append(str(s).strip())
    text = '\n'.join(p for p in text_parts if p.strip())
    docs.append({
        'title': t.text.strip(),
        'text': text
    })

打印文档提供：

[
{'title': 'PREOPERATIVE DIAGNOSIS:', 'text': 'Abdominal wall abscess.'}, 
{'title': 'POSTOPERATIVE DIAGNOSIS:', 'text': 'Abdominal wall abscess.'}, 
{'title': 'PROCEDURE:', 'text': 'Incision and drainage (I&D) of abdominal abscess, excisional debridement of nonviable and viable skin, subcutaneous tissue and muscle, then removal of foreign body.'}, 
{'title': 'ANESTHESIA:', 'text': 'LMA.'}, 
{'title': 'INDICATIONS:', 'text': 'Patient is a pleasant 60-year-old gentleman, who initially had a sigmoid colectomy for diverticular abscess, subsequently had a dehiscence with evisceration.  Came in approximately 36 hours ago with pain across his lower abdomen.  CT scan demonstrated presence of an abscess beneath the incision.  I recommended to the patient he undergo the above-named procedure.  Procedure, purpose, risks, expected benefits, potential complications, alternatives forms of therapy were discussed with him, and he was agreeable to surgery.'}, 
{'title': 'FINDINGS:', 'text': 'The patient was found to have an abscess that went down to the level of the fascia.  The anterior layer of the fascia was fibrinous and some portions necrotic.  This was excisionally debrided using the Bovie cautery, and there were multiple pieces of suture within the wound and these were removed as well.'},
{'title': 'TECHNIQUE:', 'text': 'Patient was identified, then taken into the operating room, where after induction of appropriate anesthesia, his abdomen was prepped with Betadine solution and draped in a sterile fashion.  The wound opening where it was draining was explored using a curette.  The extent of the wound marked with a marking pen and using the Bovie cautery, the abscess was opened and drained.  I then noted that there was a significant amount of undermining.  These margins were marked with a marking pen, excised with Bovie cautery; the curette was used to remove the necrotic fascia.  The wound was irrigated; cultures sent prior to irrigation and after achievement of excellent hemostasis, the wound was packed with antibiotic-soaked gauze.  A dressing was applied.  The finished wound size was 9.0 x 5.3 x 5.2 cm in size.  Patient tolerated the procedure well.  Dressing was applied, and he was taken to recovery room in stable condition.'}
]

票数 1

页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持

原文链接：

https://stackoverflow.com/questions/57024298

复制

相似问题

问如何使用BeautifulSoup从网页中获取整个正文文本？
EN

回答 1

Stack Overflow用户

社区

活动

圈层

关于

腾讯云开发者

热门产品

热门推荐

更多推荐

问如何使用BeautifulSoup从网页中获取整个正文文本？EN

回答 1

Stack Overflow用户

社区

活动

圈层

关于

腾讯云开发者

热门产品

热门推荐

更多推荐

问如何使用BeautifulSoup从网页中获取整个正文文本？
EN