文章/答案/技术大牛

发布

社区首页 >问答首页 >检查ElementTree节点是否为空故障

问检查ElementTree节点是否为空故障
EN

Stack Overflow用户

提问于 2017-04-26 02:19:03

回答 2查看 2.1K关注 0票数 0

我一直得到错误：AttributeError: 'NodeList' object has no attribute 'data'，但我只是尝试检查该节点是否为空，如果是，只需传递a -1而不是值。我的理解是temp_pub.getElementsByTagName("pages").data应该返回None。我该怎么解决这个问题？

(p.s.-我试过!= None和is None)

xmldoc = minidom.parse('pubsClean.xml')

#loop through <pub> tags to find number of pubs to grab
root = xmldoc.getElementsByTagName("root")[0]
pubs = [a.firstChild.data for a in root.getElementsByTagName("pub")]
num_pubs = len(pubs)
count = 0

while(count < num_pubs):

    temp_pages = 0
    #get data from each <pub> tag
    temp_pub = root.getElementsByTagName("pub")[count]
    temp_ID = temp_pub.getElementsByTagName("ID")[0].firstChild.data
    temp_title = temp_pub.getElementsByTagName("title")[0].firstChild.data
    temp_year = temp_pub.getElementsByTagName("year")[0].firstChild.data
    temp_booktitle = temp_pub.getElementsByTagName("booktitle")[0].firstChild.data
    #handling no value
    if temp_pub.getElementsByTagName("pages").data != None:  
        temp_pages = temp_pub.getElementsByTagName("pages")[0].firstChild.data
    else: 
        temp_pages = -1

    temp_authors = temp_pub.getElementsByTagName("authors")[0]
    temp_author_array = [a.firstChild.data for a in temp_authors.getElementsByTagName("author")]
    num_authors = len(temp_author_array)
    count = count + 1

正在处理的XML

<pub>
    <ID>5010</ID>
    <title>Model-Checking for L<sub>2</sub</title>
    <year>1997</year>
    <booktitle>Universit&auml;t Trier, Mathematik/Informatik, Forschungsbericht</booktitle>
    <pages></pages>
    <authors>
        <author>Helmut Seidl</author>
    </authors>
</pub>
<pub>
    <ID>5011</ID>
    <title>Locating Matches of Tree Patterns in Forest</title>
    <year>1998</year>
    <booktitle>Universit&auml;t Trier, Mathematik/Informatik, Forschungsbericht</booktitle>
    <pages></pages>
    <authors>
        <author>Andreas Neumann</author>
        <author>Helmut Seidl</author>
    </authors>
</pub>

从编辑(with到ElementTree)的完整代码

#for execute command to work
import sqlite3
import xml.etree.ElementTree as ET
con = sqlite3.connect("publications.db")
cur = con.cursor()

from xml.dom import minidom
#use this to clean the foreign characters
import re

def anglicise(matchobj): 
    if matchobj.group(0) == '&amp;':
        return matchobj.group(0)
    else:
        return matchobj.group(0)[1]

outputFilename = 'pubsClean.xml'

with open('test.xml') as inXML, open(outputFilename, 'w') as outXML:
    outXML.write('<root>\n')
    for line in inXML.readlines():
        if (line.find("<sub>") or line.find("</sub>")):
            newline = line.replace("<sub>", "")
            newLine = newline.replace("</sub>", "")
        outXML.write(re.sub('&[a-zA-Z]+;',anglicise,newLine))
    outXML.write('\n</root>')


tree = ET.parse('pubsClean.xml')
root = tree.getroot()

xmldoc = minidom.parse('pubsClean.xml')
#loop through <pub> tags to find number of pubs to grab
root2 = xmldoc.getElementsByTagName("root")[0]
pubs = [a.firstChild.data for a in root2.getElementsByTagName("pub")]
num_pubs = len(pubs)
count = 0

while(count < num_pubs):

    temp_pages = 0
    #get data from each <pub> tag

    temp_ID = root.find(".//ID").text
    temp_title = root.find(".//title").text
    temp_year = root.find(".//year").text
    temp_booktitle = root.find(".//booktitle").text
    #handling no value
    if root.find(".//pages").text:  
        temp_pages = root.find(".//pages").text
    else: 
        temp_pages = -1 

    temp_authors = root.find(".//authors")
    temp_author_array = [a.text for a in temp_authors.findall(".//author")]
    num_authors = len(temp_author_array)
    count = count + 1

    #process results into sqlite
    pub_params = (temp_ID, temp_title)
    cur.execute("INSERT OR IGNORE INTO publication (id, ptitle) VALUES (?, ?)", pub_params)
    cur.execute("INSERT OR IGNORE INTO journal (jtitle, pages, year, pub_id, pub_title) VALUES (?, ?, ?, ?, ?)", (temp_booktitle, temp_pages, temp_year, temp_ID, temp_title))
    x = 0
    while(x < num_authors):
        cur.execute("INSERT OR IGNORE INTO authors (name, pub_id, pub_title) VALUES (?, ?, ?)", (temp_author_array[x],temp_ID, temp_title))
        cur.execute("INSERT OR IGNORE INTO wrote (name, jtitle) VALUES (?, ?)", (temp_author_array[x], temp_booktitle))   
        x = x + 1


con.commit()
con.close()    

print("\nNumber of entries processed: ", count)

python

elementtree

回答 2

Stack Overflow用户

发布于 2017-04-26 02:37:29

您可以使用attributes方法获取类似字典的对象(文档)，然后查询字典：

if temp_pub.getElementsByTagName("pages").attributes.get('data'):

票数 0

Stack Overflow用户

发布于 2017-04-26 02:54:12

正如错误消息所示，getElementsByTagName()既不返回单个节点，也不返回None，而是返回` `NodeList。因此，您应该检查长度，以查看返回的列表是否包含任何项：

if len(temp_pub.getElementsByTagName("pages")) > 0:  
    temp_pages = temp_pub.getElementsByTagName("pages")[0].firstChild.data

或者您可以直接将列表传递给if，因为空列表是错误的：

if temp_pub.getElementsByTagName("pages"):  
    temp_pages = temp_pub.getElementsByTagName("pages")[0].firstChild.data

附带注意，尽管有这个问题的标题和标签，但您的代码表明您使用的是minidom而不是ElementTree。例如，使用ElementTree可以简化代码：

# minidom
temp_ID = temp_pub.getElementsByTagName("ID")[0].firstChild.data
# finding single element can be using elementtree's `find()`
temp_ID = temp_pub.find(".//ID").text
....
# minidom
temp_author_array = [a.firstChild.data for a in temp_authors.getElementsByTagName("author")]
# finding multiple elements using elementtree's `find_all()`
temp_author_array = [a.text for a in temp_authors.find_all(".//author")]

票数 0

页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持

原文链接：

https://stackoverflow.com/questions/43623912

复制

相似问题

问检查ElementTree节点是否为空故障
EN

回答 2

Stack Overflow用户

Stack Overflow用户

社区

活动

圈层

关于

腾讯云开发者

热门产品

热门推荐

更多推荐

问检查ElementTree节点是否为空故障EN

回答 2

Stack Overflow用户

Stack Overflow用户

社区

活动

圈层

关于

腾讯云开发者

热门产品

热门推荐

更多推荐

问检查ElementTree节点是否为空故障
EN