1

我正在尝试使用漂亮的汤来解析行表,并将每行的值保存在字典中。

一个小问题是表的结构有一些行作为节标题。因此,对于具有类“标题”的任何行,我想定义一个名为“节”的变量。这就是我所拥有的,但它不起作用,因为它在说 ['class']TypeError: string indices must be integers

这是我所拥有的:

for i in credits.contents:
    if i['class'] == 'header':
        section = i.contents
        DATA_SET[section] = {}
    else:
        DATA_SET[section]['data_point_1'] = i.find('td', {'class' : 'data_point_1'}).find('p').contents
        DATA_SET[section]['data_point_2'] = i.find('td', {'class' : 'data_point_2'}).find('p').contents
        DATA_SET[section]['data_point_3'] = i.find('td', {'class' : 'data_point_3'}).find('p').contents

数据示例:

<table class="credits">
    <tr class="header">
        <th colspan="3"><h1>HEADER NAME</h1></th>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA</p></td>
        <td class="data_point_2"><p>DATA</p></td>
        <td class="data_point_3"><p>DATA</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA</p></td>
        <td class="data_point_2"><p>DATA</p></td>
        <td class="data_point_3"><p>DATA</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA</p></td>
        <td class="data_point_2"><p>DATA</p></td>
        <td class="data_point_3"><p>DATA</p></td>
    </tr>
    <tr class="header">
        <th colspan="3"><h1>HEADER NAME</h1></th>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA</p></td>
        <td class="data_point_2"><p>DATA</p></td>
        <td class="data_point_3"><p>DATA</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA</p></td>
        <td class="data_point_2"><p>DATA</p></td>
        <td class="data_point_3"><p>DATA</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA</p></td>
        <td class="data_point_2"><p>DATA</p></td>
        <td class="data_point_3"><p>DATA</p></td>
    </tr>
</table>
4

1 回答 1

3

这是一种解决方案,对您的示例数据稍作调整,以便结果更清晰:

from BeautifulSoup import BeautifulSoup
from pprint import pprint

html = '''<body><table class="credits">
    <tr class="header">
        <th colspan="3"><h1>HEADER 1</h1></th>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA11</p></td>
        <td class="data_point_2"><p>DATA12</p></td>
        <td class="data_point_3"><p>DATA12</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA21</p></td>
        <td class="data_point_2"><p>DATA22</p></td>
        <td class="data_point_3"><p>DATA23</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA31</p></td>
        <td class="data_point_2"><p>DATA32</p></td>
        <td class="data_point_3"><p>DATA33</p></td>
    </tr>
    <tr class="header">
        <th colspan="3"><h1>HEADER 2</h1></th>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA11</p></td>
        <td class="data_point_2"><p>DATA12</p></td>
        <td class="data_point_3"><p>DATA13</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA21</p></td>
        <td class="data_point_2"><p>DATA22</p></td>
        <td class="data_point_3"><p>DATA23</p></td>
    </tr>
    <tr>
        <td class="data_point_1"><p>DATA31</p></td>
        <td class="data_point_2"><p>DATA32</p></td>
        <td class="data_point_3"><p>DATA33</p></td>
    </tr>
</table></body>'''

soup = BeautifulSoup(html)
rows = soup.findAll('tr')

section = ''
dataset = {}
for row in rows:
    if row.attrs:
        section = row.text
        dataset[section] = {}
    else:
        cells = row.findAll('td')
        for cell in cells:
            if cell['class'] in dataset[section]:
                dataset[section][ cell['class'] ].append( cell.text )
            else:
                dataset[section][ cell['class'] ] = [ cell.text ]

pprint(dataset)

产生:

{u'HEADER 1': {u'data_point_1': [u'DATA11', u'DATA21', u'DATA31'],
               u'data_point_2': [u'DATA12', u'DATA22', u'DATA32'],
               u'data_point_3': [u'DATA12', u'DATA23', u'DATA33']},
 u'HEADER 2': {u'data_point_1': [u'DATA11', u'DATA21', u'DATA31'],
               u'data_point_2': [u'DATA12', u'DATA22', u'DATA32'],
               u'data_point_3': [u'DATA13', u'DATA23', u'DATA33']}}

修改您的解决方案

您的代码很整洁,只有几个问题。你在你应该使用contents的地方使用text或者findAll- 我在下面修复了它:

soup = BeautifulSoup(html)
credits = soup.find('table')

section = ''
DATA_SET = {}

for i in credits.findAll('tr'):
    if i.get('class', '') == 'header':
        section = i.text
        DATA_SET[section] = {}
    else:
        DATA_SET[section]['data_point_1'] = i.find('td', {'class' : 'data_point_1'}).find('p').contents
        DATA_SET[section]['data_point_2'] = i.find('td', {'class' : 'data_point_2'}).find('p').contents
        DATA_SET[section]['data_point_3'] = i.find('td', {'class' : 'data_point_3'}).find('p').contents

print DATA_SET

请注意,如果连续的单元格具有相同的data_point类别,则连续的行将替换之前的行。我怀疑这在您的真实数据集中不是问题,但这就是为什么您的代码会返回这个缩写结果:

{u'HEADER 2': {'data_point_2': [u'DATA32'],
               'data_point_3': [u'DATA33'],
               'data_point_1': [u'DATA31']},
 u'HEADER 1': {'data_point_2': [u'DATA32'],
               'data_point_3': [u'DATA33'],
               'data_point_1': [u'DATA31']}}
于 2012-05-18T21:37:27.453 回答