我在这里的第一篇文章,我希望它不会太长或太详细。
当我尝试阅读和解释以下 ascii 表(从一个更大的表中简单提取)时,我遇到了一个问题。
假设文件名为“test.txt”:
A B C D E
0 992 CEN/4 -2.657293E+00 -3.309567E+01 4.697218E-01
1291 -3.368449E+00 7.837483E+00 2.311393E+00
800 -3.530800E+00 -7.392188E+01 -1.401380E+00
801 -1.952177E+00 -7.392114E+01 -1.367195E+00
1290 -1.777301E+00 7.838229E+00 2.345850E+00
0 994 CEN/4 7.270955E+00 -6.637891E+01 -1.293553E+01
1110 5.816999E+00 -3.981042E+01 -1.504738E+01
785 5.535329E+00 -9.246339E+01 -1.061554E+01
786 8.625161E+00 -9.163719E+01 -1.092563E+01
1109 9.080517E+00 -4.059749E+01 -1.523589E+01
使用 python 2.7 和 pandas (0.9.1),我可以阅读如下:
>>> r=pd.read_fwf('test.txt', widths=(3, 8, 9, 14, 14, 14),skiprows=1, header=None)
>>> print r
X0 X1 X2 X3 X4 X5
0 0 992 CEN/4 -2.657293 -33.095670 0.469722
1 NaN NaN 1291 -3.368449 7.837483 2.311393
2 NaN NaN 800 -3.530800 -73.921880 -1.401380
3 NaN NaN 801 -1.952177 -73.921140 -1.367195
4 NaN NaN 1290 -1.777301 7.838229 2.345850
5 0 994 CEN/4 7.270955 -66.378910 -12.935530
6 NaN NaN 1110 5.816999 -39.810420 -15.047380
7 NaN NaN 785 5.535329 -92.463390 -10.615540
8 NaN NaN 786 8.625161 -91.637190 -10.925630
9 NaN NaN 1109 9.080517 -40.597490 -15.235890
尝试将其直接读取为“分层表”:
>>> r=pd.read_fwf('test.txt', widths=(3, 8, 9, 14, 14, 14), skiprows=1, index_col=[0,1,2], header=None)
>>> print r
X3 X4 X5
X0 X1 X2
0 992 CEN/4 -2.657293 -33.095670 0.469722
994 1291 -3.368449 7.837483 2.311393
800 -3.530800 -73.921880 -1.401380
801 -1.952177 -73.921140 -1.367195
1290 -1.777301 7.838229 2.345850
CEN/4 7.270955 -66.378910 -12.935530
1110 5.816999 -39.810420 -15.047380
785 5.535329 -92.463390 -10.615540
786 8.625161 -91.637190 -10.925630
1109 9.080517 -40.597490 -15.235890
我的目标是获得:
>>> print r
X3 X4 X5
X0 X1 X2
0 992 CEN/4 -2.657293 -33.095670 0.469722
1291 -3.368449 7.837483 2.311393
800 -3.530800 -73.921880 -1.401380
801 -1.952177 -73.921140 -1.367195
1290 -1.777301 7.838229 2.345850
994 CEN/4 7.270955 -66.378910 -12.935530
1110 5.816999 -39.810420 -15.047380
785 5.535329 -92.463390 -10.615540
786 8.625161 -91.637190 -10.925630
1109 9.080517 -40.597490 -15.235890
有没有使用熊猫的简单方法,还是我必须在解析之前“按摩” ascii 表以获得我想要的东西?提前致谢