首页
学习
活动
专区
圈层
工具
发布
社区首页 >问答首页 >将HDF5读取为Dataframe时出错,为什么?

将HDF5读取为Dataframe时出错,为什么?
EN

Stack Overflow用户
提问于 2020-05-29 15:14:48
回答 1查看 194关注 0票数 1

1.我的问题

当我试图使用Dask读取我的HDF5文件时,我得到了下一个错误,我不知道为什么

代码语言:javascript
复制
>>> dd.read_hdf("test.h5", key="/RECORDS/STATES")
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
  File "/usr/local/lib/python3.7/site-packages/dask/dataframe/io/hdf.py", line 514, in read_hdf
    for path in paths
  File "/usr/local/lib/python3.7/site-packages/dask/dataframe/io/hdf.py", line 514, in <listcomp>
    for path in paths
  File "/usr/local/lib/python3.7/site-packages/dask/dataframe/io/hdf.py", line 382, in _read_single_hdf
    for k, s, d in zip(keys, stops, divisions)
  File "/usr/local/lib/python3.7/site-packages/dask/dataframe/multi.py", line 1071, in concat
    raise ValueError("No objects to concatenate")
ValueError: No objects to concatenate

2. HDF5文件

我要用Dask读取的文件是我使用HDF5的C生成的。如果您问一问,为了性能起见,我使用C而不是Python (numpy,HDF5 )生成,因为我需要在ASCII中解析许多GB的未格式化数据。数据以HDF5表(https://portal.hdfgroup.org/display/HDF5/Tables)的形式存储在文件中。我的文件头看起来如下:

代码语言:javascript
复制
HDF5 "rhoPimpleExtrae10TimeSteps.00.1iter.h5" {
GROUP "/" {
   ATTRIBUTE "hdf5_metadata_apps" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SCALAR
   }
   ATTRIBUTE "hdf5_metadata_date" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SCALAR
   }
   ATTRIBUTE "hdf5_metadata_hwcpu" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SIMPLE { ( 48 ) / ( 48 ) }
   }
   ATTRIBUTE "hdf5_metadata_hwnodes" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SIMPLE { ( 1 ) / ( 1 ) }
   }
   ATTRIBUTE "hdf5_metadata_name" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SCALAR
   }
   ATTRIBUTE "hdf5_metadata_nodes" {
      DATATYPE  H5T_STD_I64LE
      DATASPACE  SIMPLE { ( 1 ) / ( 1 ) }
   }
   ATTRIBUTE "hdf5_metadata_path" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SCALAR
   }
   ATTRIBUTE "hdf5_metadata_threads" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SIMPLE { ( 48 ) / ( 48 ) }
   }
   ATTRIBUTE "hdf5_metadata_time" {
      DATATYPE  H5T_STD_I64LE
      DATASPACE  SCALAR
   }
   ATTRIBUTE "hdf5_metadata_type" {
      DATATYPE  H5T_STRING {
         STRSIZE H5T_VARIABLE;
         STRPAD H5T_STR_NULLTERM;
         CSET H5T_CSET_UTF8;
         CTYPE H5T_C_S1;
      }
      DATASPACE  SCALAR
   }
   GROUP "RECORDS" {
      DATASET "COMMUNICATIONS" {
         DATATYPE  H5T_COMPOUND {
            H5T_STD_U32LE "CPU Send ID";
            H5T_STD_U32LE "Phy. Task Send ID";
            H5T_STD_U32LE "Log. Task Send ID";
            H5T_STD_U32LE "Thread Send ID";
            H5T_STD_U64LE "Log. Send Time";
            H5T_STD_U64LE "Phy. Send Time";
            H5T_STD_U32LE "CPU Receive ID";
            H5T_STD_U32LE "Phy. Task Receive ID";
            H5T_STD_U32LE "Log. Task Receive ID";
            H5T_STD_U32LE "Thread Receive ID";
            H5T_STD_U64LE "Log. Receive Time";
            H5T_STD_U64LE "Phy. Receive Time";
            H5T_STD_U64LE "Size";
            H5T_STD_U64LE "Tag";
         }
         DATASPACE  SIMPLE { ( 67574 ) / ( H5S_UNLIMITED ) }
         ATTRIBUTE "CLASS" {
            DATATYPE  H5T_STRING {
               STRSIZE 6;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_0_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 12;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_10_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 18;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_11_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 18;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_12_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 5;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_13_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 4;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_1_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 18;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_2_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 18;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_3_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 15;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_4_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 15;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_5_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 15;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_6_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 15;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_7_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 21;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_8_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 21;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_9_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 18;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "TITLE" {
            DATATYPE  H5T_STRING {
               STRSIZE 22;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "VERSION" {
            DATATYPE  H5T_STRING {
               STRSIZE 4;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
      }
      DATASET "EVENTS" {
         DATATYPE  H5T_COMPOUND {
            H5T_STD_U32LE "CPU ID";
            H5T_STD_U16LE "APP ID";
            H5T_STD_U32LE "Task ID";
            H5T_STD_U32LE "Thread ID";
            H5T_STD_U64LE "Time";
            H5T_STD_U64LE "Event Type";
            H5T_STD_U64LE "Event Value";
         }
         DATASPACE  SIMPLE { ( 3643006 ) / ( H5S_UNLIMITED ) }
         ATTRIBUTE "CLASS" {
            DATATYPE  H5T_STRING {
               STRSIZE 6;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_0_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 7;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_1_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 7;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_2_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 8;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_3_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 10;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_4_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 5;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_5_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 11;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_6_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 12;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "TITLE" {
            DATATYPE  H5T_STRING {
               STRSIZE 14;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "VERSION" {
            DATATYPE  H5T_STRING {
               STRSIZE 4;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
      }
      DATASET "STATES" {
         DATATYPE  H5T_COMPOUND {
            H5T_STD_U32LE "CPU ID";
            H5T_STD_U16LE "APP ID";
            H5T_STD_U32LE "Task ID";
            H5T_STD_U32LE "Thread ID";
            H5T_STD_U64LE "Time ini";
            H5T_STD_U64LE "Time fi";
            H5T_STD_U16LE "State";
         }
         DATASPACE  SIMPLE { ( 301496 ) / ( H5S_UNLIMITED ) }
         ATTRIBUTE "CLASS" {
            DATATYPE  H5T_STRING {
               STRSIZE 6;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_0_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 7;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_1_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 7;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_2_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 8;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_3_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 10;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_4_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 9;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_5_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 8;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "FIELD_6_NAME" {
            DATATYPE  H5T_STRING {
               STRSIZE 6;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "TITLE" {
            DATATYPE  H5T_STRING {
               STRSIZE 14;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
         ATTRIBUTE "VERSION" {
            DATATYPE  H5T_STRING {
               STRSIZE 4;
               STRPAD H5T_STR_NULLTERM;
               CSET H5T_CSET_ASCII;
               CTYPE H5T_C_S1;
            }
            DATASPACE  SCALAR
         }
      }
   }
}
}

在/RECORDS下,我基本上有3个数据集(状态、事件和通信)。我想我的HDF5没有什么奇怪的地方。我尝试过使用Pandas和Dask数组加载这些数据集,它可以工作。

3.我想知道的

我的HDF5文件有什么问题吗? Dask无法将它作为数据文件读取?

我试图在Dask文档中找到HDF5文件必须满足的要求,但是没有任何涉及这个主题的内容。如果我至少知道我的文件有什么问题,我就能解决它。

EN

回答 1

Stack Overflow用户

回答已采纳

发布于 2020-05-29 15:48:52

PR https://github.com/dask/dask/pull/6204最近被合并成了达斯克大师,幸运的是,它为你解决了这个问题。

票数 1
EN
页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持
原文链接:

https://stackoverflow.com/questions/62089195

复制
相关文章

相似问题

领券
问题归档专栏文章快讯文章归档关键词归档开发者手册归档开发者手册 Section 归档