但仍然无法解决。
我有一个带有重复ids的数据
ID Publication_type
1 Journal
1 Clinical study
1 Guideline
2 Journal
2 Letter 我想把它写得更宽,但我不知道我会有多少种出版物--也许是2种,也许是20种。因此,我不知道我需要多少栏目宽。publication_type的宽列的最大大小不能超过每个id的类型数。
预期产出
ID Publication_type1 Publication_type2 Publication_type 3 etc
1 Journal Clinical Study Guideline
2 Journal Letter NaN现在,我不需要将相同的发布类型放入同一列中。我不需要在同一栏中的所有文章。谢谢!
发布于 2021-12-22 18:11:32
您可以按ID分组,通过list进行聚合,然后根据结果创建一个新的DataFrame:
col = 'Publication_type'
new_df = pd.DataFrame(df.groupby('ID')[col].agg(lambda x: x.tolist()).tolist()).replace({None: np.nan})
new_df.columns = [f'{col}{i}' for i in new_df.columns + 1]
new_df['ID'] = df['ID'].drop_duplicates().reset_index(drop=True)输出:
>>> df
Publication_type1 Publication_type2 Publication_type3 ID
0 Journal Clinical-study Guideline 1
1 Journal Letter NaN 2https://stackoverflow.com/questions/70453310
复制相似问题