我有一个有2列的DataFrame
combined | out_dict
--------------------
"sentence 1" | {class1: 30, class2: 42, class3: 5, class4:42}
"sentence 2" | {class1: 60, class2: 10, class3: 40, class4:40}现在,我想把它转换成一个新的数据格式,如下所示。如果概率大于或等于40,则为其创建两个新列。i.e
combined | class1 | class1_score | class2 | class2_score | class3 | class3_score |class4 | class4_score
-------------------------------------------------------------------------------------------------------
"sentence 1" | Nan | NaN | 1 | 42 | Nan | Nan | 1 | 42
"sentence 2" | 1 | 60 | Nan | Nan | 1 | 40 | 1 | 40我尝试过使用一些手动逻辑,比如遍历整个数据帧,但这看起来效率很低。请指导我为这个问题创造一个最佳的解决方案。
编辑
我的解决方案:
self.df[list(ref_garm_dict.keys())] = "" # a dictionary with classes name
for _, row in self.df.iterrows():
for key, value in row["out_dict"].items():
if value >= 40:
row[key] = 1
row[key + "_score"] = value # has no effect so _score columns are not created.发布于 2022-02-08 03:42:44
您可以首先将所有字典按原样转换为列。
df = df.set_index('combined')['out_dict'].apply(pd.Series)然后根据需要添加/修改列
for col in df.columns:
score_name = f'{col}_score'
df[score_name] = df[col]
df[score_name].loc[df[col]<40] = np.nan
df[col] = df[score_name].isna().map({False: 1})最后,处理格式。
df.sort_index(axis=1).reset_index()https://stackoverflow.com/questions/71023716
复制相似问题