我有两个列、集群标题和它们所属的章节的dataframe。我想创建第三列,其中包含本章中该集群的“顺序”或位置。
因此,我想转一下以下数据:
cluster_title, chapter
"rabbits", 1
"horses", 1
"cows", 1
"trains", 2
"airplanes", 2
"ships", 2
"carrot", 3
"potato", 3
"tomato", 3变成这样的东西:
cluster_title, chapter, position_in_chapter,
"rabbits", 1, 1
"horses" 1, 2
"cows", 1, 3
"trains", 2, 1
"airplanes", 2, 2
"ships", 2, 3
"carrot", 3, 1
"potato", 3, 2
"tomato", 3, 3我尝试使用group_by函数并以某种方式使用索引,但要么我遗漏了一些显而易见的东西(很可能),要么是错误的方法,因为结果对象需要额外的步骤,这些步骤似乎使我朝着错误的方向前进。
有人能给我指明正确的方向吗?
发布于 2021-09-21 16:35:32
尝试使用groupby和cumcount
df["position_in_chapter"] = df.groupby("chapter").cumcount()+1
>>> df
cluster_title chapter position_in_chapter
0 rabbits 1 1
1 horses 1 2
2 cows 1 3
3 trains 2 1
4 airplanes 2 2
5 ships 2 3
6 carrot 3 1
7 potato 3 2
8 tomato 3 3https://stackoverflow.com/questions/69272357
复制相似问题