我正在寻找一种方法,从他们最长的重复模式中清除字符串。
我有一个大约1000个网页标题的列表,他们都有一个共同的后缀,这是网站的名称。
他们遵循这样的模式:
['art gallery - museum and visits | expand knowledge',
'lasergame - entertainment | expand knowledge',
'coffee shop - confort and food | expand knowledge',
...
]我如何能够自动从所有字符串的公共后缀" | expand knowledge"中删除?
谢谢!
编辑:对不起,我没说清楚。我事先没有关于" | expand knowledge"后缀的信息。我希望能够清除一个潜在的公共后缀字符串列表,即使我不知道它是什么。
发布于 2012-11-19 20:28:35
下面是一个在反向标题上使用os.path.commonprefix函数的解决方案:
titles = ['art gallery - museum and visits | expand knowledge',
'lasergame - entertainment | expand knowledge',
'coffee shop - confort and food | expand knowledge',
]
# Find the longest common suffix by reversing the strings and using a
# library function to find the common "prefix".
common_suffix = os.path.commonprefix([title[::-1] for title in titles])[::-1]
# Strips all titles from the number of characters in the common suffix.
stripped_titles = [title[:-len(common_suffix)] for title in titles]结果:
“美术馆-博物馆和参观”、“激光游戏-娱乐”、“咖啡厅-康福和食品”
因为它自己找到了公共后缀,所以它应该在任何一组标题上工作,即使你不知道后缀。
发布于 2012-11-19 20:27:03
如果你真的知道你想要去掉的后缀,你可以简单地做:
suffix = " | expand knowledge"
your_list = ['art gallery - museum and visits | expand knowledge',
'lasergame - entertainment | expand knowledge',
'coffee shop - confort and food | expand knowledge',
...]
new_list = [name.rstrip(suffix) for name in your_list]发布于 2012-11-19 20:25:22
如果您确信所有字符串都有公共后缀,那么这将起到以下作用:
strings = [
'art gallery - museum and visits | expand knowledge',
'lasergame - entertainment | expand knowledge']
suffixlen = len(" | expand knowledge")
print [s[:-suffixlen] for s in strings] 产出:
['art gallery - museum and visits', 'lasergame - entertainment']https://stackoverflow.com/questions/13461398
复制相似问题