首页
学习
活动
专区
圈层
工具
发布
社区首页 >问答首页 >RNA到蛋白质程序问题

RNA到蛋白质程序问题
EN

Stack Overflow用户
提问于 2015-05-12 17:40:11
回答 2查看 458关注 0票数 0

我的代码有一些问题,我希望得到一些帮助。

程序的第一部分是为了验证用户的输入,所以他们不能输入任何东西,只能输入part (或小写)。但是,如果我输入了其他内容,我会收到一条非常长的错误消息,但我只想要重新启动函数验证检查()的程序。

另外,如果用户输入了一个有效的序列,出于某种原因,我的代码没有将有效的RNA序列转换为蛋白质序列。我认为,这可能与块函数有关,该函数将input_rna中的str分隔为3个字母。

代码语言:javascript
复制
import re

input_rna = input("Type RNA sequence: ")

def chunks(l, n):
    for i in range(0, len(l), n):
        yield l[i:i+n]

def translate():
    amino_acids = {"UUU":"F", "UUC":"F", "UUA":"L", "UUG":"L",
        "UCU":"S", "UCC":"s", "UCA":"S", "UCG":"S",
        "UAU":"Y", "UAC":"Y", "UAA":"STOP", "UAG":"STOP",
        "UGU":"C", "UGC":"C", "UGA":"STOP", "UGG":"W",
        "CUU":"L", "CUC":"L", "CUA":"L", "CUG":"L",
        "CCU":"P", "CCC":"P", "CCA":"P", "CCG":"P",
        "CAU":"H", "CAC":"H", "CAA":"Q", "CAG":"Q",
        "CGU":"R", "CGC":"R", "CGA":"R", "CGG":"R",
        "AUU":"I", "AUC":"I", "AUA":"I", "AUG":"M",
        "ACU":"T", "ACC":"T", "ACA":"T", "ACG":"T",
        "AAU":"N", "AAC":"N", "AAA":"K", "AAG":"K",
        "AGU":"S", "AGC":"S", "AGA":"R", "AGG":"R",
        "GUU":"V", "GUC":"V", "GUA":"V", "GUG":"V",
        "GCU":"A", "GCC":"A", "GCA":"A", "GCG":"A",
        "GAU":"D", "GAC":"D", "GAA":"E", "GAG":"E",
        "GGU":"G", "GGC":"G", "GGA":"G", "GGG":"G",}

    translated = "".join(amino_acids[i] for i in chunks("".join(input_rna), 3))

def validation_check():
    global input_rna
    if re.match(r"[A, U, G, C, T, a, u, g, c, t]", input_rna):
        print("Correct! That is a valid sequence.")
        translate()
    else:
        print("That is not a valid RNA sequence, please try again.")
        validation_check()
validation_check()
EN

回答 2

Stack Overflow用户

回答已采纳

发布于 2015-05-12 18:21:50

除了指出的其他问题外,您的validation_check函数不允许用户再次输入字符串。这意味着您将不断地尝试一次又一次地验证它,而不会更改它。

你可能想做的事情更像是:

代码语言:javascript
复制
def validation_check():
    input_rna = raw_input("Type RNA sequence: ").upper()
    if re.match(r"^[AUGCT]+$", input_rna):
        print("Correct! That is a valid sequence.")
        print translate(input_rna)
    else:
        print("That is not a valid RNA sequence, please try again.")
        validation_check()

这避免了使用全局,允许用户重新放置,并且不会自动导致无限循环。

(尽管如此,在这里使用递归可能是不好的,所以您应该考虑将其实现为while循环。)

你会注意到另外几件事:

  • raw_input而不是input,因为后者有一个隐式eval。除非你绝对需要它,否则你想避开它。
  • .upper(),因此您可以使用标准化的字符串来验证和关闭。由于基字典只使用大写字符串,这比其他地方推荐的使用re.I更有意义。
  • 我让translate返回翻译后的蛋白质,然后打印出来。你可能还想做点别的。

我还在字典查找中添加了一个默认值:

代码语言:javascript
复制
translated = "".join(amino_acids.get(i, '!') for i in chunks("".join(rna), 3)) 

这样,如果您得到了一些奇怪的东西,就可以尝试继续处理,而不必处理KeyError (如果用户输入一个您没有键的序列,比如'CUT'),就会引发这种处理。

我还注意到您允许(但不要翻译)基本的'T'。你也许想调查一下。

总之,我最后得到的完整代码是:

代码语言:javascript
复制
import re

def chunks(l, n): 
    for i in range(0, len(l), n): 
        # print i
        chunk = l[i:i+n]
        # print chunk
        yield l[i:i+n]

def translate(rna):
    amino_acids = {"UUU":"F", "UUC":"F", "UUA":"L", "UUG":"L",
        "UCU":"S", "UCC":"s", "UCA":"S", "UCG":"S",
        "UAU":"Y", "UAC":"Y", "UAA":"STOP", "UAG":"STOP",
        "UGU":"C", "UGC":"C", "UGA":"STOP", "UGG":"W",
        "CUU":"L", "CUC":"L", "CUA":"L", "CUG":"L",
        "CCU":"P", "CCC":"P", "CCA":"P", "CCG":"P",
        "CAU":"H", "CAC":"H", "CAA":"Q", "CAG":"Q",
        "CGU":"R", "CGC":"R", "CGA":"R", "CGG":"R",
        "AUU":"I", "AUC":"I", "AUA":"I", "AUG":"M",
        "ACU":"T", "ACC":"T", "ACA":"T", "ACG":"T",
        "AAU":"N", "AAC":"N", "AAA":"K", "AAG":"K",
        "AGU":"S", "AGC":"S", "AGA":"R", "AGG":"R",
        "GUU":"V", "GUC":"V", "GUA":"V", "GUG":"V",
        "GCU":"A", "GCC":"A", "GCA":"A", "GCG":"A",
        "GAU":"D", "GAC":"D", "GAA":"E", "GAG":"E",
        "GGU":"G", "GGC":"G", "GGA":"G", "GGG":"G",}
    translated = "".join(amino_acids.get(i, '!') for i in chunks("".join(rna), 3)) 
    return translated

def validation_check():
    input_rna = raw_input("Type RNA sequence: ").upper()
    if re.match(r"^[AUGCT]+$", input_rna):
        print("Correct! That is a valid sequence.")
        print translate(input_rna)
    else:
        print("That is not a valid RNA sequence, please try again.")
        validation_check()

# in case you ever need to import this, don't always call validation_check
if __name__ == "__main__":
     validation_check()
票数 3
EN

Stack Overflow用户

发布于 2015-05-12 17:48:05

正则表达式是错误的,请尝试:

代码语言:javascript
复制
if re.match(r"^[AUGCT]+$", input_rna, re.IGNORECASE):

以下是更好的,因为在RNA中Uracil而不是Thyamine ..。

代码语言:javascript
复制
if re.match(r"^[AUGC]+$", input_rna, re.IGNORECASE):

注意:算法翻译也有问题

代码语言:javascript
复制
list(chunks("".join(input_rna), 3))

你得到:

代码语言:javascript
复制
['ACG', 'AUG', 'AGU', 'CAU', 'GCU', 'U']

最后由"ACGAUGAGUCAUGCUU“提出的问题,如果长度不是3的倍数

解决办法:

代码语言:javascript
复制
"".join(amino_acids[i] for i in chunks("".join(input_rna), 3) if len(i)==3)
票数 3
EN
页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持
原文链接:

https://stackoverflow.com/questions/30197902

复制
相关文章

相似问题

领券
问题归档专栏文章快讯文章归档关键词归档开发者手册归档开发者手册 Section 归档