首页
学习
活动
专区
圈层
工具
发布
社区首页 >问答首页 >随机指标函数(聚类性能评估)

随机指标函数(聚类性能评估)
EN

Stack Overflow用户
提问于 2018-03-31 18:28:04
回答 2查看 10.5K关注 0票数 7

据我所知,python中没有可用于Rand Index的软件包,而对于调整后的Rand Index,您可以选择使用sklearn.metrics.adjusted_rand_score(labels_true, labels_pred)

我为Rand Score编写了代码,我将把它作为这篇文章的答案与其他人分享。

EN

回答 2

Stack Overflow用户

回答已采纳

发布于 2018-03-31 18:28:04

代码语言:javascript
复制
from scipy.misc import comb
from itertools import combinations
import numpy as np

def check_clusterings(labels_true, labels_pred):
    """Check that the two clusterings matching 1D integer arrays."""
    labels_true = np.asarray(labels_true)
    labels_pred = np.asarray(labels_pred)    
    # input checks
    if labels_true.ndim != 1:
        raise ValueError(
            "labels_true must be 1D: shape is %r" % (labels_true.shape,))
    if labels_pred.ndim != 1:
        raise ValueError(
            "labels_pred must be 1D: shape is %r" % (labels_pred.shape,))
    if labels_true.shape != labels_pred.shape:
        raise ValueError(
            "labels_true and labels_pred must have same size, got %d and %d"
            % (labels_true.shape[0], labels_pred.shape[0]))
    return labels_true, labels_pred

def rand_score (labels_true, labels_pred):
"""given the true and predicted labels, it will return the Rand Index."""
    check_clusterings(labels_true, labels_pred)
    my_pair = list(combinations(range(len(labels_true)), 2)) #create list of all combinations with the length of labels.
    def is_equal(x):
        return (x[0]==x[1])
    my_a = 0
    my_b = 0
    for i in range(len(my_pair)):
            if(is_equal((labels_true[my_pair[i][0]],labels_true[my_pair[i][1]])) == is_equal((labels_pred[my_pair[i][0]],labels_pred[my_pair[i][1]])) 
               and is_equal((labels_pred[my_pair[i][0]],labels_pred[my_pair[i][1]])) == True):
                my_a += 1
            if(is_equal((labels_true[my_pair[i][0]],labels_true[my_pair[i][1]])) == is_equal((labels_pred[my_pair[i][0]],labels_pred[my_pair[i][1]])) 
               and is_equal((labels_pred[my_pair[i][0]],labels_pred[my_pair[i][1]])) == False):
                my_b += 1
    my_denom = comb(len(labels_true),2)
    ri = (my_a + my_b) / my_denom
    return ri

举个简单的例子:

代码语言:javascript
复制
labels_true = [1, 1, 0, 0, 0, 0]
labels_pred = [0, 0, 0, 1, 0, 1]
rand_score (labels_true, labels_pred)
#0.46666666666666667

可能有一些方法可以改进它,使其更具蟒蛇效应。如果你有任何建议,你可以改进它。

我找到了看起来更快的this implementation

代码语言:javascript
复制
import numpy as np
from scipy.misc import comb
def rand_index_score(clusters, classes):
    tp_plus_fp = comb(np.bincount(clusters), 2).sum()
    tp_plus_fn = comb(np.bincount(classes), 2).sum()
    A = np.c_[(clusters, classes)]
    tp = sum(comb(np.bincount(A[A[:, 0] == i, 1]), 2).sum()
             for i in set(clusters))
    fp = tp_plus_fp - tp
    fn = tp_plus_fn - tp
    tn = comb(len(A), 2) - tp - fp - fn
    return (tp + tn) / (tp + fp + fn + tn)

举个简单的例子:

代码语言:javascript
复制
labels_true = [1, 1, 0, 0, 0, 0]
labels_pred = [0, 0, 0, 1, 0, 1]
rand_index_score (labels_true, labels_pred)
#0.46666666666666667
票数 6
EN

Stack Overflow用户

发布于 2021-11-25 10:55:24

从scikit Learn0.24.0开始,添加了sklearn.metrics.rand_score函数,实现了(未调整的)兰德索引。请检查changelog

你要做的就是:

代码语言:javascript
复制
from sklearn.metrics import rand_score

rand_score(labels_true, labels_pred)

labels_truelabels_pred可以在不同的域中具有值。例如:

代码语言:javascript
复制
>>> rand_score(['a', 'b', 'c'], [5, 6, 7])
1.0
票数 1
EN
页面原文内容由Stack Overflow提供。腾讯云小微IT领域专用引擎提供翻译支持
原文链接:

https://stackoverflow.com/questions/49586742

复制
相关文章

相似问题

领券
问题归档专栏文章快讯文章归档关键词归档开发者手册归档开发者手册 Section 归档