MCPcopy Create free account
hub / github.com/Dod-o/Statistical-Learning-Method_Code / do_lsa

Function do_lsa

LSA/LSA.py:78–108  ·  view source on GitHub ↗

INPUT: X - (array) 单词-文本矩阵 k - (int) 设定的话题数 words - (list) 单词列表 OUTPUT: topics - (list) 生成的话题列表

(X, k, words)

Source from the content-addressed store, hash-verified

76
77#定义潜在语义分析函数
78def do_lsa(X, k, words):
79 '''
80 INPUT:
81 X - (array) 单词-文本矩阵
82 k - (int) 设定的话题数
83 words - (list) 单词列表
84
85 OUTPUT:
86 topics - (list) 生成的话题列表
87
88 '''
89 w, v = np.linalg.eig(np.matmul(X.T, X)) #计算Sx的特征值和特征向量,其中Sx=X.T*X,Sx的特征值w即为X的奇异值分解的奇异值,v即为对应的奇异向量
90 sort_inds = np.argsort(w)[::-1] #对特征值降序排列后取出对应的索引值
91 w = np.sort(w)[::-1] #对特征值降序排列
92 V_T = [] #用来保存矩阵V的转置
93 for ind in sort_inds:
94 V_T.append(v[ind]/np.linalg.norm(v[ind])) #将降序排列后各特征值对应的特征向量单位化后保存到V_T中
95 V_T = np.array(V_T) #将V_T转换为数组,方便之后的操作
96 Sigma = np.diag(np.sqrt(w)) #将特征值数组w转换为对角矩阵,即得到SVD分解中的Sigma
97 U = np.zeros((len(words), k)) #用来保存SVD分解中的矩阵U
98 for i in range(k):
99 ui = np.matmul(X, V_T.T[:, i]) / Sigma[i][i] #计算矩阵U的第i个列向量
100 U[:, i] = ui #保存到矩阵U中
101 topics = [] #用来保存k个话题
102 for i in range(k):
103 inds = np.argsort(U[:, i])[::-1] #U的每个列向量表示一个话题向量,话题向量的长度为m,其中每个值占向量值之和的比重表示对应单词在当前话题中所占的比重,这里对第i个话题向量的值降序排列后取出对应的索引值
104 topic = [] #用来保存第i个话题
105 for j in range(10):
106 topic.append(words[inds[j]]) #根据索引inds取出当前话题中比重最大的10个单词作为第i个话题
107 topics.append(' '.join(topic)) #保存话题i
108 return topics
109
110
111if __name__ == "__main__":

Callers 1

LSA.pyFile · 0.85

Calls

no outgoing calls

Tested by

no test coverage detected