MCPcopy Create free account
hub / github.com/BIT-DataLab/LakeBench / pre_multi_process

Function pre_multi_process

join/LSH/LSH_benchmark_wlq.py:147–164  ·  view source on GitHub ↗
(file_list, q, id)

Source from the content-addressed store, hash-verified

145
146
147def pre_multi_process(file_list, q, id):
148 # print("{} start".format(id))
149 sizes = []
150 for i, sets_file in enumerate(file_list):
151 try:
152 df = pd.read_csv(sets_file, dtype='str', lineterminator='\n').dropna()
153 except pd.errors.ParserError as e:
154 print(sets_file)
155 data = df.values.T.tolist()
156 for vals in data:
157 # 需要对value去重
158 sizes.append(len(set(vals)))
159 if id==0:
160 sys.stdout.write("\rId 0 Process Read and pre {}/{} files".format(i+1, len(file_list)))
161 if id==0:
162 sys.stdout.write("\n")
163 q.put((sizes,))
164 # print("{} end,size={}".format(id,len(sizes)))
165
166def minhash_multi_process(file_list, q, param_dic, id):
167 # print("{} start".format(id))

Callers

nothing calls this directly

Calls 1

writeMethod · 0.80

Tested by

no test coverage detected