MCPcopy Create free account
hub / github.com/cosdata/cosdata / pre_process_vector

Function pre_process_vector

tests/test-dataset-streaming.py:449–455  ·  view source on GitHub ↗
(id, values)

Source from the content-addressed store, hash-verified

447 return results
448
449def pre_process_vector(id, values):
450 corrected_values = [float(v) for v in values]
451 return {
452 "id": str(id), # Keep as string for server compatibility
453 "dense_values": corrected_values,
454 "document_id": f"doc_{id//10}" # Group vectors into documents
455 }
456
457def read_dataset_from_parquet(dataset_name, max_vectors=50000):
458 """Read dataset from parquet files with a strict limit to prevent memory issues."""

Callers 3

read_single_parquet_fileFunction · 0.70

Calls

no outgoing calls

Tested by

no test coverage detected