MCPcopy Create free account
hub / github.com/Tencent/WeKnora / ExcelParser

Class ExcelParser

docreader/parser/excel_parser.py:45–120  ·  view source on GitHub ↗

Parser for Excel files (.xlsx, .xls). This parser extracts text content from Excel files by processing all sheets and converting each row into a structured text format. Each row becomes a separate chunk with key-value pairs. Features: - Supports multiple sheets in a

Source from the content-addressed store, hash-verified

43
44
45class ExcelParser(BaseParser):
46 """Parser for Excel files (.xlsx, .xls).
47
48 This parser extracts text content from Excel files by processing all sheets
49 and converting each row into a structured text format. Each row becomes a
50 separate chunk with key-value pairs.
51
52 Features:
53 - Supports multiple sheets in a single Excel file
54 - Automatically removes completely empty rows
55 - Converts each row to "column: value" format
56 - Creates individual chunks for each row for better granularity
57
58 Example:
59 >>> parser = ExcelParser()
60 >>> with open("data.xlsx", "rb") as f:
61 ... content = f.read()
62 ... document = parser.parse_into_text(content)
63 >>> print(document.content)
64 Name: John,Age: 30,City: NYC
65 Name: Jane,Age: 25,City: LA
66 """
67
68 def parse_into_text(self, content: bytes) -> Document:
69 """Parse Excel file bytes into a Document object.
70
71 Args:
72 content: Raw bytes of the Excel file
73
74 Returns:
75 Document: Parsed document containing:
76 - content: Full text with all rows from all sheets
77 - chunks: List of Chunk objects, one per row
78
79 Note:
80 - Empty rows (all NaN values) are automatically skipped
81 - Each row is formatted as: "col1: val1,col2: val2,..."
82 - Chunks maintain sequential ordering across all sheets
83 """
84 chunks: List[Chunk] = []
85 text: List[str] = []
86 start, end = 0, 0
87
88 excel_file = _open_excel_file(content, file_type=self.file_type)
89
90 # Process each sheet in the Excel file
91 for excel_sheet_name in excel_file.sheet_names:
92 df = _read_sheet_dataframe(excel_file, excel_sheet_name)
93 # Remove rows where all values are NaN (completely empty rows)
94 df.dropna(how="all", inplace=True)
95
96 # Process each row in the DataFrame
97 for _, row in df.iterrows():
98 page_content = []
99 # Build key-value pairs for non-null values
100 for k, v in row.items():
101 if pd.notna(v) and not _is_image_function(v):
102 page_content.append(f"{k}: {v}")

Calls

no outgoing calls