Parser for Excel files (.xlsx, .xls). This parser extracts text content from Excel files by processing all sheets and converting each row into a structured text format. Each row becomes a separate chunk with key-value pairs. Features: - Supports multiple sheets in a
| 43 | |
| 44 | |
| 45 | class ExcelParser(BaseParser): |
| 46 | """Parser for Excel files (.xlsx, .xls). |
| 47 | |
| 48 | This parser extracts text content from Excel files by processing all sheets |
| 49 | and converting each row into a structured text format. Each row becomes a |
| 50 | separate chunk with key-value pairs. |
| 51 | |
| 52 | Features: |
| 53 | - Supports multiple sheets in a single Excel file |
| 54 | - Automatically removes completely empty rows |
| 55 | - Converts each row to "column: value" format |
| 56 | - Creates individual chunks for each row for better granularity |
| 57 | |
| 58 | Example: |
| 59 | >>> parser = ExcelParser() |
| 60 | >>> with open("data.xlsx", "rb") as f: |
| 61 | ... content = f.read() |
| 62 | ... document = parser.parse_into_text(content) |
| 63 | >>> print(document.content) |
| 64 | Name: John,Age: 30,City: NYC |
| 65 | Name: Jane,Age: 25,City: LA |
| 66 | """ |
| 67 | |
| 68 | def parse_into_text(self, content: bytes) -> Document: |
| 69 | """Parse Excel file bytes into a Document object. |
| 70 | |
| 71 | Args: |
| 72 | content: Raw bytes of the Excel file |
| 73 | |
| 74 | Returns: |
| 75 | Document: Parsed document containing: |
| 76 | - content: Full text with all rows from all sheets |
| 77 | - chunks: List of Chunk objects, one per row |
| 78 | |
| 79 | Note: |
| 80 | - Empty rows (all NaN values) are automatically skipped |
| 81 | - Each row is formatted as: "col1: val1,col2: val2,..." |
| 82 | - Chunks maintain sequential ordering across all sheets |
| 83 | """ |
| 84 | chunks: List[Chunk] = [] |
| 85 | text: List[str] = [] |
| 86 | start, end = 0, 0 |
| 87 | |
| 88 | excel_file = _open_excel_file(content, file_type=self.file_type) |
| 89 | |
| 90 | # Process each sheet in the Excel file |
| 91 | for excel_sheet_name in excel_file.sheet_names: |
| 92 | df = _read_sheet_dataframe(excel_file, excel_sheet_name) |
| 93 | # Remove rows where all values are NaN (completely empty rows) |
| 94 | df.dropna(how="all", inplace=True) |
| 95 | |
| 96 | # Process each row in the DataFrame |
| 97 | for _, row in df.iterrows(): |
| 98 | page_content = [] |
| 99 | # Build key-value pairs for non-null values |
| 100 | for k, v in row.items(): |
| 101 | if pd.notna(v) and not _is_image_function(v): |
| 102 | page_content.append(f"{k}: {v}") |
no outgoing calls