MCPcopy Create free account
hub / github.com/Tencent/WeKnora / WebParser

Class WebParser

docreader/parser/web_parser.py:261–270  ·  view source on GitHub ↗

Web parser using pipeline pattern. This parser chains StdWebParser (for web scraping and HTML to markdown conversion) with MarkdownParser (for markdown processing). The pipeline processes content sequentially through both parsers.

Source from the content-addressed store, hash-verified

259
260
261class WebParser(PipelineParser):
262 """Web parser using pipeline pattern.
263
264 This parser chains StdWebParser (for web scraping and HTML to markdown conversion)
265 with MarkdownParser (for markdown processing). The pipeline processes content
266 sequentially through both parsers.
267 """
268
269 # Parser classes to be executed in sequence
270 _parser_cls = (StdWebParser, MarkdownParser)
271
272
273if __name__ == "__main__":

Callers 2

parse_urlMethod · 0.90
web_parser.pyFile · 0.85

Calls

no outgoing calls

Tested by

no test coverage detected