MCPcopy Create free account
hub / github.com/Tencent/WeKnora / MarkdownImageUtil

Class MarkdownImageUtil

docreader/parser/markdown_parser.py:227–413  ·  view source on GitHub ↗

Utility class for handling images in Markdown. This class provides functionality to: - Extract base64-encoded images from Markdown - Extract image paths from Markdown - Replace image paths with new URLs - Convert base64 images to binary format Supported formats: - Base6

Source from the content-addressed store, hash-verified

225
226
227class MarkdownImageUtil:
228 """Utility class for handling images in Markdown.
229
230 This class provides functionality to:
231 - Extract base64-encoded images from Markdown
232 - Extract image paths from Markdown
233 - Replace image paths with new URLs
234 - Convert base64 images to binary format
235
236 Supported formats:
237 - Base64 embedded images: ![alt](data:image/png;base64,iVBORw0...)
238 - Regular image links: ![alt](path/to/image.png)
239 """
240
241 def __init__(self):
242 # Pattern to match base64 embedded images
243 # Captures: (1) alt text, (2) image format, (3) base64 data
244 # Alt text uses .*? (non-greedy) to allow literal ] (e.g. Windows paths).
245 # MIME subtype uses [^;]+ to handle types with hyphens like x-emf.
246 self.b64_pattern = re.compile(
247 r"!\[(.*?)\]\(data:image/([^;]+);base64,([^\)]+)\)"
248 )
249 # Pattern to match regular image syntax (alt text allows ])
250 self.image_pattern = re.compile(r"!\[(.*?)\]\(([^)]+)\)")
251 # Pattern for replacing image paths
252 self.replace_pattern = re.compile(r"!\[(.*?)\]\(([^)]+)\)")
253
254 def extract_image(
255 self,
256 content: str,
257 path_prefix: Optional[str] = None,
258 replace: bool = True,
259 ) -> Tuple[str, List[str]]:
260 """Extract image paths from Markdown content.
261
262 Args:
263 content: Markdown text containing images
264 path_prefix: Optional prefix to add to image paths
265 replace: Whether to replace image syntax in content
266
267 Returns:
268 Tuple of (processed_text, list_of_image_paths)
269
270 Example:
271 >>> util = MarkdownImageUtil()
272 >>> text, images = util.extract_image("![logo](img/logo.png)")
273 >>> print(images)
274 ['img/logo.png']
275 """
276 # List to store extracted image paths
277 images: List[str] = []
278
279 def repl(match: Match[str]) -> str:
280 """Replacement function for each image match."""
281 title = match.group(1) # Alt text
282 image_path = match.group(2) # Image path
283
284 # Add prefix if specified

Callers 2

_self_testMethod · 0.85
__init__Method · 0.85

Calls

no outgoing calls

Tested by

no test coverage detected