MCPcopy Create free account
hub / github.com/nlweb-ai/NLWeb / main

Function main

AskAgent/python/scraping/expBackOffCrawl.py:335–363  ·  view source on GitHub ↗
()

Source from the content-addressed store, hash-verified

333 self.stats.print_failure_summary()
334
335def main():
336 if len(sys.argv) < 3:
337 print("Usage: python crawlUrls.py <input_file> <target_directory> [max_retries]")
338 sys.exit(1)
339
340 input_file = sys.argv[1]
341 target_dir = sys.argv[2]
342
343 # Parse additional arguments
344 max_retries = 3 # Default
345
346 # Check for max_retries
347 if len(sys.argv) > 3:
348 max_retries = int(sys.argv[3])
349
350 # Read URLs
351 with open(input_file) as f:
352 urls = f.readlines()
353
354 # Create and run crawler
355 crawler = SimpleCrawler(
356 target_dir=target_dir,
357 max_retries=max_retries
358 )
359
360 print("Starting crawler with sequential processing")
361 print(f"Max retries: {max_retries}")
362
363 crawler.crawl_urls(urls)
364
365if __name__ == "__main__":
366 main()

Callers 1

expBackOffCrawl.pyFile · 0.70

Calls 2

crawl_urlsMethod · 0.95
SimpleCrawlerClass · 0.85

Tested by

no test coverage detected